system

The system addresses the challenge of unified and efficient user inquiry response by integrating AI chatbots, avatars, and video calls, enhancing user interaction and satisfaction.

JP2026072856APending Publication Date: 2026-05-01SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-18
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Conventional systems fail to provide unified and efficient responses to user inquiries.

Method used

A system comprising a reception unit for receiving inquiries via text chat, an answering unit using an AI chatbot for text-based responses, a voice answering unit using an AI avatar for voice responses, and a call unit for video calls with real operators, leveraging natural language processing, speech synthesis, and WebRTC technology for efficient and diverse user interaction.

Benefits of technology

The system efficiently responds to user inquiries through multiple communication channels, providing quick, accurate, and personalized support, reducing operator burden and improving customer satisfaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026072856000001_ABST
    Figure 2026072856000001_ABST
Patent Text Reader

Abstract

The system according to this embodiment aims to respond to user inquiries in an efficient and diverse manner. [Solution] The system according to this embodiment comprises a reception unit, an answering unit, a voice answering unit, and a call unit. The reception unit receives inquiries from users via text chat. The answering unit uses an AI chatbot to provide answers to inquiries received by the reception unit. The voice answering unit uses an AI avatar to provide voice answers to inquiries received by the reception unit. The call unit conducts video calls with real operators in response to inquiries received by the reception unit.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance that responds to the user utterance.

Prior Art Documents

Patent Documents

[0003] [[ID=2l]]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In the conventional technology, there is a problem that the response to inquiries from users is not unified and efficient support is not provided.

[0005] The system according to the embodiment aims to respond to inquiries from users in an efficient and diverse manner.

Means for Solving the Problems

[0006] The system according to this embodiment comprises a reception unit, an answering unit, a voice answering unit, and a call unit. The reception unit receives inquiries from users via text chat. The answering unit uses an AI chatbot to provide answers to inquiries received by the reception unit. The voice answering unit uses an AI avatar to provide voice answers to inquiries received by the reception unit. The call unit conducts video calls with real operators in response to inquiries received by the reception unit. [Effects of the Invention]

[0007] The system according to this embodiment can respond to user inquiries in an efficient and diverse manner. [Brief explanation of the drawing]

[0008] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10]This shows an emotion map where multiple emotions are mapped. [Modes for carrying out the invention]

[0009] Hereinafter, an example of an embodiment of the system relating to the technology of this disclosure will be described with reference to the attached drawings.

[0010] First, let's explain the terminology used in the following explanation.

[0011] In the following embodiments, the signed processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Furthermore, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose computing on Graphics Processing Units), APU (Accelerated Processing Unit), or TPU (Tensor Processing Unit).

[0012] In the following embodiments, signed RAM (Random Access Memory) is a memory that temporarily stores information and is used as work memory by the processor.

[0013] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.

[0014] In the following embodiments, the numbered communication I / F (Interface) is an interface including a communication processor, an antenna, etc. The communication I / F manages communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0015] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B". That is, "A and / or B" means that it may be only A, only B, or a combination of A and B. Also, in this specification, when expressing three or more matters connected by "and / or", the same concept as "A and / or B" is applied.

[0016] [First Embodiment] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0017] As shown in FIG. 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0018] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. Also, the database 24 and the communication I / F 26 are connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0019] The smart device 14 comprises a computer 36, a receiving device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The receiving device 38, output device 40, and camera 42 are also connected to the bus 52.

[0020] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, and accepts user input. The touch panel 38A accepts user input via touch by detecting contact with an object (e.g., a pen or finger). The microphone 38B accepts user input via voice by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 (see Figure 2) acquires the data indicating the user input.

[0021] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user by outputting the data in a form perceptible to the user (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0022] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0023] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0024] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0025] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.

[0026] In the smart device 14, specific processing is performed by the processor 46. The storage 50 stores a specific processing program 60. The specific processing program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the specific processing program 60 from the storage 50 and executes the read specific processing program 60 on the RAM 48. The specific processing is realized by the processor 46 operating as a control unit 46A according to the specific processing program 60 executed on the RAM 48. The smart device 14 also has a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.

[0027] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device (e.g., a generation server) may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device having the data generation model 58. The data processing device 12 may also be a server device or a terminal device owned by a user (e.g., a mobile phone, robot, home appliance, etc.). Next, an example of processing by the data processing system 10 according to the first embodiment will be described.

[0028] (Example of form 1) The customer support system according to an embodiment of the present invention is a system that provides customer support through a smartphone application. This system is designed to be more convenient for users to use when making inquiries about services. Specifically, it has the following functions. First, as a chat function, users can make inquiries via text chat, and an AI chatbot will respond and provide answers to basic questions. Next, as a voice chat function, voice chat using an AI avatar is possible, reducing the burden on real operators. Users ask questions by voice, and the AI ​​avatar answers. Furthermore, as a video call function, video calls with real operators are possible as needed, and it can handle complex problems and situations requiring detailed explanations. This system is designed for smartphones and can be used on various devices. For example, it can be used on smartwatches and tablets. In the future, it is envisioned that dedicated communication devices (such as smart displays and conversational robots) will be developed to provide services to users who do not have smartphones. The introduction of this system is expected to improve the efficiency of customer support operations and alleviate the problem of labor shortages. In addition, since users can easily make inquiries at any time, an improvement in customer satisfaction is expected. As a result, the customer support system will be able to efficiently receive user inquiries and provide answers, voice answers, and video calls.

[0029] The customer support system according to this embodiment comprises a reception unit, an answering unit, a voice answering unit, and a call unit. The reception unit receives inquiries from users via text chat. User inquiries include, but are not limited to, text, voice, and video. The reception unit receives user inquiries, for example, through text chat. The reception unit can also receive user inquiries through voice chat or video calls. For example, the reception unit receives inquiries using web chat, SMS, messaging apps, etc. The answering unit provides answers to inquiries received by the reception unit using an AI chatbot. The AI ​​chatbot analyzes user inquiries using, for example, natural language processing technology and generates appropriate answers. For example, the AI ​​chatbot provides answers to user inquiries using machine learning algorithms. The AI ​​chatbot can also generate answers to user inquiries based on pre-trained data. For example, the AI ​​chatbot provides answers by referring to an FAQ database. The voice answering unit provides voice answers to inquiries received by the reception unit using an AI avatar. AI avatars can, for example, provide voice responses to user inquiries using speech synthesis technology. AI avatars can also provide visual feedback to users using animation technology. Furthermore, AI avatars can provide real-time voice responses to user voice inquiries. For example, AI avatars can analyze user voice inquiries using speech recognition technology and generate appropriate voice responses. The call unit conducts video calls with real operators in response to inquiries received by the reception unit. The call unit can implement video calls between users and real operators using, for example, WebRTC technology. For example, the call unit can provide high-quality video calls using video codecs. The call unit can also conduct two-way video calls between users and real operators. For example, the call unit can share screens in real time during video calls.As a result, the customer support system according to this embodiment can efficiently receive and respond to user inquiries, providing voice responses and video calls.

[0030] The reception desk accepts user inquiries via text chat. These inquiries may include, but are not limited to, text, voice, and video. The reception desk can accept user inquiries via text chat, voice chat, or video calls. For example, it may use web chat, SMS, or messaging apps. Specifically, the reception desk provides an interface accessible to users through a website or mobile app, making it easy for users to initiate an inquiry. In the case of text chat, users can type their questions into the chat window and receive real-time responses. In voice chat, users input their questions using a microphone, the system converts them to text using speech recognition technology, and routes them to the appropriate department. In the case of video calls, users can send a video feed using their camera and interact with an operator in real-time. Furthermore, the reception desk has the ability to automatically categorize user inquiries and route them to the appropriate department or person. For example, inquiries can be categorized into different categories such as product questions, technical support, and billing inquiries, allowing for prompt responses from the appropriate specialist. Furthermore, the reception desk can refer to the user's past inquiry history to provide more personalized support. This allows users to receive consistent support and improves the efficiency of problem resolution.

[0031] The response department uses an AI chatbot to provide answers to inquiries received by the reception department. The AI ​​chatbot analyzes user inquiries using natural language processing technology, for example, and generates appropriate responses. Specifically, the AI ​​chatbot analyzes the user's input text, understands keywords and context, and selects the optimal response. For example, if a user enters "I want to check the status of my order," the AI ​​chatbot will ask for the order number, retrieve the latest status from the system based on that information, and provide the answer. The AI ​​chatbot also uses machine learning algorithms to provide answers to user inquiries. For example, it can learn from past inquiry data and predict the optimal answer to similar questions. Furthermore, the AI ​​chatbot can generate answers to user inquiries based on previously learned data. For example, it can refer to an FAQ database to provide answers. This allows users to obtain quick and accurate answers. The AI ​​chatbot can collect user feedback and continuously improve the accuracy and quality of its answers. For example, it has a function to allow users to rate their satisfaction with the provided answers, and retrains for low-rated answers. The AI ​​chatbot is also multilingual, accommodating users who speak different languages. This allows us to provide consistent support to our global user base.

[0032] The voice response unit uses an AI avatar to provide voice responses to inquiries received by the reception unit. The AI ​​avatar, for example, uses speech synthesis technology to provide voice answers to user inquiries. Specifically, the AI ​​avatar analyzes the user's text and voice input to generate appropriate voice responses. For example, if a user asks "How do I use this product?", the AI ​​avatar uses speech recognition technology to analyze the content and speech synthesis technology to generate an answer. The AI ​​avatar also provides visual feedback to the user using animation technology. For example, the AI ​​avatar uses facial expressions and gestures to provide a user-friendly interface. Furthermore, the AI ​​avatar can provide real-time voice responses to user voice inquiries. For example, the AI ​​avatar uses speech recognition technology to analyze the user's voice inquiry and generate an appropriate voice response. This allows users to receive support in a natural conversational format. In addition, the voice response unit can analyze the user's voice data to understand their emotions and intentions. For example, if a user is dissatisfied, the unit can detect this emotion and provide a more attentive response. Furthermore, the voice response unit can learn the user's accent and speaking style, enabling it to provide more natural voice responses. This allows the voice response unit to provide more personalized support to the user, thereby improving satisfaction.

[0033] The call center conducts video calls with real operators in response to inquiries received by the reception department. The call center uses, for example, WebRTC technology to facilitate video calls between users and real operators. Specifically, if a user requests a video call, the call center provides high-quality video calls using WebRTC technology. For example, the call center provides high-quality video calls using video codecs. The call center can also conduct two-way video calls between users and real operators. For example, the call center can share screens in real time during video calls. This allows users to resolve issues while directly interacting with the operator. Furthermore, the call center has a video call recording function, allowing users to save call content for later reference. For example, important support sessions and training sessions can be recorded and reviewed later. The call center also provides a chat function during video calls, allowing users to ask additional questions via text. This enables users to receive effective support by combining multiple communication methods. The call center also has a function to monitor the user's network status and optimize the quality of video calls. For example, if network bandwidth decreases, it automatically adjusts the video resolution to maintain smooth calls. Furthermore, the call unit uses encryption technology to protect call data in order to safeguard user privacy. This allows the call unit to provide secure and reliable video calls, thereby improving user satisfaction.

[0034] The response section uses an AI chatbot to provide text-based answers to user inquiries. For example, the AI ​​chatbot might analyze user inquiries using natural language processing techniques and generate appropriate responses. Alternatively, the AI ​​chatbot could use machine learning algorithms to provide answers. Furthermore, the AI ​​chatbot could generate answers based on pre-trained data. For instance, the AI ​​chatbot might refer to an FAQ database to provide answers. This enables rapid text-based responses through the use of an AI chatbot.

[0035] The voice response unit provides voice responses to user voice inquiries using an AI avatar. For example, the voice response unit uses speech synthesis technology to provide voice responses to user inquiries. Alternatively, the voice response unit can use animation technology to provide visual feedback to the user. Furthermore, the voice response unit can provide real-time voice responses to user voice inquiries. For example, the voice response unit can use speech recognition technology to analyze the user's voice inquiry and generate an appropriate voice response. This enables rapid voice responses through the use of an AI avatar.

[0036] The call unit enables users to conduct video calls with real operators. For example, the call unit uses WebRTC technology to facilitate video calls between users and real operators. For example, the call unit provides high-quality video calls using video codecs. Furthermore, the call unit can conduct two-way video calls between users and real operators. For example, the call unit can share screens in real time during video calls. This enables video calls with real operators and allows for the resolution of complex issues.

[0037] The reception desk accepts inquiries from various devices such as smartphones, smartwatches, and tablets. For example, the reception desk can accept user inquiries using a smartphone. For example, the reception desk can also accept user inquiries using a smartwatch. Furthermore, the reception desk can accept user inquiries using a tablet. For example, the reception desk can accept user inquiries through a smartphone app. This improves user convenience by accepting inquiries from various devices.

[0038] The reception desk accepts inquiries from dedicated communication devices. For example, the reception desk may use a dedicated smart display to receive user inquiries. For example, the reception desk may also use a conversational robot to receive user inquiries. Furthermore, the reception desk may use devices specialized for specific purposes to receive user inquiries. For example, the reception desk may accept user inquiries through a dedicated communication device. This allows the system to serve users who do not have smartphones by accepting inquiries from dedicated communication devices.

[0039] The reception department analyzes the user's past inquiry history and selects the most suitable method of contact. For example, the reception department may prioritize suggesting inquiry methods that the user has frequently used in the past. For example, the reception department may prepare relevant information in advance based on the user's past inquiry content. Furthermore, the reception department may suggest the most suitable method of contact for a specific time of day based on the user's past inquiry history. For example, the reception department may use machine learning algorithms to analyze the user's past inquiry history. This allows the reception department to select the most suitable method of contact by analyzing past inquiry history.

[0040] The reception department filters inquiries based on the user's current contract details and usage history. For example, the reception department only accepts inquiries relevant to the user's contract details. For example, the reception department can also provide appropriate support considering the user's usage history. Furthermore, the reception department can select inquiries to prioritize based on the user's contract details and usage history. For example, the reception department uses machine learning algorithms to filter inquiries based on the user's contract details and usage history. This allows for the provision of appropriate support by filtering inquiries based on the user's contract details and usage history.

[0041] The reception desk prioritizes receiving inquiries based on their relevance, taking into account the user's geographical location. For example, if a user is in a specific region, the reception desk will prioritize inquiries related to that region. The reception desk can also provide optimal support based on the user's location. Furthermore, the reception desk can prioritize processing relevant inquiries by considering the user's geographical location. For instance, the reception desk uses machine learning algorithms to filter inquiries based on the user's geographical location. This allows the reception desk to prioritize receiving inquiries that are highly relevant by considering the user's geographical location.

[0042] The reception desk analyzes the user's social media activity when receiving inquiries and accepts relevant inquiries. For example, the reception desk can analyze the user's social media activity and prioritize receiving relevant inquiries. For example, the reception desk can provide appropriate support based on the user's social media posts. The reception desk can also prioritize processing relevant inquiries by considering the user's social media activity. For example, the reception desk uses machine learning algorithms to filter based on the user's social media activity. This allows the reception desk to accept relevant inquiries by analyzing the user's social media activity.

[0043] The response unit adjusts the level of detail in its response based on the importance of the inquiry. For example, it provides detailed answers to high-priority inquiries, and concise answers to low-priority inquiries. It can also provide answers with an appropriate level of detail depending on the importance of the inquiry. For example, it uses a machine learning algorithm to evaluate the importance of the inquiry. This allows it to provide appropriate answers by adjusting the level of detail based on the importance of the inquiry.

[0044] The response unit applies different response algorithms depending on the category of the inquiry when providing a response. For example, it applies a specialized response algorithm to technical inquiries. For example, it can also apply a simpler response algorithm to general inquiries. Furthermore, the response unit can select the most suitable response algorithm depending on the category of the inquiry. For example, it uses a machine learning algorithm to classify the category of the inquiry. This allows it to provide an appropriate answer by applying the most suitable response algorithm according to the category of the inquiry.

[0045] The response system prioritizes responses based on when the inquiry was submitted. For example, it provides a quick response to urgent inquiries. For example, it can also provide responses to regular inquiries with normal priority. Furthermore, it can provide responses with appropriate priority depending on when the inquiry was submitted. For example, the response system uses a machine learning algorithm to evaluate when an inquiry was submitted. This allows it to provide appropriate responses by prioritizing responses based on when the inquiry was submitted.

[0046] The response unit adjusts the order of responses based on the relevance of the inquiries. For example, it prioritizes responses to highly relevant inquiries. For example, it can also provide responses to less relevant inquiries in the normal order. Furthermore, it can provide responses in an appropriate order depending on the relevance of the inquiries. For example, the response unit uses machine learning algorithms to evaluate the relevance of inquiries. This allows it to provide appropriate responses by adjusting the order of responses based on the relevance of the inquiries.

[0047] The voice response unit adjusts the level of detail in the voice response based on the content of the inquiry. For example, it provides a detailed voice response for important inquiries. For example, it can also provide a concise voice response for general inquiries. Furthermore, the voice response unit can provide a voice response with an appropriate level of detail depending on the content of the inquiry. For example, the voice response unit uses a machine learning algorithm to evaluate the content of the inquiry. This allows it to provide an appropriate voice response by adjusting the level of detail in the voice response based on the content of the inquiry.

[0048] The voice response unit applies different voice response algorithms depending on the category of the inquiry when providing a voice response. For example, the voice response unit applies a specialized voice response algorithm to technical inquiries. For example, the voice response unit can also apply a simpler voice response algorithm to general inquiries. Furthermore, the voice response unit can select the most suitable voice response algorithm depending on the category of the inquiry. For example, the voice response unit uses a machine learning algorithm to classify the category of the inquiry. This allows for the application of the most suitable voice response algorithm according to the category of the inquiry, thereby providing an appropriate voice response.

[0049] The voice response unit determines the priority of voice responses based on when the inquiry was submitted. For example, it provides a rapid voice response to urgent inquiries. For example, it can also provide voice responses to regular inquiries with a normal priority. Furthermore, it can provide voice responses with an appropriate priority depending on when the inquiry was submitted. For example, the voice response unit uses a machine learning algorithm to evaluate when an inquiry was submitted. This allows it to provide an appropriate voice response by determining the priority of voice responses based on when the inquiry was submitted.

[0050] The voice response unit adjusts the order of voice responses based on the relevance of the inquiries. For example, it prioritizes providing voice responses to highly relevant inquiries. For example, it can also provide voice responses to less relevant inquiries in the normal order. Furthermore, the voice response unit can provide voice responses in an appropriate order depending on the relevance of the inquiries. For example, the voice response unit uses machine learning algorithms to evaluate the relevance of inquiries. This allows it to provide appropriate voice responses by adjusting the order of voice responses based on the relevance of the inquiries.

[0051] The call function adjusts the level of detail during video calls based on the content of the inquiry. For example, the call function provides detailed video calls for important inquiries. For example, the call function can also provide concise video calls for general inquiries. Furthermore, the call function can provide video calls with an appropriate level of detail depending on the content of the inquiry. For example, the call function uses machine learning algorithms to evaluate the content of inquiries. This allows for the provision of appropriate video calls by adjusting the level of detail based on the content of the inquiry.

[0052] The call function applies different call algorithms during video calls depending on the category of the inquiry. For example, it applies a specialized call algorithm to technical inquiries. For example, it can also apply a simpler call algorithm to general inquiries. Furthermore, the call function can select the optimal call algorithm depending on the category of the inquiry. For example, it uses a machine learning algorithm to classify the category of the inquiry. This allows for the application of the most appropriate call algorithm according to the category of the inquiry, thereby providing a suitable video call.

[0053] The call department prioritizes video calls based on when the inquiry was submitted. For example, it can provide a video call quickly for urgent inquiries. For example, it can also provide a video call with normal priority for regular inquiries. Furthermore, it can provide a video call with an appropriate priority depending on when the inquiry was submitted. For example, the call department uses a machine learning algorithm to evaluate when an inquiry was submitted. This allows it to provide an appropriate video call by prioritizing calls based on when the inquiry was submitted.

[0054] The call system adjusts the order of video calls based on the relevance of the inquiries. For example, the call system prioritizes video calls for highly relevant inquiries. For example, the call system may also provide video calls in the normal order for less relevant inquiries. Furthermore, the call system can provide video calls in an appropriate order depending on the relevance of the inquiries. For example, the call system may use machine learning algorithms to evaluate the relevance of inquiries. This allows for the provision of appropriate video calls by adjusting the order of calls based on the relevance of the inquiries.

[0055] The system according to the embodiment is not limited to the example described above, and various modifications are possible, for example, as follows.

[0056] A customer support system can analyze a user's past inquiry history and provide the most appropriate answer. For example, if a user has made a similar inquiry in the past, the support system can refer to that history and provide a quick response. Furthermore, if a user has previously requested a detailed explanation for a specific issue, the support system can use that information to provide a more detailed answer. In addition, it can suggest the most appropriate response method for a particular time of day based on the user's past inquiry history. This allows for the provision of more appropriate answers by leveraging past inquiry history.

[0057] Customer support systems can provide highly relevant answers by considering the user's geographical location. For example, if a user is in a specific region, the system can prioritize providing information relevant to that region. It can also provide optimal support based on the user's location. Furthermore, it can prioritize processing inquiries related to the user's geographical location. This allows for more appropriate answers to be provided by leveraging the user's geographical location.

[0058] A customer support system can analyze a user's social media activity and provide relevant answers. For example, it can analyze a user's social media activity and prioritize providing relevant information. It can also provide appropriate support based on a user's social media posts. Furthermore, it can prioritize processing relevant inquiries by considering the user's social media activity. In this way, by leveraging the user's social media activity, more appropriate answers can be provided.

[0059] The customer support system can filter inquiries based on the user's current contract details and usage history. For example, it can accept only inquiries relevant to the user's contract. It can also provide appropriate support considering the user's usage history. Furthermore, it can select inquiries to prioritize based on the user's contract details and usage history. In this way, appropriate support can be provided by filtering inquiries based on the user's contract details and usage history.

[0060] A customer support system can analyze a user's past inquiry history when receiving an inquiry and select the most appropriate method of handling it. For example, it can prioritize suggesting inquiry methods that the user has frequently used in the past. It can also prepare relevant information in advance based on the user's past inquiries. Furthermore, it can suggest the most suitable method of handling an inquiry for a specific time of day based on the user's past inquiry history. In this way, the system can select the most appropriate method of handling an inquiry by analyzing past inquiry history.

[0061] The following briefly describes the processing flow for example form 1.

[0062] Step 1: The reception desk receives user inquiries via text chat. User inquiries may include, but are not limited to, text, voice, and video. The reception desk can receive user inquiries via text chat, for example. The reception desk may also receive user inquiries via voice chat or video call. For example, the reception desk can receive inquiries using web chat, SMS, messaging apps, etc. Step 2: The response department uses an AI chatbot to provide answers to inquiries received by the reception department. The AI ​​chatbot analyzes user inquiries using, for example, natural language processing technology and generates appropriate answers. For example, the AI ​​chatbot provides answers to user inquiries using machine learning algorithms. The AI ​​chatbot can also generate answers to user inquiries based on pre-trained data. For example, the AI ​​chatbot provides answers by referring to an FAQ database. Step 3: The voice response unit uses an AI avatar to provide voice responses to inquiries received by the reception unit. The AI ​​avatar provides voice responses to user inquiries using, for example, speech synthesis technology. For example, the AI ​​avatar provides visual feedback to the user using animation technology. The AI ​​avatar can also provide voice responses to user voice inquiries in real time. For example, the AI ​​avatar analyzes the user's voice inquiry using speech recognition technology and generates an appropriate voice response. Step 4: The call center conducts a video call with a real operator in response to an inquiry received by the reception center. The call center can, for example, use WebRTC technology to enable video calls between the user and the real operator. For example, the call center can provide high-quality video calls using video codecs. The call center can also conduct two-way video calls between the user and the real operator. For example, the call center can share its screen in real time during a video call.

[0063] (Example of form 2) The customer support system according to an embodiment of the present invention is a system that provides customer support through a smartphone application. This system is designed to be more convenient for users to use when making inquiries about services. Specifically, it has the following functions. First, as a chat function, users can make inquiries via text chat, and an AI chatbot will respond and provide answers to basic questions. Next, as a voice chat function, voice chat using an AI avatar is possible, reducing the burden on real operators. Users ask questions by voice, and the AI ​​avatar answers. Furthermore, as a video call function, video calls with real operators are possible as needed, and it can handle complex problems and situations requiring detailed explanations. This system is designed for smartphones and can be used on various devices. For example, it can be used on smartwatches and tablets. In the future, it is envisioned that dedicated communication devices (such as smart displays and conversational robots) will be developed to provide services to users who do not have smartphones. The introduction of this system is expected to improve the efficiency of customer support operations and alleviate the problem of labor shortages. In addition, since users can easily make inquiries at any time, an improvement in customer satisfaction is expected. As a result, the customer support system will be able to efficiently receive user inquiries and provide answers, voice answers, and video calls.

[0064] The customer support system according to this embodiment comprises a reception unit, an answering unit, a voice answering unit, and a call unit. The reception unit receives inquiries from users via text chat. User inquiries include, but are not limited to, text, voice, and video. The reception unit receives user inquiries, for example, through text chat. The reception unit can also receive user inquiries through voice chat or video calls. For example, the reception unit receives inquiries using web chat, SMS, messaging apps, etc. The answering unit provides answers to inquiries received by the reception unit using an AI chatbot. The AI ​​chatbot analyzes user inquiries using, for example, natural language processing technology and generates appropriate answers. For example, the AI ​​chatbot provides answers to user inquiries using machine learning algorithms. The AI ​​chatbot can also generate answers to user inquiries based on pre-trained data. For example, the AI ​​chatbot provides answers by referring to an FAQ database. The voice answering unit provides voice answers to inquiries received by the reception unit using an AI avatar. AI avatars can, for example, provide voice responses to user inquiries using speech synthesis technology. AI avatars can also provide visual feedback to users using animation technology. Furthermore, AI avatars can provide real-time voice responses to user voice inquiries. For example, AI avatars can analyze user voice inquiries using speech recognition technology and generate appropriate voice responses. The call unit conducts video calls with real operators in response to inquiries received by the reception unit. The call unit can implement video calls between users and real operators using, for example, WebRTC technology. For example, the call unit can provide high-quality video calls using video codecs. The call unit can also conduct two-way video calls between users and real operators. For example, the call unit can share screens in real time during video calls.As a result, the customer support system according to this embodiment can efficiently receive and respond to user inquiries, providing voice responses and video calls.

[0065] The reception desk accepts user inquiries via text chat. These inquiries may include, but are not limited to, text, voice, and video. The reception desk can accept user inquiries via text chat, voice chat, or video calls. For example, it may use web chat, SMS, or messaging apps. Specifically, the reception desk provides an interface accessible to users through a website or mobile app, making it easy for users to initiate an inquiry. In the case of text chat, users can type their questions into the chat window and receive real-time responses. In voice chat, users input their questions using a microphone, the system converts them to text using speech recognition technology, and routes them to the appropriate department. In the case of video calls, users can send a video feed using their camera and interact with an operator in real-time. Furthermore, the reception desk has the ability to automatically categorize user inquiries and route them to the appropriate department or person. For example, inquiries can be categorized into different categories such as product questions, technical support, and billing inquiries, allowing for prompt responses from the appropriate specialist. Furthermore, the reception desk can refer to the user's past inquiry history to provide more personalized support. This allows users to receive consistent support and improves the efficiency of problem resolution.

[0066] The response department uses an AI chatbot to provide answers to inquiries received by the reception department. The AI ​​chatbot analyzes user inquiries using natural language processing technology, for example, and generates appropriate responses. Specifically, the AI ​​chatbot analyzes the user's input text, understands keywords and context, and selects the optimal response. For example, if a user enters "I want to check the status of my order," the AI ​​chatbot will ask for the order number, retrieve the latest status from the system based on that information, and provide the answer. The AI ​​chatbot also uses machine learning algorithms to provide answers to user inquiries. For example, it can learn from past inquiry data and predict the optimal answer to similar questions. Furthermore, the AI ​​chatbot can generate answers to user inquiries based on previously learned data. For example, it can refer to an FAQ database to provide answers. This allows users to obtain quick and accurate answers. The AI ​​chatbot can collect user feedback and continuously improve the accuracy and quality of its answers. For example, it has a function to allow users to rate their satisfaction with the provided answers, and retrains for low-rated answers. The AI ​​chatbot is also multilingual, accommodating users who speak different languages. This allows us to provide consistent support to our global user base.

[0067] The voice response unit uses an AI avatar to provide voice responses to inquiries received by the reception unit. The AI ​​avatar, for example, uses speech synthesis technology to provide voice answers to user inquiries. Specifically, the AI ​​avatar analyzes the user's text and voice input to generate appropriate voice responses. For example, if a user asks "How do I use this product?", the AI ​​avatar uses speech recognition technology to analyze the content and speech synthesis technology to generate an answer. The AI ​​avatar also provides visual feedback to the user using animation technology. For example, the AI ​​avatar uses facial expressions and gestures to provide a user-friendly interface. Furthermore, the AI ​​avatar can provide real-time voice responses to user voice inquiries. For example, the AI ​​avatar uses speech recognition technology to analyze the user's voice inquiry and generate an appropriate voice response. This allows users to receive support in a natural conversational format. In addition, the voice response unit can analyze the user's voice data to understand their emotions and intentions. For example, if a user is dissatisfied, the unit can detect this emotion and provide a more attentive response. Furthermore, the voice response unit can learn the user's accent and speaking style, enabling it to provide more natural voice responses. This allows the voice response unit to provide more personalized support to the user, thereby improving satisfaction.

[0068] The call center conducts video calls with real operators in response to inquiries received by the reception department. The call center uses, for example, WebRTC technology to facilitate video calls between users and real operators. Specifically, if a user requests a video call, the call center provides high-quality video calls using WebRTC technology. For example, the call center provides high-quality video calls using video codecs. The call center can also conduct two-way video calls between users and real operators. For example, the call center can share screens in real time during video calls. This allows users to resolve issues while directly interacting with the operator. Furthermore, the call center has a video call recording function, allowing users to save call content for later reference. For example, important support sessions and training sessions can be recorded and reviewed later. The call center also provides a chat function during video calls, allowing users to ask additional questions via text. This enables users to receive effective support by combining multiple communication methods. The call center also has a function to monitor the user's network status and optimize the quality of video calls. For example, if network bandwidth decreases, it automatically adjusts the video resolution to maintain smooth calls. Furthermore, the call unit uses encryption technology to protect call data in order to safeguard user privacy. This allows the call unit to provide secure and reliable video calls, thereby improving user satisfaction.

[0069] The response section uses an AI chatbot to provide text-based answers to user inquiries. For example, the AI ​​chatbot might analyze user inquiries using natural language processing techniques and generate appropriate responses. Alternatively, the AI ​​chatbot could use machine learning algorithms to provide answers. Furthermore, the AI ​​chatbot could generate answers based on pre-trained data. For instance, the AI ​​chatbot might refer to an FAQ database to provide answers. This enables rapid text-based responses through the use of an AI chatbot.

[0070] The voice response unit provides voice responses to user voice inquiries using an AI avatar. For example, the voice response unit uses speech synthesis technology to provide voice responses to user inquiries. Alternatively, the voice response unit can use animation technology to provide visual feedback to the user. Furthermore, the voice response unit can provide real-time voice responses to user voice inquiries. For example, the voice response unit can use speech recognition technology to analyze the user's voice inquiry and generate an appropriate voice response. This enables rapid voice responses through the use of an AI avatar.

[0071] The call unit enables users to conduct video calls with real operators. For example, the call unit uses WebRTC technology to facilitate video calls between users and real operators. For example, the call unit provides high-quality video calls using video codecs. Furthermore, the call unit can conduct two-way video calls between users and real operators. For example, the call unit can share screens in real time during video calls. This enables video calls with real operators and allows for the resolution of complex issues.

[0072] The reception desk accepts inquiries from various devices such as smartphones, smartwatches, and tablets. For example, the reception desk can accept user inquiries using a smartphone. For example, the reception desk can also accept user inquiries using a smartwatch. Furthermore, the reception desk can accept user inquiries using a tablet. For example, the reception desk can accept user inquiries through a smartphone app. This improves user convenience by accepting inquiries from various devices.

[0073] The reception desk accepts inquiries from dedicated communication devices. For example, the reception desk may use a dedicated smart display to receive user inquiries. For example, the reception desk may also use a conversational robot to receive user inquiries. Furthermore, the reception desk may use devices specialized for specific purposes to receive user inquiries. For example, the reception desk may accept user inquiries through a dedicated communication device. This allows the system to serve users who do not have smartphones by accepting inquiries from dedicated communication devices.

[0074] The reception desk estimates the user's emotions and prioritizes inquiries based on those emotions. For example, if the user is stressed, the reception desk will process the inquiry with the highest priority. For example, if the user is relaxed, the reception desk may process the inquiry with the normal priority. The reception desk can also raise the priority of inquiries to respond quickly if the user is in a hurry. For example, the reception desk may use facial recognition technology to estimate the user's emotions. For example, the reception desk may also use voice analysis technology to estimate the user's emotions. This allows for more appropriate responses by prioritizing inquiries based on the user's emotions. Emotion estimation is achieved using emotion estimation functions, such as emotion engines or generative AI. Generative AI includes, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI.

[0075] The reception department analyzes the user's past inquiry history and selects the most suitable method of contact. For example, the reception department may prioritize suggesting inquiry methods that the user has frequently used in the past. For example, the reception department may prepare relevant information in advance based on the user's past inquiry content. Furthermore, the reception department may suggest the most suitable method of contact for a specific time of day based on the user's past inquiry history. For example, the reception department may use machine learning algorithms to analyze the user's past inquiry history. This allows the reception department to select the most suitable method of contact by analyzing past inquiry history.

[0076] The reception department filters inquiries based on the user's current contract details and usage history. For example, the reception department only accepts inquiries relevant to the user's contract details. For example, the reception department can also provide appropriate support considering the user's usage history. Furthermore, the reception department can select inquiries to prioritize based on the user's contract details and usage history. For example, the reception department uses machine learning algorithms to filter inquiries based on the user's contract details and usage history. This allows for the provision of appropriate support by filtering inquiries based on the user's contract details and usage history.

[0077] The reception desk estimates the user's emotions and adjusts the timing of the reception based on the estimated emotions. For example, if the user is stressed, the reception desk will process the user quickly. For example, if the user is relaxed, the reception desk can process the user at a normal time. The reception desk can also process the user immediately if the user is in a hurry. For example, the reception desk may use facial recognition technology to estimate the user's emotions. For example, the reception desk may use voice analysis technology to estimate the user's emotions. This allows for more appropriate timing of the reception by adjusting the timing based on the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI includes, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI.

[0078] The reception desk prioritizes receiving inquiries based on their relevance, taking into account the user's geographical location. For example, if a user is in a specific region, the reception desk will prioritize inquiries related to that region. The reception desk can also provide optimal support based on the user's location. Furthermore, the reception desk can prioritize processing relevant inquiries by considering the user's geographical location. For instance, the reception desk uses machine learning algorithms to filter inquiries based on the user's geographical location. This allows the reception desk to prioritize receiving inquiries that are highly relevant by considering the user's geographical location.

[0079] The reception desk analyzes the user's social media activity when receiving inquiries and accepts relevant inquiries. For example, the reception desk can analyze the user's social media activity and prioritize receiving relevant inquiries. For example, the reception desk can provide appropriate support based on the user's social media posts. The reception desk can also prioritize processing relevant inquiries by considering the user's social media activity. For example, the reception desk uses machine learning algorithms to filter based on the user's social media activity. This allows the reception desk to accept relevant inquiries by analyzing the user's social media activity.

[0080] The response unit estimates the user's emotions and adjusts the way it expresses its response based on those emotions. For example, if the user is stressed, the response unit provides a concise and easy-to-understand response. For example, if the user is relaxed, the response unit may provide a response that includes detailed explanations. Also, if the user is in a hurry, the response unit may provide a quick and to-the-point response. For example, the response unit may use facial recognition technology to estimate the user's emotions. For example, the response unit may use speech analysis technology to estimate the user's emotions. This allows for the provision of more appropriate responses by adjusting the way the response is expressed based on the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI includes, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI.

[0081] The response unit adjusts the level of detail in its response based on the importance of the inquiry. For example, it provides detailed answers to high-priority inquiries, and concise answers to low-priority inquiries. It can also provide answers with an appropriate level of detail depending on the importance of the inquiry. For example, it uses a machine learning algorithm to evaluate the importance of the inquiry. This allows it to provide appropriate answers by adjusting the level of detail based on the importance of the inquiry.

[0082] The response unit applies different response algorithms depending on the category of the inquiry when providing a response. For example, it applies a specialized response algorithm to technical inquiries. For example, it can also apply a simpler response algorithm to general inquiries. Furthermore, the response unit can select the most suitable response algorithm depending on the category of the inquiry. For example, it uses a machine learning algorithm to classify the category of the inquiry. This allows it to provide an appropriate answer by applying the most suitable response algorithm according to the category of the inquiry.

[0083] The response unit estimates the user's emotions and adjusts the length of the response based on the estimated emotions. For example, if the user is stressed, the response unit will provide a short, to-the-point response. For example, if the user is relaxed, the response unit may provide a longer response with more detailed explanations. Also, if the user is in a hurry, the response unit may provide a quick, to-the-point short response. For example, the response unit may use facial recognition technology to estimate the user's emotions. For example, the response unit may use speech analysis technology to estimate the user's emotions. This allows for the provision of more appropriate responses by adjusting the length of the response based on the user's emotions. Emotion estimation is achieved using emotion estimation functions, such as emotion engines or generative AI. Generative AI includes, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI.

[0084] The response system prioritizes responses based on when the inquiry was submitted. For example, it provides a quick response to urgent inquiries. For example, it can also provide responses to regular inquiries with normal priority. Furthermore, it can provide responses with appropriate priority depending on when the inquiry was submitted. For example, the response system uses a machine learning algorithm to evaluate when an inquiry was submitted. This allows it to provide appropriate responses by prioritizing responses based on when the inquiry was submitted.

[0085] The response unit adjusts the order of responses based on the relevance of the inquiries. For example, it prioritizes responses to highly relevant inquiries. For example, it can also provide responses to less relevant inquiries in the normal order. Furthermore, it can provide responses in an appropriate order depending on the relevance of the inquiries. For example, the response unit uses machine learning algorithms to evaluate the relevance of inquiries. This allows it to provide appropriate responses by adjusting the order of responses based on the relevance of the inquiries.

[0086] The voice response unit estimates the user's emotions and adjusts the tone and speed of the voice response based on the estimated emotions. For example, if the user is stressed, the voice response unit will respond slowly in a calm tone. For example, if the user is relaxed, the voice response unit can respond in a bright tone and at a normal speed. Also, if the user is in a hurry, the voice response unit can respond quickly in a rapid tone. For example, the voice response unit may use facial recognition technology to estimate the user's emotions. For example, the voice response unit may also use voice analysis technology to estimate the user's emotions. This allows for the provision of more appropriate voice responses by adjusting the tone and speed of the voice response based on the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI includes, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI.

[0087] The voice response unit adjusts the level of detail in the voice response based on the content of the inquiry. For example, it provides a detailed voice response for important inquiries. For example, it can also provide a concise voice response for general inquiries. Furthermore, the voice response unit can provide a voice response with an appropriate level of detail depending on the content of the inquiry. For example, the voice response unit uses a machine learning algorithm to evaluate the content of the inquiry. This allows it to provide an appropriate voice response by adjusting the level of detail in the voice response based on the content of the inquiry.

[0088] The voice response unit applies different voice response algorithms depending on the category of the inquiry when providing a voice response. For example, the voice response unit applies a specialized voice response algorithm to technical inquiries. For example, the voice response unit can also apply a simpler voice response algorithm to general inquiries. Furthermore, the voice response unit can select the most suitable voice response algorithm depending on the category of the inquiry. For example, the voice response unit uses a machine learning algorithm to classify the category of the inquiry. This allows for the application of the most suitable voice response algorithm according to the category of the inquiry, thereby providing an appropriate voice response.

[0089] The voice response unit estimates the user's emotions and adjusts the order of voice responses based on the estimated emotions. For example, if the user is stressed, the voice response unit will provide a voice response with the highest priority. For example, if the user is relaxed, the voice response unit can provide voice responses in the normal order. Also, if the user is in a hurry, the voice response unit can provide voice responses quickly. For example, the voice response unit may use facial recognition technology to estimate the user's emotions. For example, the voice response unit may also use voice analysis technology to estimate the user's emotions. This allows for the provision of more appropriate voice responses by adjusting the order of voice responses based on the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI.

[0090] The voice response unit determines the priority of voice responses based on when the inquiry was submitted. For example, it provides a rapid voice response to urgent inquiries. For example, it can also provide voice responses to regular inquiries with a normal priority. Furthermore, it can provide voice responses with an appropriate priority depending on when the inquiry was submitted. For example, the voice response unit uses a machine learning algorithm to evaluate when an inquiry was submitted. This allows it to provide an appropriate voice response by determining the priority of voice responses based on when the inquiry was submitted.

[0091] The voice response unit adjusts the order of voice responses based on the relevance of the inquiries. For example, it prioritizes providing voice responses to highly relevant inquiries. For example, it can also provide voice responses to less relevant inquiries in the normal order. Furthermore, the voice response unit can provide voice responses in an appropriate order depending on the relevance of the inquiries. For example, the voice response unit uses machine learning algorithms to evaluate the relevance of inquiries. This allows it to provide appropriate voice responses by adjusting the order of voice responses based on the relevance of the inquiries.

[0092] The call unit estimates the user's emotions and adjusts the timing of the video call based on the estimated emotions. For example, if the user is feeling stressed, the call unit will start the video call quickly. For example, if the user is relaxed, the call unit can start the video call at the normal time. Also, if the user is in a hurry, the call unit can start the video call immediately. For example, the call unit may use facial recognition technology to estimate the user's emotions. For example, the call unit may use voice analysis technology to estimate the user's emotions. This allows for more appropriate timing of the video call by adjusting the timing of the video call based on the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI.

[0093] The call function adjusts the level of detail during video calls based on the content of the inquiry. For example, the call function provides detailed video calls for important inquiries. For example, the call function can also provide concise video calls for general inquiries. Furthermore, the call function can provide video calls with an appropriate level of detail depending on the content of the inquiry. For example, the call function uses machine learning algorithms to evaluate the content of inquiries. This allows for the provision of appropriate video calls by adjusting the level of detail based on the content of the inquiry.

[0094] The call function applies different call algorithms during video calls depending on the category of the inquiry. For example, it applies a specialized call algorithm to technical inquiries. For example, it can also apply a simpler call algorithm to general inquiries. Furthermore, the call function can select the optimal call algorithm depending on the category of the inquiry. For example, it uses a machine learning algorithm to classify the category of the inquiry. This allows for the application of the most appropriate call algorithm according to the category of the inquiry, thereby providing a suitable video call.

[0095] The call unit estimates the user's emotions and adjusts the order of video calls based on the estimated emotions. For example, if the user is stressed, the call unit will provide the video call with the highest priority. For example, if the user is relaxed, the call unit may provide the video call in the normal order. The call unit can also provide the video call quickly if the user is in a hurry. For example, the call unit may use facial recognition technology to estimate the user's emotions. For example, the call unit may use voice analysis technology to estimate the user's emotions. This allows for more appropriate video calls to be provided by adjusting the order of video calls based on the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI.

[0096] The call department prioritizes video calls based on when the inquiry was submitted. For example, it can provide a video call quickly for urgent inquiries. For example, it can also provide a video call with normal priority for regular inquiries. Furthermore, it can provide a video call with an appropriate priority depending on when the inquiry was submitted. For example, the call department uses a machine learning algorithm to evaluate when an inquiry was submitted. This allows it to provide an appropriate video call by prioritizing calls based on when the inquiry was submitted.

[0097] The call system adjusts the order of video calls based on the relevance of the inquiries. For example, the call system prioritizes video calls for highly relevant inquiries. For example, the call system may also provide video calls in the normal order for less relevant inquiries. Furthermore, the call system can provide video calls in an appropriate order depending on the relevance of the inquiries. For example, the call system may use machine learning algorithms to evaluate the relevance of inquiries. This allows for the provision of appropriate video calls by adjusting the order of calls based on the relevance of the inquiries.

[0098] The system according to the embodiment is not limited to the example described above, and various modifications are possible, for example, as follows.

[0099] Customer support systems can estimate a user's emotions and adjust the tone of their responses based on that estimation. For example, if a user is stressed, the system can provide a calm response. If the user is relaxed, the system can provide a cheerful response. Furthermore, if the user is in a hurry, the system can provide a quick response. This allows for responses in an appropriate tone according to the user's emotions, which is expected to improve user satisfaction.

[0100] A customer support system can analyze a user's past inquiry history and provide the most appropriate answer. For example, if a user has made a similar inquiry in the past, the support system can refer to that history and provide a quick response. Furthermore, if a user has previously requested a detailed explanation for a specific issue, the support system can use that information to provide a more detailed answer. In addition, it can suggest the most appropriate response method for a particular time of day based on the user's past inquiry history. This allows for the provision of more appropriate answers by leveraging past inquiry history.

[0101] A customer support system can estimate a user's emotions and adjust the level of detail in its response based on that estimation. For example, if a user is stressed, the system can provide a concise and to-the-point answer. If the user is relaxed, the system can provide a more detailed explanation. Furthermore, if the user is in a hurry, the system can provide a quick and concise answer. This allows for responses with the appropriate level of detail according to the user's emotions, which is expected to improve user satisfaction.

[0102] Customer support systems can provide highly relevant answers by considering the user's geographical location. For example, if a user is in a specific region, the system can prioritize providing information relevant to that region. It can also provide optimal support based on the user's location. Furthermore, it can prioritize processing inquiries related to the user's geographical location. This allows for more appropriate answers to be provided by leveraging the user's geographical location.

[0103] A customer support system can estimate a user's emotions and adjust the order of responses based on those emotions. For example, if a user is stressed, a response can be provided as a top priority. If the user is relaxed, responses can be provided in the normal order. Furthermore, if the user is in a hurry, a response can be provided quickly. This allows for responses to be provided in an appropriate order according to the user's emotions, which is expected to improve user satisfaction.

[0104] A customer support system can analyze a user's social media activity and provide relevant answers. For example, it can analyze a user's social media activity and prioritize providing relevant information. It can also provide appropriate support based on a user's social media posts. Furthermore, it can prioritize processing relevant inquiries by considering the user's social media activity. In this way, by leveraging the user's social media activity, more appropriate answers can be provided.

[0105] The customer support system can estimate the user's emotions and adjust the tone and speed of the voice response based on those estimates. For example, if the user is stressed, the voice response unit can respond in a calm tone and slowly. If the user is relaxed, the voice response unit can respond in a cheerful tone and at a normal speed. Furthermore, if the user is in a hurry, the voice response unit can respond quickly in a rapid tone. This allows for voice responses with appropriate tone and speed according to the user's emotions, which is expected to improve user satisfaction.

[0106] The customer support system can filter inquiries based on the user's current contract details and usage history. For example, it can accept only inquiries relevant to the user's contract. It can also provide appropriate support considering the user's usage history. Furthermore, it can select inquiries to prioritize based on the user's contract details and usage history. In this way, appropriate support can be provided by filtering inquiries based on the user's contract details and usage history.

[0107] The customer support system can estimate the user's emotions and adjust the timing of video call initiation based on those emotions. For example, if the user is feeling stressed, the video call can be started quickly. If the user is relaxed, the video call can be started at the usual time. Furthermore, if the user is in a hurry, the video call can be started immediately. This enables video calls at the appropriate time according to the user's emotions, which is expected to improve user satisfaction.

[0108] A customer support system can analyze a user's past inquiry history when receiving an inquiry and select the most appropriate method of handling it. For example, it can prioritize suggesting inquiry methods that the user has frequently used in the past. It can also prepare relevant information in advance based on the user's past inquiries. Furthermore, it can suggest the most suitable method of handling an inquiry for a specific time of day based on the user's past inquiry history. In this way, the system can select the most appropriate method of handling an inquiry by analyzing past inquiry history.

[0109] The following briefly describes the processing flow for example form 2.

[0110] Step 1: The reception desk receives user inquiries via text chat. User inquiries may include, but are not limited to, text, voice, and video. The reception desk can receive user inquiries via text chat, for example. The reception desk may also receive user inquiries via voice chat or video call. For example, the reception desk can receive inquiries using web chat, SMS, messaging apps, etc. Step 2: The response department uses an AI chatbot to provide answers to inquiries received by the reception department. The AI ​​chatbot analyzes user inquiries using, for example, natural language processing technology and generates appropriate answers. For example, the AI ​​chatbot provides answers to user inquiries using machine learning algorithms. The AI ​​chatbot can also generate answers to user inquiries based on pre-trained data. For example, the AI ​​chatbot provides answers by referring to an FAQ database. Step 3: The voice response unit uses an AI avatar to provide voice responses to inquiries received by the reception unit. The AI ​​avatar provides voice responses to user inquiries using, for example, speech synthesis technology. For example, the AI ​​avatar provides visual feedback to the user using animation technology. The AI ​​avatar can also provide voice responses to user voice inquiries in real time. For example, the AI ​​avatar analyzes the user's voice inquiry using speech recognition technology and generates an appropriate voice response. Step 4: The call center conducts a video call with a real operator in response to an inquiry received by the reception center. The call center can, for example, use WebRTC technology to enable video calls between the user and the real operator. For example, the call center can provide high-quality video calls using video codecs. The call center can also conduct two-way video calls between the user and the real operator. For example, the call center can share its screen in real time during a video call.

[0111] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0112] Data generation model 58 is a form of so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> Examples of generative AI include text generation AI, image generation AI, and multimodal generation AI. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats from audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVMs), k-means clustering, convolutional neural networks (CNNs), recurrent neural networks (RNNs), generative adversarial networks (GANs), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI ​​may be an AI agent. Furthermore, when the processing of each of the above parts is performed by the AI, the processing may be performed by the AI ​​in part or in whole, but is not limited to this example.Furthermore, processing performed by AI, including generative AI, may be replaced with rule-based processing, and rule-based processing may be replaced with processing performed by AI, including generative AI.

[0113] Furthermore, the processing performed by the data processing system 10 described above is carried out by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart device 14, but it may also be carried out by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart device 14. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the smart device 14 or an external device, and the smart device 14 acquires or collects information necessary for processing from the data processing device 12 or an external device.

[0114] Each of the multiple elements described above, including the reception unit, response unit, voice response unit, and call unit, is implemented in at least one of the smart device 14 and the data processing device 12. For example, the reception unit is implemented by the control unit 46A of the smart device 14 and receives user inquiries via text chat, voice chat, or video call. The response unit is implemented by the specific processing unit 290 of the data processing device 12 and provides answers to user inquiries using an AI chatbot. The voice response unit is implemented by the control unit 46A of the smart device 14 and provides voice answers using an AI avatar. The call unit is implemented by the specific processing unit 290 of the data processing device 12 and conducts video calls with a real operator using WebRTC technology. The correspondence between each unit and the device or control unit is not limited to the examples described above and can be modified in various ways.

[0115] [Second Embodiment] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0116] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0117] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.

[0118] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0119] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0120] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).

[0121] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0122] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing by the processor 28. The storage 32 stores the specific processing program 56.

[0123] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0124] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.

[0125] In the smart glasses 214, specific processing is performed by the processor 46. The storage 50 stores a specific processing program 60. The processor 46 reads the specific processing program 60 from the storage 50 and executes the read specific processing program 60 on the RAM 48. The specific processing is realized by the processor 46 acting as a control unit 46A according to the specific processing program 60 executed on the RAM 48. The smart glasses 214 also have a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.

[0126] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).

[0127] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0128] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI ​​may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI ​​in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.

[0129] The data processing system 210 according to the second embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 210 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart glasses 214, but it may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart glasses 214. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the smart glasses 214 or an external device, and the smart glasses 214 acquires or collects information necessary for processing from the data processing device 12 or an external device.

[0130] Each of the multiple elements described above, including the reception unit, response unit, voice response unit, and call unit, is implemented in at least one of the smart glasses 214 and the data processing device 12. For example, the reception unit is implemented by the control unit 46A of the smart glasses 214 and receives user inquiries via text chat, voice chat, or video call. The response unit is implemented by the specific processing unit 290 of the data processing device 12 and provides answers to user inquiries using an AI chatbot. The voice response unit is implemented by the control unit 46A of the smart glasses 214 and provides voice answers using an AI avatar. The call unit is implemented by the specific processing unit 290 of the data processing device 12 and conducts video calls with a real operator using WebRTC technology. The correspondence between each unit and the device or control unit is not limited to the examples described above and can be modified in various ways.

[0131] [Third Embodiment] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0132] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0133] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.

[0134] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0135] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0136] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).

[0137] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0138] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0139] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0140] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.

[0141] In the headset terminal 314, specific processing is performed by the processor 46. The storage 50 stores a specific program 60. The processor 46 reads the specific program 60 from the storage 50 and executes the read specific program 60 on the RAM 48. The specific processing is realized by the processor 46 acting as a control unit 46A according to the specific program 60 executed on the RAM 48. The headset terminal 314 also has a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.

[0142] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).

[0143] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0144] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI ​​may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI ​​in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.

[0145] The data processing system 310 according to the third embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 310 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the headset terminal 314, but may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the headset terminal 314. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the headset terminal 314 or an external device, and the headset terminal 314 acquires or collects information necessary for processing from the data processing device 12 or an external device.

[0146] Each of the multiple elements described above, including the reception unit, response unit, voice response unit, and call unit, is implemented in at least one of the following: the headset terminal 314 and the data processing unit 12. For example, the reception unit is implemented by the control unit 46A of the headset terminal 314 and receives user inquiries via text chat, voice chat, or video call. The response unit is implemented by the specific processing unit 290 of the data processing unit 12 and provides answers to user inquiries using an AI chatbot. The voice response unit is implemented by the control unit 46A of the headset terminal 314 and provides voice answers using an AI avatar. The call unit is implemented by the specific processing unit 290 of the data processing unit 12 and conducts video calls with a real operator using WebRTC technology. The correspondence between each unit and the device or control unit is not limited to the examples described above and can be modified in various ways.

[0147] [Fourth Embodiment] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0148] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0149] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.

[0150] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0151] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0152] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS image sensor or CCD image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).

[0153] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0154] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. The robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0155] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0156] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0157] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.

[0158] In robot 414, specific processing is performed by processor 46. A specific program 60 is stored in storage 50. Processor 46 reads the specific program 60 from storage 50 and executes it on RAM 48. The specific processing is achieved by processor 46 acting as a control unit 46A according to the specific program 60 executed on RAM 48. Robot 414 also has data generation model 58 and emotion identification model 59, similar to those of the robot, and can perform processing similar to that of the specific processing unit 290 using these models.

[0159] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).

[0160] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0161] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI ​​may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI ​​in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.

[0162] The data processing system 410 according to the fourth embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 410 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the robot 414, but it may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the robot 414. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the robot 414 or an external device, and the robot 414 acquires or collects information necessary for processing from the data processing device 12 or an external device.

[0163] Each of the multiple elements described above, including the reception unit, response unit, voice response unit, and call unit, is implemented by, for example, at least one of the robot 414 and the data processing unit 12. For example, the reception unit is implemented by the control unit 46A of the robot 414 and receives user inquiries via text chat, voice chat, or video call. The response unit is implemented by, for example, the specific processing unit 290 of the data processing unit 12 and provides answers to user inquiries using an AI chatbot. The voice response unit is implemented by, for example, the control unit 46A of the robot 414 and provides voice answers using an AI avatar. The call unit is implemented by, for example, the specific processing unit 290 of the data processing unit 12 and conducts video calls with a real operator using WebRTC technology. The correspondence between each unit and the device or control unit is not limited to the examples described above and can be modified in various ways.

[0164] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0165] Figure 9 shows the emotion map 400, in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0166] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0167] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0168] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, and motorcycles, emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated based, for example, on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0169] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0170] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0171] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing method for the specific process may be used, which includes multiple computers, including computer 22.

[0172] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0173] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0174] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0175] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0176] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0177] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0178] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0179] Furthermore, although the above-described examples were divided into four embodiments, some or all of these embodiments may be combined. Also, the smart device 14, smart glasses 214, headset terminal 314, and robot 414 are just examples, and they may be combined, or other devices may be used. Also, although the above-described examples were divided into two embodiments, Embodiment 1 and Embodiment 2, these may be combined.

[0180] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and other things that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0181] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.

[0182] (Note 1) A reception department that handles user inquiries via text chat, The response unit provides answers to inquiries received by the reception unit via an AI chatbot, The voice response unit provides voice responses to inquiries received by the aforementioned reception unit, and The system includes a call unit that conducts video calls with real operators in response to inquiries received by the reception unit. A system characterized by the following features. (Note 2) The aforementioned response section is, Using an AI chatbot to provide text-based answers to user inquiries. The system described in Appendix 1, characterized by the features described herein. (Note 3) The aforementioned voice response unit is Using an AI avatar, it provides voice responses to user voice inquiries. The system described in Appendix 1, characterized by the features described herein. (Note 4) The aforementioned communication unit is, Enables users to have video calls with real operators. The system described in Appendix 1, characterized by the features described herein. (Note 5) The aforementioned reception unit is We accept inquiries from various devices such as smartphones, smartwatches, and tablets. The system described in Appendix 1, characterized by the features described herein. (Note 6) The aforementioned reception unit is We accept inquiries via a dedicated communication device. The system described in Appendix 1, characterized by the features described herein. (Note 7) The aforementioned reception unit is It estimates the user's emotions and prioritizes inquiries based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 8) The aforementioned reception unit is Analyze the user's past inquiry history and select the most suitable contact method. The system described in Appendix 1, characterized by the features described herein. (Note 9) The aforementioned reception unit is When receiving an inquiry, filtering is performed based on the user's current contract details and usage status. The system described in Appendix 1, characterized by the features described herein. (Note 10) The aforementioned reception unit is The system estimates the user's emotions and adjusts the timing of the reception based on those emotions. The system described in Appendix 1, characterized by the features described herein. (Note 11) The aforementioned reception unit is When receiving inquiries, the system prioritizes inquiries that are highly relevant, taking into account the user's geographical location. The system described in Appendix 1, characterized by the features described herein. (Note 12) The aforementioned reception unit is When receiving an inquiry, the system analyzes the user's social media activity and selects relevant inquiries. The system described in Appendix 1, characterized by the features described herein. (Note 13) The aforementioned response section is, It estimates the user's emotions and adjusts the way responses are expressed based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 14) The aforementioned response section is, When responding, adjust the level of detail in your response based on the importance of the inquiry. The system described in Appendix 1, characterized by the features described herein. (Note 15) The aforementioned response section is, When responding, different response algorithms are applied depending on the category of the inquiry. The system described in Appendix 1, characterized by the features described herein. (Note 16) The aforementioned response section is, It estimates the user's emotions and adjusts the length of the response based on the estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 17) The aforementioned response section is, When responding, we will prioritize responses based on when the inquiry was submitted. The system described in Appendix 1, characterized by the features described herein. (Note 18) The aforementioned response section is, When responding, we adjust the order of responses based on the relevance of the inquiries. The system described in Appendix 1, characterized by the features described herein. (Note 19) The aforementioned voice response unit is It estimates the user's emotions and adjusts the tone and speed of the voice response based on the estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 20) The aforementioned voice response unit is When providing a voice response, adjust the level of detail in the voice response based on the content of the inquiry. The system described in Appendix 1, characterized by the features described herein. (Note 21) The aforementioned voice response unit is When providing voice responses, different voice response algorithms are applied depending on the category of the inquiry. The system described in Appendix 1, characterized by the features described herein. (Note 22) The aforementioned voice response unit is The system estimates the user's emotions and adjusts the order of voice responses based on the estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 23) The aforementioned voice response unit is When providing an audio response, we prioritize the audio response based on when the inquiry was submitted. The system described in Appendix 1, characterized by the features described herein. (Note 24) The aforementioned voice response unit is When providing voice responses, the order of voice responses will be adjusted based on the relevance of the inquiry. The system described in Appendix 1, characterized by the features described herein. (Note 25) The aforementioned communication unit is, It estimates the user's emotions and adjusts the timing of the video call start based on the estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 26) The aforementioned communication unit is, During video calls, the level of detail of the call will be adjusted based on the content of the inquiry. The system described in Appendix 1, characterized by the features described herein. (Note 27) The aforementioned communication unit is, During video calls, different call algorithms are applied depending on the category of the inquiry. The system described in Appendix 1, characterized by the features described herein. (Note 28) The aforementioned communication unit is, It estimates the user's emotions and adjusts the order of video calls based on the estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 29) The aforementioned communication unit is, During video calls, call priorities are determined based on when the inquiry was submitted. The system described in Appendix 1, characterized by the features described herein. (Note 30) The aforementioned communication unit is, During video calls, the order of calls will be adjusted based on the relevance of the inquiries. The system described in Appendix 1, characterized by the features described herein. [Explanation of Symbols]

[0183] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots

Claims

1. A reception department that handles user inquiries via text chat, The response unit provides answers to inquiries received by the reception unit via an AI chatbot, The voice response unit provides voice responses to inquiries received by the aforementioned reception unit, and The system includes a call unit that conducts video calls with real operators in response to inquiries received by the reception unit. A system characterized by the following features.

2. The aforementioned response section is, Using an AI chatbot, we provide text-based answers to user inquiries. The system according to feature 1.

3. The aforementioned voice response unit is Using an AI avatar, we provide voice responses to user voice inquiries. The system according to feature 1.

4. The aforementioned communication unit is, Enables users to have video calls with real operators. The system according to feature 1.

5. The aforementioned reception unit is We accept inquiries from various devices such as smartphones, smartwatches, and tablets. The system according to feature 1.

6. The aforementioned reception unit is We accept inquiries via a dedicated communication device. The system according to feature 1.

7. The aforementioned reception unit is It estimates the user's emotions and prioritizes inquiries based on those estimated emotions. The system according to feature 1.

8. The aforementioned reception unit is Analyze the user's past inquiry history and select the most suitable contact method. The system according to feature 1.

9. The aforementioned reception unit is When receiving an inquiry, filtering is performed based on the user's current contract details and usage status. The system according to feature 1.

10. The aforementioned reception unit is The system estimates the user's emotions and adjusts the timing of the reception based on those emotions. The system according to feature 1.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A