system

The system addresses the lack of effective avatar generation by collecting and analyzing user messages to create avatars that engage in dialogue, ensuring user characteristics are reflected and enabling appropriate responses, particularly in emergencies.

JP7852004B2Active Publication Date: 2026-04-27SOFTBANK GROUP CORP
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-09-19
Publication Date
2026-04-27

AI Technical Summary

Technical Problem

Conventional technologies have not sufficiently addressed the generation of avatars based on user messages and communication exchanges, limiting their ability to engage in meaningful conversations.

Method used

A system comprising a collection unit, analysis unit, and generation unit that collects user messages, analyzes user characteristics, and generates avatars using AI to engage in dialogue on behalf of the user, employing natural language processing and text generation to reflect the user's speaking style and expressions.

Benefits of technology

The system effectively generates avatars that reflect user characteristics, enabling them to engage in conversations and take appropriate actions on behalf of the user, particularly in emergencies, thereby enhancing user peace of mind and daily life continuity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007852004000001
    Figure 0007852004000001
  • Figure 0007852004000002
    Figure 0007852004000002
  • Figure 0007852004000003
    Figure 0007852004000003
Patent Text Reader

Abstract

To generate avatars based on user messages and call exchanges by the system in accordance with the embodiment, and to interact with them.SOLUTION: A system according to the embodiment has a collection unit, an analysis unit, a generation unit, and a dialogue unit. The collection unit collects user messages or call exchanges. The analysis unit analyzes the information collected by the collection unit and extracts user characteristics. The generation unit generates avatars based on the features extracted by the analysis unit. The dialog unit is where avatars generated by the generation unit interact with each other.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In the conventional technology, avatar generation based on user messages and communication exchanges has not been sufficiently performed, and there is room for improvement.

[0005] The system according to the embodiment aims to generate an avatar based on user messages and communication exchanges and conduct conversations.

Means for Solving the Problems

[0006] The system according to this embodiment comprises a collection unit, an analysis unit, a generation unit, and a dialogue unit. The collection unit collects user messages or call exchanges. The analysis unit analyzes the information collected by the collection unit and extracts user characteristics. The generation unit generates an avatar based on the characteristics extracted by the analysis unit. The dialogue unit allows the avatar generated by the generation unit to engage in dialogue. [Effects of the Invention]

[0007] The system according to this embodiment can generate an avatar based on the user's messages and call exchanges, and engage in dialogue with it. [Brief explanation of the drawing]

[0008] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Modes for carrying out the invention]

[0009] Hereinafter, an example of an embodiment of the system relating to the technology of this disclosure will be described with reference to the attached drawings.

[0010] First, let's explain the terminology used in the following explanation.

[0011] In the following embodiments, the signed processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Furthermore, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose computing on Graphics Processing Units), APU (Accelerated Processing Unit), or TPU (Tensor Processing Unit).

[0012] In the following embodiments, signed RAM (Random Access Memory) is a memory that temporarily stores information and is used as work memory by the processor.

[0013] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.

[0014] In the following embodiments, the labeled communication I / F (Interface) is an interface including a communication processor, an antenna, etc. The communication I / F manages communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0015] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B". That is, "A and / or B" means that it may be only A, only B, or a combination of A and B. Also, in this specification, when expressing three or more matters connected by "and / or", the same concept as "A and / or B" is applied.

[0016] [First Embodiment] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0017] As shown in FIG. 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0018] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. Also, the database 24 and the communication I / F 26 are connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0019] The smart device 14 comprises a computer 36, a receiving device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The receiving device 38, output device 40, and camera 42 are also connected to the bus 52.

[0020] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, and accepts user input. The touch panel 38A accepts user input via touch by detecting contact with an object (e.g., a pen or finger). The microphone 38B accepts user input via voice by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 (see Figure 2) acquires the data indicating the user input.

[0021] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user by outputting the data in a form perceptible to the user (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0022] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0023] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0024] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0025] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.

[0026] In the smart device 14, specific processing is performed by the processor 46. The storage 50 stores a specific processing program 60. The specific processing program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the specific processing program 60 from the storage 50 and executes the read specific processing program 60 on the RAM 48. The specific processing is realized by the processor 46 operating as a control unit 46A according to the specific processing program 60 executed on the RAM 48. The smart device 14 also has a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.

[0027] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device (e.g., a generation server) may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device having the data generation model 58. The data processing device 12 may also be a server device or a terminal device owned by a user (e.g., a mobile phone, robot, home appliance, etc.). Next, an example of processing by the data processing system 10 according to the first embodiment will be described.

[0028] (Example of form 1) The interactive avatar system according to an embodiment of the present invention is a system that collects user messages and call exchanges, and a generating AI analyzes the user's characteristics to generate an anthropomorphic avatar. This interactive avatar system can generate an interactive avatar that reflects the user's characteristics and can engage in conversations on behalf of the user in times of emergency. For example, it collects user messages and call exchanges, and the generating AI analyzes the user's speaking style and expression characteristics. Next, the generating AI generates an avatar based on the analyzed characteristics, and this avatar engages in conversations on behalf of the user. This allows the user to live their daily life with peace of mind. Thus, the interactive avatar system can generate an avatar that reflects the user's characteristics and can engage in conversations on behalf of the user in times of emergency.

[0029] The interactive avatar system according to this embodiment comprises a collection unit, an analysis unit, a generation unit, and a dialogue unit. The collection unit collects user messages or call exchanges. The collection unit can collect data such as text messages, voice calls, and video calls. The collection unit collects the content of messages and calls sent by the user in detail and provides this data to the analysis unit. The analysis unit analyzes the collected data and extracts user characteristics. The analysis unit analyzes the characteristics of the user's speech and expression using, for example, natural language processing technology. The analysis unit extracts words, phrases, and speaking tone that the user frequently uses. The generation unit generates an avatar based on the characteristics extracted by the analysis unit. The generation unit generates an avatar that reflects the user's characteristics using a generation AI. The generation AI generates the avatar's dialogue content using, for example, a text generation AI (e.g., LLM). The generation unit can generate 3D models, 2D characters, voice-only avatars, etc., that reflect the user's characteristics. The dialogue unit has the generated avatar engage in dialogue on behalf of the user. The dialogue unit conducts conversations using methods such as text chat, voice chat, and video chat. The dialogue unit can use generative AI to conduct conversations that reflect the user's characteristics. As a result, the dialogue avatar system according to this embodiment can generate an avatar that reflects the user's characteristics and conduct conversations on behalf of the user in times of emergency. For example, the dialogue unit can have the generated avatar conduct a conversation on behalf of the user and take appropriate action. This allows the user to live their daily life with peace of mind.

[0030] The data collection unit collects user messages and call interactions. For example, the collection unit can collect data such as text messages, voice calls, and video calls. Specifically, in the case of text messages, it acquires the text information sent by the user in real time; in the case of voice calls, it records the call content and stores it as audio data; and in the case of video calls, it collects both video and audio to obtain detailed information such as the user's facial expressions and tone of voice. This data is centrally managed by the data collection unit and provided to the analysis unit. The data collection unit can use encryption technology to protect privacy when collecting data. For example, collected data is encrypted at the time of collection to prevent unauthorized access by third parties until it is transmitted to the analysis unit. Furthermore, the data collection unit has a mechanism to collect data only after obtaining the user's consent, and can stop collection if the user refuses data collection. This allows the data collection unit to efficiently collect necessary data while protecting user privacy. In addition, the data collection unit is equipped with high-speed data transfer technology to provide collected data to the analysis unit in real time, thereby improving the overall system response speed.

[0031] The analysis unit analyzes the collected data and extracts user characteristics. For example, the analysis unit uses natural language processing technology to analyze the characteristics of the user's speech and expression. Specifically, in the case of text messages, it performs morphological and grammatical analysis to identify words and phrases that the user frequently uses. In the case of voice calls, it uses speech recognition technology to convert voice data into text and analyzes that text data. It can also extract features such as voice tone, speaking speed, and emotional expression from the voice data. In the case of video calls, it uses video analysis technology to analyze the user's facial expressions and gestures, and infers the user's emotions and intentions based on this information. The analysis unit integrates these analysis results to comprehensively understand the user's characteristics. Furthermore, the analysis unit can learn user behavior patterns and preferences by utilizing past data and user history information. For example, based on past conversation history, it can predict how the user will react in specific situations and provide information to realize more natural conversations. In addition, the analysis unit can use anomaly detection algorithms to detect unusual user behavior or statements and issue warnings as needed. This allows the analysis unit to accurately understand user characteristics and improve the overall dialogue quality of the system.

[0032] The generation unit generates avatars based on features extracted by the analysis unit. The generation unit uses a generation AI to generate avatars that reflect the user's characteristics. Specifically, it uses a text generation AI (e.g., LLM) to generate the avatar's dialogue. The generation AI learns the user's speaking style and expressive characteristics and can generate natural dialogue based on that. For example, it can generate dialogue that incorporates words and phrases frequently used by the user and create avatar statements that resemble the user's speaking style. The generation unit can also generate 3D models, 2D characters, and voice-only avatars that reflect the user's characteristics. In the case of 3D models, it can generate avatars that reproduce the user's appearance and facial expressions and can operate them in real time. In the case of 2D characters, it generates illustrations and animations that reflect the user's characteristics and displays them during dialogue. In the case of voice-only avatars, it uses speech synthesis technology that reflects the characteristics of the user's voice to achieve natural voice dialogue. The generation unit provides these generation results to the dialogue unit, preparing it to engage in dialogue on behalf of the user. Furthermore, the generation unit can evaluate the quality of the generated avatars and make corrections or improvements as needed. For example, if the generated dialogue is unnatural or does not accurately reflect the user's characteristics, it is regenerated to improve quality. This allows the generation unit to produce high-quality avatars that accurately reflect the user's characteristics, thereby improving the overall dialogue quality of the system.

[0033] The dialogue unit uses a generated avatar to interact on behalf of the user. The dialogue unit engages in conversations using methods such as text chat, voice chat, and video chat. Specifically, in the case of text chat, it sends text messages generated by a generative AI on behalf of the user and engages in conversation with the other party. In the case of voice chat, it plays voice generated using speech synthesis technology and engages in conversation on behalf of the user. In the case of video chat, it displays a generated 3D model or 2D character and engages in conversation while it operates in real time. The dialogue unit can use generative AI to conduct conversations that reflect the user's characteristics. For example, it can generate dialogue content incorporating words and phrases frequently used by the user and create avatar statements that resemble the user's speaking style. Furthermore, the dialogue unit can incorporate new dialogue content generated by the generative AI in real time as the conversation progresses. This enables the dialogue unit to achieve natural and smooth conversations and respond appropriately on behalf of the user. In addition, the dialogue unit can record the content and progress of the conversation and provide this information to the analysis and generation units later. This allows for continuous improvement of the overall dialogue quality of the system through feedback. For example, if the content of the dialogue is inappropriate or does not accurately reflect the user's characteristics, the analysis and generation units can correct or improve it based on that information. This allows the dialogue unit to conduct high-quality dialogue on behalf of the user, improving the overall reliability and satisfaction of the system.

[0034] The collection unit can collect user messages or call content. For example, the collection unit can collect data such as text messages, voice calls, and video calls. The collection unit meticulously collects the content of messages and calls sent by the user and provides this data to the analysis unit. This allows for the understanding of user characteristics by collecting user messages and call content. Some or all of the above processing in the collection unit may be performed using AI, or without AI. For example, the collection unit can input user messages and call content into an AI, which can then analyze the data.

[0035] The analysis unit can analyze the collected data and extract characteristics of the user's speech or expression. For example, the analysis unit may use natural language processing techniques to analyze the characteristics of the user's speech and expression. The analysis unit may extract words, phrases, and tone of voice that the user frequently uses. By extracting characteristics of the user's speech and expression, it is possible to generate an avatar that reflects the user's characteristics. Some or all of the above processing in the analysis unit may be performed using AI, for example, or without AI. For example, the analysis unit may input the collected data into an AI, which can then analyze the data.

[0036] The generation unit can generate avatars based on extracted features. The generation unit uses a generation AI to generate avatars that reflect the user's characteristics. The generation AI uses, for example, a text generation AI (e.g., LLM) to generate the avatar's dialogue. The generation unit can generate 3D models, 2D characters, or voice-only avatars that reflect the user's characteristics. In this way, by generating avatars based on extracted features, it is possible to provide avatars that reflect the user's characteristics. Some or all of the above-described processes in the generation unit may be performed using, for example, a generation AI, or without using a generation AI. For example, the generation unit can input extracted features into a generation AI, and the generation AI can generate an avatar.

[0037] The dialogue unit allows a generated avatar to engage in dialogue on behalf of the user. The dialogue unit engages in dialogue using methods such as text chat, voice dialogue, and video dialogue. The dialogue unit can use a generation AI to engage in dialogue that reflects the user's characteristics. This allows the generated avatar to engage in dialogue on behalf of the user, enabling appropriate responses in times of emergency. Some or all of the above-described processes in the dialogue unit may be performed using AI, or not using AI. For example, the dialogue unit can input the generated avatar into an AI, which can then engage in dialogue.

[0038] The dialogue unit can use algorithms to respond in the event of an emergency. For example, the dialogue unit can take appropriate action in the event of an emergency using emergency response scenarios and the AI ​​technology it employs. This allows it to take appropriate action on behalf of the user by using algorithms designed for appropriate responses in emergencies. Some or all of the above-described processes in the dialogue unit may be performed using AI, or not using AI. For example, the dialogue unit can input emergency response scenarios into the AI, which can then take action.

[0039] The data collection unit can analyze the user's past messages and call history and select a collection method. For example, the data collection unit can prioritize collecting data from messaging and calling apps that the user has frequently used in the past. Furthermore, the data collection unit can concentrate data collection on specific time periods based on the user's past messages and call history. In addition, the data collection unit can analyze the user's past communication patterns and select the most effective collection method. This allows the optimal collection method to be selected by analyzing the user's past messages and call history. Some or all of the above processing in the data collection unit may be performed using AI, or not. For example, the data collection unit can input the user's past messages and call history into an AI, which can then select a collection method.

[0040] The data collection unit can filter messages and call content based on the user's current lifestyle and areas of interest. For example, if the user is at work, the data collection unit can prioritize collecting work-related messages and call content. Similarly, if the user is engaged in a hobby, the data collection unit can collect messages and call content related to that hobby. Furthermore, if the user is traveling, the data collection unit can collect travel-related messages and call content. This allows for the collection of more relevant data by filtering based on the user's lifestyle and areas of interest. Some or all of the processing described above in the data collection unit may be performed using AI, or not. For example, the data collection unit can input the user's lifestyle and areas of interest into an AI, which can then perform the filtering.

[0041] The data collection unit can prioritize the collection of highly relevant content based on the user's geographical location when collecting messages and call content. For example, if the user is in a specific region, the data collection unit can prioritize the collection of messages and call content related to that region. Furthermore, if the user is traveling, the data collection unit can prioritize the collection of messages and call content related to their travel destination. Additionally, if the user is at home, the data collection unit can prioritize the collection of messages and call content related to their home. This allows for the priority collection of highly relevant data by considering the user's geographical location. Some or all of the above processing in the data collection unit may be performed using AI, or without AI. For example, the data collection unit can input the user's geographical location information into AI, which can then prioritize the collection of highly relevant content.

[0042] The data collection unit can analyze the user's social media activity and collect relevant content when collecting messages and call content. For example, the data collection unit can prioritize collecting messages and call content with people the user frequently interacts with on social media. It can also collect messages and call content related to topics the user shows interest in on social media. Furthermore, the data collection unit can collect messages and call content related to groups and communities the user participates in on social media. In this way, relevant content can be collected by analyzing the user's social media activity. Some or all of the above processing in the data collection unit may be performed using AI, for example, or not. For example, the data collection unit can input the user's social media activity into AI, which can then collect relevant content.

[0043] The analysis unit can adjust the level of detail of the analysis based on the importance of the message or call content during the analysis. For example, the analysis unit can perform a detailed analysis of high-importance messages or call content to extract features down to the smallest detail. Conversely, the analysis unit can perform a simplified analysis of low-importance messages or call content to extract only the main features. Furthermore, the analysis unit can perform a rapid analysis of urgent messages or call content to immediately extract features. In this way, by adjusting the level of detail of the analysis based on the importance of the message or call content, the analysis can be performed efficiently. Some or all of the above processing in the analysis unit may be performed using AI, for example, or without AI. For example, the analysis unit can input the importance of the message or call content into the AI, and the AI ​​can adjust the level of detail of the analysis.

[0044] The analysis unit can apply different analysis algorithms depending on the category of the message or call content during analysis. For example, the analysis unit can apply an algorithm that analyzes formal expressions and technical terms to business-related messages or call content. It can also apply an algorithm that analyzes casual expressions and everyday language to private messages or call content. Furthermore, the analysis unit can apply an algorithm for rapid response to urgent messages or call content. By applying different analysis algorithms depending on the category of the message or call content, more accurate analysis can be performed. Some or all of the above-described processes in the analysis unit may be performed using AI, for example, or without AI. For example, the analysis unit can input the category of the message or call content into the AI, which can then apply different analysis algorithms.

[0045] The analysis unit can determine the priority of analysis based on the submission timing of messages and call content during the analysis process. For example, the analysis unit can prioritize the analysis of recently sent messages and call content. It can also prioritize the analysis of messages and call content sent during a specific time period. Furthermore, it can prioritize the analysis of messages and call content of high urgency. This allows for efficient analysis by determining the priority of analysis based on the submission timing of messages and call content. Some or all of the above-described processes in the analysis unit may be performed using AI, for example, or without AI. For example, the analysis unit can input the submission timing of messages and call content into the AI, which can then determine the priority of analysis.

[0046] The analysis unit can adjust the order of analysis based on the relevance of messages and call content during the analysis process. For example, the analysis unit can prioritize the analysis of highly relevant messages and call content. It can also postpone the analysis of less relevant messages and call content. Furthermore, the analysis unit can prioritize the analysis of messages and call content related to specific topics. This allows for efficient analysis by adjusting the order of analysis based on the relevance of messages and call content. Some or all of the above-described processes in the analysis unit may be performed using AI, for example, or without AI. For example, the analysis unit can input the relevance of messages and call content into the AI, which can then adjust the order of analysis.

[0047] The generation unit can adjust the level of detail of the avatar based on the importance of the user's features during avatar generation. For example, if the user's features are clear, the generation unit can generate a detailed avatar. If the user's features are ambiguous, the generation unit can generate a simplified avatar. Furthermore, if the user's features are diverse, the generation unit can generate an avatar that emphasizes the main features. By adjusting the level of detail of the avatar based on the importance of the user's features, a more appropriate avatar can be generated. Some or all of the above processing in the generation unit may be performed using AI, for example, or without AI. For example, the generation unit can input the importance of the user's features into the AI, and the AI ​​can adjust the level of detail of the avatar.

[0048] The generation unit can apply different generation algorithms depending on the user's category when generating avatars. For example, the generation unit can apply a generation algorithm with formal expressions and actions to avatars for business use. It can also apply a generation algorithm with casual expressions and actions to avatars for private use. Furthermore, the generation unit can apply a generation algorithm for rapid response to avatars in emergencies. By applying different generation algorithms according to the user's category, a more appropriate avatar can be generated. Some or all of the above-described processes in the generation unit may be performed using AI, for example, or without AI. For example, the generation unit can input the user's category into the AI, which can then apply different generation algorithms.

[0049] The generation unit can determine the generation priority based on the user's past avatar usage history when generating avatars. For example, the generation unit can prioritize the reflection of features of avatars that the user has frequently used in the past. Furthermore, the generation unit can predict and generate avatars that the user will use at specific times based on their past avatar usage history. In addition, the generation unit can analyze the user's past avatar usage patterns and generate the most effective avatars. This allows for the generation of more appropriate avatars by determining the generation priority based on the user's past avatar usage history. Some or all of the above-described processes in the generation unit may be performed using AI, or not. For example, the generation unit can input the user's past avatar usage history into an AI, which can then determine the generation priority.

[0050] The generation unit can improve the accuracy of avatar generation by referring to the user's relevant data. For example, the generation unit can refer to the user's social media profile and posts to adjust the avatar's appearance and personality. It can also refer to the user's past messages and call content to adjust the avatar's speaking style and expressions. Furthermore, the generation unit can customize the avatar's characteristics based on the user's hobbies and interests. This allows for improved generation accuracy by referring to the user's relevant data. Some or all of the above-described processes in the generation unit may be performed using AI, for example, or without AI. For example, the generation unit can input the user's relevant data into the AI, which can then improve the generation accuracy.

[0051] The dialogue unit can select a dialogue method by referring to the user's past dialogue history during a conversation. For example, the dialogue unit can incorporate expressions and phrases that the user has previously used frequently into the conversation. Furthermore, the dialogue unit can select a dialogue method for a specific topic based on the user's past dialogue history. In addition, the dialogue unit can analyze the user's past dialogue patterns and select the most effective dialogue method. This allows the optimal dialogue method to be selected by referring to the user's past dialogue history. Some or all of the above processing in the dialogue unit may be performed using AI, or not. For example, the dialogue unit can input the user's past dialogue history into an AI, which can then select a dialogue method.

[0052] The dialogue unit can customize the means of dialogue based on the user's current situation during a conversation. For example, if the user is at work, the dialogue unit can engage in a businesslike conversation. If the user is relaxed, the dialogue unit can engage in a casual conversation. Furthermore, if the user is in an emergency, the dialogue unit can engage in a quick and concise conversation. In this way, by customizing the means of dialogue based on the user's current situation, a more appropriate conversation can be conducted. Some or all of the above processing in the dialogue unit may be performed using AI, for example, or without AI. For example, the dialogue unit can input the user's current situation into the AI, and the AI ​​can customize the means of dialogue.

[0053] The dialogue unit can select the optimal dialogue method during a conversation by considering the user's geographical location information. For example, if the user is in a specific region, the dialogue unit can provide information related to that region. If the user is traveling, the dialogue unit can provide information related to the travel destination. Furthermore, if the user is at home, the dialogue unit can provide information related to home. In this way, the optimal dialogue method can be selected by considering the user's geographical location information. Some or all of the above processing in the dialogue unit may be performed using AI, for example, or without AI. For example, the dialogue unit can input the user's geographical location information into the AI, and the AI ​​can select the optimal dialogue method.

[0054] The dialogue unit can analyze the user's social media activity during a conversation and suggest a suitable dialogue method. For example, the dialogue unit can prioritize conversations with people the user frequently interacts with on social media. It can also conduct conversations related to topics the user has shown interest in on social media. Furthermore, it can conduct conversations related to groups and communities the user participates in on social media. In this way, by analyzing the user's social media activity, it can suggest the most suitable dialogue method. Some or all of the above processing in the dialogue unit may be performed using AI, for example, or without AI. For example, the dialogue unit can input the user's social media activity into AI, which can then suggest a dialogue method.

[0055] The system according to the embodiment is not limited to the example described above, and various modifications are possible, for example, as follows.

[0056] The interactive avatar system can also collect user health data and reflect it in the avatar's dialogue. For example, the data collection unit can collect heart rate and sleep data from the user's wearable device. The analysis unit can analyze this data to understand the user's health status. The generation unit can have the avatar provide health advice based on the user's health status. The dialogue unit can engage in dialogue to promote relaxation or encourage exercise, depending on the user's health status. This allows for more personalized support by providing dialogue that takes the user's health status into consideration.

[0057] The data collection unit can collect user environmental data and reflect it in the avatar's dialogue. For example, the data collection unit can collect environmental data such as temperature, humidity, and noise level around the user. The analysis unit can analyze this data to understand the user's environmental conditions. The generation unit can provide advice to help the avatar maintain a comfortable environment based on the user's environmental conditions. The dialogue unit can engage in appropriate dialogue according to the user's environmental conditions. In this way, by engaging in dialogue that takes the user's environmental conditions into consideration, it is possible to support a more comfortable life.

[0058] The analysis unit can analyze the user's hobbies and interests and reflect them in the avatar's dialogue. For example, the analysis unit can analyze data on keywords the user has searched for in the past and websites they have visited. Based on this data, the generation unit can have the avatar provide topics related to the user's hobbies and interests. The dialogue unit can suggest relevant information and activities according to the user's hobbies and interests. This allows for more interesting dialogue by taking the user's hobbies and interests into consideration.

[0059] The generation unit can customize the avatar's dialogue by referring to the user's past dialogue history. For example, the generation unit can incorporate expressions and phrases that the user has previously used frequently into the avatar's dialogue. The dialogue unit can select a dialogue method for a specific topic from the user's past dialogue history. Furthermore, the generation unit can analyze the user's past dialogue patterns and select the most effective dialogue method. This allows for more personalized dialogue by referring to the user's past dialogue history.

[0060] The dialogue unit can customize the avatar's conversation content by taking into account the user's geographical location. For example, if the user is in a specific region, the dialogue unit can provide conversations that offer information related to that region. If the user is traveling, the dialogue unit can provide conversations that offer information related to their travel destination. Furthermore, if the user is at home, the dialogue unit can provide conversations that offer information related to their home. This allows for more relevant conversations by considering the user's geographical location.

[0061] The following briefly describes the processing flow for example form 1.

[0062] Step 1: The collection unit collects user messages or call interactions. The collection unit can collect data such as text messages, voice calls, and video calls. The collection unit collects detailed information about the messages and calls sent by the user and provides this data to the analysis unit. Step 2: The analysis unit analyzes the collected data and extracts user characteristics. For example, the analysis unit uses natural language processing technology to analyze the characteristics of the user's speech and expressions. The analysis unit extracts words, phrases, and speaking tone that the user frequently uses. Step 3: The generation unit generates an avatar based on the features extracted by the analysis unit. The generation unit uses a generation AI to generate an avatar that reflects the user's characteristics. The generation AI generates the avatar's dialogue using, for example, a text generation AI (e.g., LLM). The generation unit can generate 3D models, 2D characters, or voice-only avatars that reflect the user's characteristics. Step 4: The dialogue unit has the generated avatar interact on behalf of the user. The dialogue unit interacts using methods such as text chat, voice dialogue, or video dialogue. The dialogue unit can use a generation AI to perform dialogue that reflects the user's characteristics. As a result, the dialogue avatar system according to the embodiment can generate an avatar that reflects the user's characteristics and perform dialogue on behalf of the user in times of emergency.

[0063] (Example of form 2) The interactive avatar system according to an embodiment of the present invention is a system that collects user messages and call exchanges, and a generating AI analyzes the user's characteristics to generate an anthropomorphic avatar. This interactive avatar system can generate an interactive avatar that reflects the user's characteristics and can engage in conversations on behalf of the user in times of emergency. For example, it collects user messages and call exchanges, and the generating AI analyzes the user's speaking style and expression characteristics. Next, the generating AI generates an avatar based on the analyzed characteristics, and this avatar engages in conversations on behalf of the user. This allows the user to live their daily life with peace of mind. Thus, the interactive avatar system can generate an avatar that reflects the user's characteristics and can engage in conversations on behalf of the user in times of emergency.

[0064] The interactive avatar system according to this embodiment comprises a collection unit, an analysis unit, a generation unit, and a dialogue unit. The collection unit collects user messages or call exchanges. The collection unit can collect data such as text messages, voice calls, and video calls. The collection unit collects the content of messages and calls sent by the user in detail and provides this data to the analysis unit. The analysis unit analyzes the collected data and extracts user characteristics. The analysis unit analyzes the characteristics of the user's speech and expression using, for example, natural language processing technology. The analysis unit extracts words, phrases, and speaking tone that the user frequently uses. The generation unit generates an avatar based on the characteristics extracted by the analysis unit. The generation unit generates an avatar that reflects the user's characteristics using a generation AI. The generation AI generates the avatar's dialogue content using, for example, a text generation AI (e.g., LLM). The generation unit can generate 3D models, 2D characters, voice-only avatars, etc., that reflect the user's characteristics. The dialogue unit has the generated avatar engage in dialogue on behalf of the user. The dialogue unit conducts conversations using methods such as text chat, voice chat, and video chat. The dialogue unit can use generative AI to conduct conversations that reflect the user's characteristics. As a result, the dialogue avatar system according to this embodiment can generate an avatar that reflects the user's characteristics and conduct conversations on behalf of the user in times of emergency. For example, the dialogue unit can have the generated avatar conduct a conversation on behalf of the user and take appropriate action. This allows the user to live their daily life with peace of mind.

[0065] The data collection unit collects user messages and call interactions. For example, the collection unit can collect data such as text messages, voice calls, and video calls. Specifically, in the case of text messages, it acquires the text information sent by the user in real time; in the case of voice calls, it records the call content and stores it as audio data; and in the case of video calls, it collects both video and audio to obtain detailed information such as the user's facial expressions and tone of voice. This data is centrally managed by the data collection unit and provided to the analysis unit. The data collection unit can use encryption technology to protect privacy when collecting data. For example, collected data is encrypted at the time of collection to prevent unauthorized access by third parties until it is transmitted to the analysis unit. Furthermore, the data collection unit has a mechanism to collect data only after obtaining the user's consent, and can stop collection if the user refuses data collection. This allows the data collection unit to efficiently collect necessary data while protecting user privacy. In addition, the data collection unit is equipped with high-speed data transfer technology to provide collected data to the analysis unit in real time, thereby improving the overall system response speed.

[0066] The analysis unit analyzes the collected data and extracts user characteristics. For example, the analysis unit uses natural language processing technology to analyze the characteristics of the user's speech and expression. Specifically, in the case of text messages, it performs morphological and grammatical analysis to identify words and phrases that the user frequently uses. In the case of voice calls, it uses speech recognition technology to convert voice data into text and analyzes that text data. It can also extract features such as voice tone, speaking speed, and emotional expression from the voice data. In the case of video calls, it uses video analysis technology to analyze the user's facial expressions and gestures, and infers the user's emotions and intentions based on this information. The analysis unit integrates these analysis results to comprehensively understand the user's characteristics. Furthermore, the analysis unit can learn user behavior patterns and preferences by utilizing past data and user history information. For example, based on past conversation history, it can predict how the user will react in specific situations and provide information to realize more natural conversations. In addition, the analysis unit can use anomaly detection algorithms to detect unusual user behavior or statements and issue warnings as needed. This allows the analysis unit to accurately understand user characteristics and improve the overall dialogue quality of the system.

[0067] The generation unit generates avatars based on features extracted by the analysis unit. The generation unit uses a generation AI to generate avatars that reflect the user's characteristics. Specifically, it uses a text generation AI (e.g., LLM) to generate the avatar's dialogue. The generation AI learns the user's speaking style and expressive characteristics and can generate natural dialogue based on that. For example, it can generate dialogue that incorporates words and phrases frequently used by the user and create avatar statements that resemble the user's speaking style. The generation unit can also generate 3D models, 2D characters, and voice-only avatars that reflect the user's characteristics. In the case of 3D models, it can generate avatars that reproduce the user's appearance and facial expressions and can operate them in real time. In the case of 2D characters, it generates illustrations and animations that reflect the user's characteristics and displays them during dialogue. In the case of voice-only avatars, it uses speech synthesis technology that reflects the characteristics of the user's voice to achieve natural voice dialogue. The generation unit provides these generation results to the dialogue unit, preparing it to engage in dialogue on behalf of the user. Furthermore, the generation unit can evaluate the quality of the generated avatars and make corrections or improvements as needed. For example, if the generated dialogue is unnatural or does not accurately reflect the user's characteristics, it is regenerated to improve quality. This allows the generation unit to produce high-quality avatars that accurately reflect the user's characteristics, thereby improving the overall dialogue quality of the system.

[0068] The dialogue unit uses a generated avatar to interact on behalf of the user. The dialogue unit engages in conversations using methods such as text chat, voice chat, and video chat. Specifically, in the case of text chat, it sends text messages generated by a generative AI on behalf of the user and engages in conversation with the other party. In the case of voice chat, it plays voice generated using speech synthesis technology and engages in conversation on behalf of the user. In the case of video chat, it displays a generated 3D model or 2D character and engages in conversation while it operates in real time. The dialogue unit can use generative AI to conduct conversations that reflect the user's characteristics. For example, it can generate dialogue content incorporating words and phrases frequently used by the user and create avatar statements that resemble the user's speaking style. Furthermore, the dialogue unit can incorporate new dialogue content generated by the generative AI in real time as the conversation progresses. This enables the dialogue unit to achieve natural and smooth conversations and respond appropriately on behalf of the user. In addition, the dialogue unit can record the content and progress of the conversation and provide this information to the analysis and generation units later. This allows for continuous improvement of the overall dialogue quality of the system through feedback. For example, if the content of the dialogue is inappropriate or does not accurately reflect the user's characteristics, the analysis and generation units can correct or improve it based on that information. This allows the dialogue unit to conduct high-quality dialogue on behalf of the user, improving the overall reliability and satisfaction of the system.

[0069] The collection unit can collect user messages or call content. For example, the collection unit can collect data such as text messages, voice calls, and video calls. The collection unit meticulously collects the content of messages and calls sent by the user and provides this data to the analysis unit. This allows for the understanding of user characteristics by collecting user messages and call content. Some or all of the above processing in the collection unit may be performed using AI, or without AI. For example, the collection unit can input user messages and call content into an AI, which can then analyze the data.

[0070] The analysis unit can analyze the collected data and extract characteristics of the user's speech or expression. For example, the analysis unit may use natural language processing techniques to analyze the characteristics of the user's speech and expression. The analysis unit may extract words, phrases, and tone of voice that the user frequently uses. By extracting characteristics of the user's speech and expression, it is possible to generate an avatar that reflects the user's characteristics. Some or all of the above processing in the analysis unit may be performed using AI, for example, or without AI. For example, the analysis unit may input the collected data into an AI, which can then analyze the data.

[0071] The generation unit can generate avatars based on extracted features. The generation unit uses a generation AI to generate avatars that reflect the user's characteristics. The generation AI uses, for example, a text generation AI (e.g., LLM) to generate the avatar's dialogue. The generation unit can generate 3D models, 2D characters, or voice-only avatars that reflect the user's characteristics. In this way, by generating avatars based on extracted features, it is possible to provide avatars that reflect the user's characteristics. Some or all of the above-described processes in the generation unit may be performed using, for example, a generation AI, or without using a generation AI. For example, the generation unit can input extracted features into a generation AI, and the generation AI can generate an avatar.

[0072] The dialogue unit allows a generated avatar to engage in dialogue on behalf of the user. The dialogue unit engages in dialogue using methods such as text chat, voice dialogue, and video dialogue. The dialogue unit can use a generation AI to engage in dialogue that reflects the user's characteristics. This allows the generated avatar to engage in dialogue on behalf of the user, enabling appropriate responses in times of emergency. Some or all of the above-described processes in the dialogue unit may be performed using AI, or not using AI. For example, the dialogue unit can input the generated avatar into an AI, which can then engage in dialogue.

[0073] The dialogue unit can use algorithms to respond in the event of an emergency. For example, the dialogue unit can take appropriate action in the event of an emergency using emergency response scenarios and the AI ​​technology it employs. This allows it to take appropriate action on behalf of the user by using algorithms designed for appropriate responses in emergencies. Some or all of the above-described processes in the dialogue unit may be performed using AI, or not using AI. For example, the dialogue unit can input emergency response scenarios into the AI, which can then take action.

[0074] The data collection unit can estimate the user's emotions and adjust the timing of message and call content collection based on the estimated emotions. The data collection unit estimates the user's emotions using technologies such as voice analysis, facial recognition, and text analysis. If the user is stressed, the data collection unit can temporarily stop collecting messages and call content and resume collection when the user is relaxed. If the user is relaxed, the data collection unit can actively collect messages and call content to obtain detailed data. Furthermore, if the user is in a hurry, the data collection unit can prioritize collecting important messages and call content in a short amount of time. By adjusting the collection timing based on the user's emotions, data can be collected at a more appropriate time. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the data collection unit may be performed using AI, or not using AI. For example, the data collection unit can input user emotion data into the generative AI, which can then adjust the collection timing.

[0075] The data collection unit can analyze the user's past messages and call history and select a collection method. For example, the data collection unit can prioritize collecting data from messaging and calling apps that the user has frequently used in the past. Furthermore, the data collection unit can concentrate data collection on specific time periods based on the user's past messages and call history. In addition, the data collection unit can analyze the user's past communication patterns and select the most effective collection method. This allows the optimal collection method to be selected by analyzing the user's past messages and call history. Some or all of the above processing in the data collection unit may be performed using AI, or not. For example, the data collection unit can input the user's past messages and call history into an AI, which can then select a collection method.

[0076] The data collection unit can filter messages and call content based on the user's current lifestyle and areas of interest. For example, if the user is at work, the data collection unit can prioritize collecting work-related messages and call content. Similarly, if the user is engaged in a hobby, the data collection unit can collect messages and call content related to that hobby. Furthermore, if the user is traveling, the data collection unit can collect travel-related messages and call content. This allows for the collection of more relevant data by filtering based on the user's lifestyle and areas of interest. Some or all of the processing described above in the data collection unit may be performed using AI, or not. For example, the data collection unit can input the user's lifestyle and areas of interest into an AI, which can then perform the filtering.

[0077] The data collection unit can estimate the user's emotions and determine the priority of messages and call content to collect based on the estimated emotions. The data collection unit estimates the user's emotions using technologies such as voice analysis, facial recognition, and text analysis. If the user is stressed, the data collection unit can prioritize collecting high-priority messages and call content. If the user is relaxed, the data collection unit can collect all messages and call content equally. Furthermore, if the user is in a hurry, the data collection unit can prioritize collecting urgent messages and call content. In this way, important data can be collected preferentially by determining the priority of messages and call content to collect based on the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the data collection unit may be performed using AI, or not using AI. For example, the data collection unit can input user emotion data into a generative AI, which can then determine the collection priority.

[0078] The data collection unit can prioritize the collection of highly relevant content based on the user's geographical location when collecting messages and call content. For example, if the user is in a specific region, the data collection unit can prioritize the collection of messages and call content related to that region. Furthermore, if the user is traveling, the data collection unit can prioritize the collection of messages and call content related to their travel destination. Additionally, if the user is at home, the data collection unit can prioritize the collection of messages and call content related to their home. This allows for the priority collection of highly relevant data by considering the user's geographical location. Some or all of the above processing in the data collection unit may be performed using AI, or without AI. For example, the data collection unit can input the user's geographical location information into AI, which can then prioritize the collection of highly relevant content.

[0079] The data collection unit can analyze the user's social media activity and collect relevant content when collecting messages and call content. For example, the data collection unit can prioritize collecting messages and call content with people the user frequently interacts with on social media. It can also collect messages and call content related to topics the user shows interest in on social media. Furthermore, the data collection unit can collect messages and call content related to groups and communities the user participates in on social media. In this way, relevant content can be collected by analyzing the user's social media activity. Some or all of the above processing in the data collection unit may be performed using AI, for example, or not. For example, the data collection unit can input the user's social media activity into AI, which can then collect relevant content.

[0080] The analysis unit can estimate the user's emotions and analyze the characteristics of their speech and expression based on the estimated emotions. The analysis unit estimates the user's emotions using technologies such as speech analysis, facial recognition, and text analysis. When the user is relaxed, the analysis unit can analyze a calm tone and slow speaking style. When the user is stressed, the analysis unit can analyze a fast pace and strong tone of speech. Furthermore, when the user is excited, the analysis unit can analyze emotionally rich expressions and emphasized words. This allows for the extraction of more accurate characteristics by analyzing the characteristics of speech and expression based on the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the analysis unit may be performed using AI, for example, or without AI. For example, the analysis unit can input the user's emotion data into a generative AI, which can then analyze the characteristics of their speech and expression.

[0081] The analysis unit can adjust the level of detail of the analysis based on the importance of the message or call content during the analysis. For example, the analysis unit can perform a detailed analysis of high-importance messages or call content to extract features down to the smallest detail. Conversely, the analysis unit can perform a simplified analysis of low-importance messages or call content to extract only the main features. Furthermore, the analysis unit can perform a rapid analysis of urgent messages or call content to immediately extract features. In this way, by adjusting the level of detail of the analysis based on the importance of the message or call content, the analysis can be performed efficiently. Some or all of the above processing in the analysis unit may be performed using AI, for example, or without AI. For example, the analysis unit can input the importance of the message or call content into the AI, and the AI ​​can adjust the level of detail of the analysis.

[0082] The analysis unit can apply different analysis algorithms depending on the category of the message or call content during analysis. For example, the analysis unit can apply an algorithm that analyzes formal expressions and technical terms to business-related messages or call content. It can also apply an algorithm that analyzes casual expressions and everyday language to private messages or call content. Furthermore, the analysis unit can apply an algorithm for rapid response to urgent messages or call content. By applying different analysis algorithms depending on the category of the message or call content, more accurate analysis can be performed. Some or all of the above-described processes in the analysis unit may be performed using AI, for example, or without AI. For example, the analysis unit can input the category of the message or call content into the AI, which can then apply different analysis algorithms.

[0083] The analysis unit can estimate the user's emotions and determine the priority of analysis based on the estimated emotions. The analysis unit estimates the user's emotions using technologies such as voice analysis, facial recognition, and text analysis. If the user is stressed, the analysis unit can prioritize the analysis of high-priority messages and call content. If the user is relaxed, the analysis unit can analyze all messages and call content equally. Furthermore, if the user is in a hurry, the analysis unit can prioritize the analysis of urgent messages and call content. In this way, by determining the priority of analysis based on the user's emotions, important data can be analyzed preferentially. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the analysis unit may be performed using AI, or not using AI. For example, the analysis unit can input user emotion data into a generative AI, which can then determine the priority of analysis.

[0084] The analysis unit can determine the priority of analysis based on the submission timing of messages and call content during the analysis process. For example, the analysis unit can prioritize the analysis of recently sent messages and call content. It can also prioritize the analysis of messages and call content sent during a specific time period. Furthermore, it can prioritize the analysis of messages and call content of high urgency. This allows for efficient analysis by determining the priority of analysis based on the submission timing of messages and call content. Some or all of the above-described processes in the analysis unit may be performed using AI, for example, or without AI. For example, the analysis unit can input the submission timing of messages and call content into the AI, which can then determine the priority of analysis.

[0085] The analysis unit can adjust the order of analysis based on the relevance of messages and call content during the analysis process. For example, the analysis unit can prioritize the analysis of highly relevant messages and call content. It can also postpone the analysis of less relevant messages and call content. Furthermore, the analysis unit can prioritize the analysis of messages and call content related to specific topics. This allows for efficient analysis by adjusting the order of analysis based on the relevance of messages and call content. Some or all of the above-described processes in the analysis unit may be performed using AI, for example, or without AI. For example, the analysis unit can input the relevance of messages and call content into the AI, which can then adjust the order of analysis.

[0086] The generation unit can estimate the user's emotions and adjust the avatar's expression based on the estimated emotions. The generation unit estimates the user's emotions using technologies such as voice analysis, facial recognition, and text analysis. If the user is relaxed, the generation unit can generate an avatar with calm facial expressions and movements. If the user is stressed, the generation unit can generate an avatar with calm facial expressions and movements. Furthermore, if the user is excited, the generation unit can generate an avatar with lively facial expressions and movements. By adjusting the avatar's expression based on the user's emotions, a more appropriate avatar can be generated. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or a generation AI. The generation AI is, but is not limited to, a text generation AI (e.g., LLM) or a multimodal generation AI. Some or all of the above processing in the generation unit may be performed using AI, for example, or without AI. For example, the generation unit can input the user's emotion data into the generation AI, which can then adjust the avatar's expression.

[0087] The generation unit can adjust the level of detail of the avatar based on the importance of the user's features during avatar generation. For example, if the user's features are clear, the generation unit can generate a detailed avatar. If the user's features are ambiguous, the generation unit can generate a simplified avatar. Furthermore, if the user's features are diverse, the generation unit can generate an avatar that emphasizes the main features. By adjusting the level of detail of the avatar based on the importance of the user's features, a more appropriate avatar can be generated. Some or all of the above processing in the generation unit may be performed using AI, for example, or without AI. For example, the generation unit can input the importance of the user's features into the AI, and the AI ​​can adjust the level of detail of the avatar.

[0088] The generation unit can apply different generation algorithms depending on the user's category when generating avatars. For example, the generation unit can apply a generation algorithm with formal expressions and actions to avatars for business use. It can also apply a generation algorithm with casual expressions and actions to avatars for private use. Furthermore, the generation unit can apply a generation algorithm for rapid response to avatars in emergencies. By applying different generation algorithms according to the user's category, a more appropriate avatar can be generated. Some or all of the above-described processes in the generation unit may be performed using AI, for example, or without AI. For example, the generation unit can input the user's category into the AI, which can then apply different generation algorithms.

[0089] The generation unit can estimate the user's emotions and adjust the avatar's appearance and actions based on the estimated emotions. The generation unit estimates the user's emotions using technologies such as voice analysis, facial recognition, and text analysis. If the user is relaxed, the generation unit can generate an avatar with calm facial expressions and actions. If the user is stressed, the generation unit can generate an avatar with calm facial expressions and actions. Furthermore, if the user is excited, the generation unit can generate an avatar with lively facial expressions and actions. By adjusting the avatar's appearance and actions based on the user's emotions, a more appropriate avatar can be generated. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or a generation AI. The generation AI is, but is not limited to, a text generation AI (e.g., LLM) or a multimodal generation AI. Some or all of the above-described processes in the generation unit may be performed using AI, or not using AI. For example, the generation unit can input the user's emotion data into the generation AI, which can then adjust the avatar's appearance and actions.

[0090] The generation unit can determine the generation priority based on the user's past avatar usage history when generating avatars. For example, the generation unit can prioritize the reflection of features of avatars that the user has frequently used in the past. Furthermore, the generation unit can predict and generate avatars that the user will use at specific times based on their past avatar usage history. In addition, the generation unit can analyze the user's past avatar usage patterns and generate the most effective avatars. This allows for the generation of more appropriate avatars by determining the generation priority based on the user's past avatar usage history. Some or all of the above-described processes in the generation unit may be performed using AI, or not. For example, the generation unit can input the user's past avatar usage history into an AI, which can then determine the generation priority.

[0091] The generation unit can improve the accuracy of avatar generation by referring to the user's relevant data. For example, the generation unit can refer to the user's social media profile and posts to adjust the avatar's appearance and personality. It can also refer to the user's past messages and call content to adjust the avatar's speaking style and expressions. Furthermore, the generation unit can customize the avatar's characteristics based on the user's hobbies and interests. This allows for improved generation accuracy by referring to the user's relevant data. Some or all of the above-described processes in the generation unit may be performed using AI, for example, or without AI. For example, the generation unit can input the user's relevant data into the AI, which can then improve the generation accuracy.

[0092] The dialogue unit can estimate the user's emotions and adjust the way the dialogue is expressed based on those estimated emotions. The dialogue unit estimates the user's emotions using technologies such as speech analysis, facial recognition, and text analysis. If the user is relaxed, the dialogue unit can engage in dialogue in a calm tone. If the user is stressed, the dialogue unit can engage in dialogue in a calm tone. Furthermore, if the user is excited, the dialogue unit can engage in dialogue in an active tone. This allows for more appropriate dialogue by adjusting the way the dialogue is expressed based on the user's emotions. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or generative AI. Generative AI may be, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above-described processes in the dialogue unit may be performed using AI, or not. For example, the dialogue unit can input user emotion data into a generative AI, which can then adjust the way the dialogue is expressed.

[0093] The dialogue unit can select a dialogue method by referring to the user's past dialogue history during a conversation. For example, the dialogue unit can incorporate expressions and phrases that the user has previously used frequently into the conversation. Furthermore, the dialogue unit can select a dialogue method for a specific topic based on the user's past dialogue history. In addition, the dialogue unit can analyze the user's past dialogue patterns and select the most effective dialogue method. This allows the optimal dialogue method to be selected by referring to the user's past dialogue history. Some or all of the above processing in the dialogue unit may be performed using AI, or not. For example, the dialogue unit can input the user's past dialogue history into an AI, which can then select a dialogue method.

[0094] The dialogue unit can customize the means of dialogue based on the user's current situation during a conversation. For example, if the user is at work, the dialogue unit can engage in a businesslike conversation. If the user is relaxed, the dialogue unit can engage in a casual conversation. Furthermore, if the user is in an emergency, the dialogue unit can engage in a quick and concise conversation. In this way, by customizing the means of dialogue based on the user's current situation, a more appropriate conversation can be conducted. Some or all of the above processing in the dialogue unit may be performed using AI, for example, or without AI. For example, the dialogue unit can input the user's current situation into the AI, and the AI ​​can customize the means of dialogue.

[0095] The dialogue unit can estimate the user's emotions and determine the priority of conversations based on the estimated emotions. The dialogue unit estimates the user's emotions using technologies such as voice analysis, facial recognition, and text analysis. If the user is stressed, the dialogue unit can prioritize high-importance conversations. If the user is relaxed, the dialogue unit can distribute all conversations equally. Furthermore, if the user is in a hurry, the dialogue unit can prioritize urgent conversations. In this way, important conversations can be prioritized by determining the priority of conversations based on the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the dialogue unit may be performed using AI, or not using AI. For example, the dialogue unit can input user emotion data into a generative AI, which can then determine the priority of conversations.

[0096] The dialogue unit can select the optimal dialogue method during a conversation by considering the user's geographical location information. For example, if the user is in a specific region, the dialogue unit can provide information related to that region. If the user is traveling, the dialogue unit can provide information related to the travel destination. Furthermore, if the user is at home, the dialogue unit can provide information related to home. In this way, the optimal dialogue method can be selected by considering the user's geographical location information. Some or all of the above processing in the dialogue unit may be performed using AI, for example, or without AI. For example, the dialogue unit can input the user's geographical location information into the AI, and the AI ​​can select the optimal dialogue method.

[0097] The dialogue unit can analyze the user's social media activity during a conversation and suggest a suitable dialogue method. For example, the dialogue unit can prioritize conversations with people the user frequently interacts with on social media. It can also conduct conversations related to topics the user has shown interest in on social media. Furthermore, it can conduct conversations related to groups and communities the user participates in on social media. In this way, by analyzing the user's social media activity, it can suggest the most suitable dialogue method. Some or all of the above processing in the dialogue unit may be performed using AI, for example, or without AI. For example, the dialogue unit can input the user's social media activity into AI, which can then suggest a dialogue method.

[0098] The system according to the embodiment is not limited to the example described above, and various modifications are possible, for example, as follows.

[0099] The interactive avatar system can also collect user health data and reflect it in the avatar's dialogue. For example, the data collection unit can collect heart rate and sleep data from the user's wearable device. The analysis unit can analyze this data to understand the user's health status. The generation unit can have the avatar provide health advice based on the user's health status. The dialogue unit can engage in dialogue to promote relaxation or encourage exercise, depending on the user's health status. This allows for more personalized support by providing dialogue that takes the user's health status into consideration.

[0100] The data collection unit can collect user environmental data and reflect it in the avatar's dialogue. For example, the data collection unit can collect environmental data such as temperature, humidity, and noise level around the user. The analysis unit can analyze this data to understand the user's environmental conditions. The generation unit can provide advice to help the avatar maintain a comfortable environment based on the user's environmental conditions. The dialogue unit can engage in appropriate dialogue according to the user's environmental conditions. In this way, by engaging in dialogue that takes the user's environmental conditions into consideration, it is possible to support a more comfortable life.

[0101] The analysis unit can analyze the user's hobbies and interests and reflect them in the avatar's dialogue. For example, the analysis unit can analyze data on keywords the user has searched for in the past and websites they have visited. Based on this data, the generation unit can have the avatar provide topics related to the user's hobbies and interests. The dialogue unit can suggest relevant information and activities according to the user's hobbies and interests. This allows for more interesting dialogue by taking the user's hobbies and interests into consideration.

[0102] The generation unit can customize the avatar's dialogue by referring to the user's past dialogue history. For example, the generation unit can incorporate expressions and phrases that the user has previously used frequently into the avatar's dialogue. The dialogue unit can select a dialogue method for a specific topic from the user's past dialogue history. Furthermore, the generation unit can analyze the user's past dialogue patterns and select the most effective dialogue method. This allows for more personalized dialogue by referring to the user's past dialogue history.

[0103] The dialogue unit can customize the avatar's conversation content by taking into account the user's geographical location. For example, if the user is in a specific region, the dialogue unit can provide conversations that offer information related to that region. If the user is traveling, the dialogue unit can provide conversations that offer information related to their travel destination. Furthermore, if the user is at home, the dialogue unit can provide conversations that offer information related to their home. This allows for more relevant conversations by considering the user's geographical location.

[0104] The data collection unit can estimate the user's emotions and adjust the type of data collected based on those emotions. For example, if the user is stressed, the unit can collect music or video data to help them relax. If the user is relaxed, the unit can collect data that is helpful for learning or working. Furthermore, if the user is excited, the unit can collect entertainment-related data. By adjusting the type of data collected based on the user's emotions, the system can provide more relevant data.

[0105] The analysis unit can estimate the user's emotions and adjust the level of detail of the analysis based on the estimated emotions. For example, if the user is stressed, the analysis unit can perform a detailed analysis to identify the cause of the stress. If the user is relaxed, the analysis unit can perform a simplified analysis and extract only the main features. Furthermore, if the user is excited, the analysis unit can perform a rapid analysis and immediately extract features. By adjusting the level of detail of the analysis based on the user's emotions, a more appropriate analysis can be performed.

[0106] The generation unit can estimate the user's emotions and adjust the avatar's appearance and actions based on those emotions. For example, if the user is relaxed, the generation unit can generate an avatar with a calm expression and actions. If the user is stressed, the generation unit can generate an avatar with a calm expression and actions. Furthermore, if the user is excited, the generation unit can generate an avatar with an energetic expression and actions. By adjusting the avatar's appearance and actions based on the user's emotions, a more appropriate avatar can be generated.

[0107] The dialogue unit can estimate the user's emotions and adjust the way it expresses itself based on those emotions. For example, if the user is relaxed, the dialogue unit can use a calm tone. If the user is stressed, it can use a calm tone. Furthermore, if the user is excited, it can use a lively tone. By adjusting the way it expresses itself based on the user's emotions, it can provide more appropriate dialogue.

[0108] The dialogue unit can estimate the user's emotions and prioritize conversations based on those emotions. For example, if the user is stressed, the dialogue unit can prioritize high-priority conversations. If the user is relaxed, the dialogue unit can distribute all conversations equally. Furthermore, if the user is in a hurry, the dialogue unit can prioritize urgent conversations. In this way, by prioritizing conversations based on the user's emotions, important conversations can be given priority.

[0109] The following briefly describes the processing flow for example form 2.

[0110] Step 1: The collection unit collects user messages or call interactions. The collection unit can collect data such as text messages, voice calls, and video calls. The collection unit collects detailed information about the messages and calls sent by the user and provides this data to the analysis unit. Step 2: The analysis unit analyzes the collected data and extracts user characteristics. For example, the analysis unit uses natural language processing technology to analyze the characteristics of the user's speech and expressions. The analysis unit extracts words, phrases, and speaking tone that the user frequently uses. Step 3: The generation unit generates an avatar based on the features extracted by the analysis unit. The generation unit uses a generation AI to generate an avatar that reflects the user's characteristics. The generation AI generates the avatar's dialogue using, for example, a text generation AI (e.g., LLM). The generation unit can generate 3D models, 2D characters, or voice-only avatars that reflect the user's characteristics. Step 4: The dialogue unit has the generated avatar interact on behalf of the user. The dialogue unit interacts using methods such as text chat, voice dialogue, or video dialogue. The dialogue unit can use a generation AI to perform dialogue that reflects the user's characteristics. As a result, the dialogue avatar system according to the embodiment can generate an avatar that reflects the user's characteristics and perform dialogue on behalf of the user in times of emergency.

[0111] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0112] Data generation model 58 is a form of so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> Examples of generative AI include text generation AI, image generation AI, and multimodal generation AI. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats from audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVMs), k-means clustering, convolutional neural networks (CNNs), recurrent neural networks (RNNs), generative adversarial networks (GANs), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI ​​may be an AI agent. Furthermore, when the processing of each of the above parts is performed by the AI, the processing may be performed by the AI ​​in part or in whole, but is not limited to this example.Furthermore, processing performed by AI, including generative AI, may be replaced with rule-based processing, and rule-based processing may be replaced with processing performed by AI, including generative AI.

[0113] Furthermore, the processing performed by the data processing system 10 described above is carried out by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart device 14, but it may also be carried out by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart device 14. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the smart device 14 or an external device, and the smart device 14 acquires or collects information necessary for processing from the data processing device 12 or an external device.

[0114] Each of the multiple elements described above, including the collection unit, analysis unit, generation unit, and dialogue unit, is implemented in at least one of the smart device 14 and the data processing unit 12. For example, the collection unit collects user messages and call interactions using the camera 42 and microphone 38B of the smart device 14. The analysis unit is implemented in the identification processing unit 290 of the data processing unit 12, for example, and analyzes the collected data to extract user characteristics. The generation unit is implemented in the identification processing unit 290 of the data processing unit 12, for example, and generates an avatar that reflects the user's characteristics using generation AI. The dialogue unit is implemented in the control unit 46A of the smart device 14, for example, and the generated avatar engages in dialogue on behalf of the user. The correspondence between each unit and the device or control unit is not limited to the example described above and can be modified in various ways.

[0115] [Second Embodiment] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0116] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0117] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.

[0118] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0119] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0120] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).

[0121] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0122] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing by the processor 28. The storage 32 stores the specific processing program 56.

[0123] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0124] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.

[0125] In the smart glasses 214, specific processing is performed by the processor 46. The storage 50 stores a specific processing program 60. The processor 46 reads the specific processing program 60 from the storage 50 and executes the read specific processing program 60 on the RAM 48. The specific processing is realized by the processor 46 acting as a control unit 46A according to the specific processing program 60 executed on the RAM 48. The smart glasses 214 also have a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.

[0126] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).

[0127] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0128] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI ​​may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI ​​in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.

[0129] The data processing system 210 according to the second embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 210 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart glasses 214, but it may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart glasses 214. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the smart glasses 214 or an external device, and the smart glasses 214 acquires or collects information necessary for processing from the data processing device 12 or an external device.

[0130] Each of the multiple elements described above, including the collection unit, analysis unit, generation unit, and dialogue unit, is implemented in at least one of the smart glasses 214 and the data processing unit 12. For example, the collection unit collects user messages and call interactions using the camera 42 and microphone 238 of the smart glasses 214. The analysis unit is implemented in the identification processing unit 290 of the data processing unit 12, for example, and analyzes the collected data to extract user characteristics. The generation unit is implemented in the identification processing unit 290 of the data processing unit 12, for example, and generates an avatar that reflects the user's characteristics using generation AI. The dialogue unit is implemented in the control unit 46A of the smart glasses 214, for example, and the generated avatar engages in dialogue on behalf of the user. The correspondence between each unit and the device or control unit is not limited to the example described above and can be modified in various ways.

[0131] [Third Embodiment] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0132] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0133] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.

[0134] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0135] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0136] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).

[0137] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0138] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0139] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0140] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.

[0141] In the headset terminal 314, specific processing is performed by the processor 46. The storage 50 stores a specific program 60. The processor 46 reads the specific program 60 from the storage 50 and executes the read specific program 60 on the RAM 48. The specific processing is realized by the processor 46 acting as a control unit 46A according to the specific program 60 executed on the RAM 48. The headset terminal 314 also has a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.

[0142] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).

[0143] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0144] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI ​​may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI ​​in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.

[0145] The data processing system 310 according to the third embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 310 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the headset terminal 314, but may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the headset terminal 314. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the headset terminal 314 or an external device, and the headset terminal 314 acquires or collects information necessary for processing from the data processing device 12 or an external device.

[0146] Each of the multiple elements described above, including the collection unit, analysis unit, generation unit, and dialogue unit, is implemented in at least one of the headset terminal 314 and the data processing unit 12. For example, the collection unit collects user messages and call exchanges using the camera 42 and microphone 238 of the headset terminal 314. The analysis unit is implemented in the identification processing unit 290 of the data processing unit 12, for example, and analyzes the collected data to extract user characteristics. The generation unit is implemented in the identification processing unit 290 of the data processing unit 12, for example, and generates an avatar that reflects the user's characteristics using a generation AI. The dialogue unit is implemented in the control unit 46A of the headset terminal 314, for example, and the generated avatar engages in dialogue on behalf of the user. The correspondence between each unit and the device or control unit is not limited to the example described above, and various modifications are possible.

[0147] [Fourth Embodiment] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0148] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0149] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.

[0150] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0151] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0152] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS image sensor or CCD image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).

[0153] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0154] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. The robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0155] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0156] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0157] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.

[0158] In robot 414, specific processing is performed by processor 46. A specific program 60 is stored in storage 50. Processor 46 reads the specific program 60 from storage 50 and executes it on RAM 48. The specific processing is achieved by processor 46 acting as a control unit 46A according to the specific program 60 executed on RAM 48. Robot 414 also has data generation model 58 and emotion identification model 59, similar to those of the robot, and can perform processing similar to that of the specific processing unit 290 using these models.

[0159] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).

[0160] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0161] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI ​​may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI ​​in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.

[0162] The data processing system 410 according to the fourth embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 410 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the robot 414, but it may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the robot 414. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the robot 414 or an external device, and the robot 414 acquires or collects information necessary for processing from the data processing device 12 or an external device.

[0163] Each of the multiple elements described above, including the collection unit, analysis unit, generation unit, and dialogue unit, is implemented in at least one of the following: the robot 414 and the data processing unit 12. For example, the collection unit collects user messages and call exchanges using the camera 42 and microphone 238 of the robot 414. The analysis unit is implemented, for example, by the identification processing unit 290 of the data processing unit 12, which analyzes the collected data to extract user characteristics. The generation unit is implemented, for example, by the identification processing unit 290 of the data processing unit 12, which generates an avatar that reflects the user's characteristics using a generation AI. The dialogue unit is implemented, for example, by the control unit 46A of the robot 414, and the generated avatar engages in dialogue on behalf of the user. The correspondence between each unit and the device or control unit is not limited to the example described above, and various modifications are possible.

[0164] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0165] Figure 9 shows the emotion map 400, in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0166] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0167] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0168] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, and motorcycles, emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated based, for example, on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0169] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0170] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0171] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing method for the specific process may be used, which includes computer 22 and multiple other computers.

[0172] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0173] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0174] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0175] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0176] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0177] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0178] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0179] Furthermore, although the above-described examples were divided into four embodiments, some or all of these embodiments may be combined. Also, the smart device 14, smart glasses 214, headset terminal 314, and robot 414 are just examples, and they may be combined, or other devices may be used. Also, although the above-described examples were divided into two embodiments, Embodiment 1 and Embodiment 2, these may be combined.

[0180] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and other things that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0181] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.

[0182] (Note 1) A collection unit that collects user messages or call exchanges, An analysis unit analyzes the information collected by the aforementioned collection unit and extracts user characteristics, A generation unit generates an avatar based on the features extracted by the analysis unit, The system comprises a dialogue unit in which the avatar generated by the generation unit engages in dialogue. system. (Note 2) The aforementioned collection unit is Collect user messages or call content. The system described in Appendix 1, characterized by the features described herein. (Note 3) The aforementioned analysis unit, The collected data is analyzed to extract characteristics of the user's speech or expression. The system described in Appendix 1, characterized by the features described herein. (Note 4) The generating unit is Generate an avatar based on extracted features. The system described in Appendix 1, characterized by the features described herein. (Note 5) The aforementioned dialogue unit, The generated avatar interacts on behalf of the user. The system described in Appendix 1, characterized by the features described herein. (Note 6) The aforementioned dialogue unit, Use an algorithm to respond in the event of an emergency. The system described in Appendix 1, characterized by the features described herein. (Note 7) The aforementioned collection unit is The system estimates the user's emotions and adjusts the timing of message and call content collection based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 8) The aforementioned collection unit is Analyze the user's past messages and call history to select the best collection method. The system described in Appendix 1, characterized by the features described herein. (Note 9) The aforementioned collection unit is When collecting messages and call content, filtering is performed based on the user's current lifestyle and areas of interest. The system described in Appendix 1, characterized by the features described herein. (Note 10) The aforementioned collection unit is It estimates the user's emotions and determines the priority of messages and call content to collect based on the estimated user emotions. The system described in Appendix 1, characterized by the features described herein. (Note 11) The aforementioned collection unit is When collecting messages and call content, the system prioritizes collecting relevant content based on the user's geographical location. The system described in Appendix 1, characterized by the features described herein. (Note 12) The aforementioned collection unit is When collecting messages and call content, the system analyzes the user's social media activity and collects relevant information. The system described in Appendix 1, characterized by the features described herein. (Note 13) The aforementioned analysis unit, It estimates the user's emotions and analyzes the characteristics of their speech and expressions based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 14) The aforementioned analysis unit, During analysis, the level of detail is adjusted based on the importance of the messages and call content. The system described in Appendix 1, characterized by the features described herein. (Note 15) The aforementioned analysis unit, During analysis, different analysis algorithms are applied depending on the category of the message or call content. The system described in Appendix 1, characterized by the features described herein. (Note 16) The aforementioned analysis unit, The system estimates the user's emotions and determines the priority of analysis based on the estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 17) The aforementioned analysis unit, During analysis, the priority of the analysis is determined based on when the messages and call content were submitted. The system described in Appendix 1, characterized by the features described herein. (Note 18) The aforementioned analysis unit, During analysis, the order of analysis is adjusted based on the relevance of messages and call content. The system described in Appendix 1, characterized by the features described herein. (Note 19) The generating unit is It estimates the user's emotions and adjusts the avatar's representation based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 20) The generating unit is When generating an avatar, the level of detail is adjusted based on the importance of the user's features. The system described in Appendix 1, characterized by the features described herein. (Note 21) The generating unit is When generating avatars, different generation algorithms are applied depending on the user's category. The system described in Appendix 1, characterized by the features described herein. (Note 22) The generating unit is It estimates the user's emotions and adjusts the avatar's appearance and behavior based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 23) The generating unit is When generating avatars, the system prioritizes their creation based on the user's past avatar usage history. The system described in Appendix 1, characterized by the features described herein. (Note 24) The generating unit is When generating avatars, the system improves the accuracy of the generation process by referencing the user's relevant data. The system described in Appendix 1, characterized by the features described herein. (Note 25) The aforementioned dialogue unit, It estimates the user's emotions and adjusts the way the dialogue is expressed based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 26) The aforementioned dialogue unit, During a conversation, the system selects the appropriate conversation method by referring to the user's past conversation history. The system described in Appendix 1, characterized by the features described herein. (Note 27) The aforementioned dialogue unit, During a conversation, customize the conversational methods based on the user's current situation. The system described in Appendix 1, characterized by the features described herein. (Note 28) The aforementioned dialogue unit, It estimates the user's emotions and determines the priority of the conversation based on the estimated user emotions. The system described in Appendix 1, characterized by the features described herein. (Note 29) The aforementioned dialogue unit, During the interaction, the system selects the optimal interaction method, taking into account the user's geographical location. The system described in Appendix 1, characterized by the features described herein. (Note 30) The aforementioned dialogue unit, During the conversation, we analyze the user's social media activity and suggest ways to communicate. The system described in Appendix 1, characterized by the features described herein. [Explanation of symbols]

[0183] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots

Claims

1. A collection unit that collects user messages or call exchanges, An analysis unit analyzes the information collected by the aforementioned collection unit and extracts user characteristics, A generation unit generates an avatar based on the features extracted by the analysis unit, The system comprises a dialogue unit in which the avatar generated by the generation unit engages in dialogue, The analysis unit estimates the user's emotions and performs a detailed analysis if it estimates the user is feeling stressed, and performs a simplified analysis if it estimates the user is relaxed. system.

2. The aforementioned collection unit is Collect user messages or call content. The system according to feature 1.

3. The aforementioned analysis unit, The collected data is analyzed to extract characteristics of the user's speech or expression. The system according to feature 1.

4. The aforementioned dialogue unit, Use an algorithm to respond in the event of an emergency. The system according to feature 1.

5. The aforementioned collection unit is The system estimates the user's emotions and adjusts the timing of message and call content collection based on those estimated emotions. The system according to feature 1.

6. The aforementioned collection unit is Analyze the user's past messages and call history to select the best collection method. The system according to feature 1.

Citation Information

Patent Citations

  • Behavior information recording apparatus

    JP2012168862A

  • Sound information processing device and system

    JP2016059765A

  • Information collection program, information collection system, and information collection method

    JP2016212597A

  • Program and information processing device

    JP2020166360A

  • Persona chatbot control method and system

    JP2022180282A