system

A system with a learning, proxy, and cloning unit addresses the challenge of personalization by learning user patterns and creating clones for personalized communication.

JP2026072667APending Publication Date: 2026-05-01SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-18
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing systems struggle to meet the individual needs and personalities of users, failing to provide a personalized experience.

Method used

A system comprising a learning unit, proxy unit, and cloning unit that learns a user's language and thought patterns, functions as a proxy, and creates a clone to provide personalized communication experiences.

Benefits of technology

The system effectively responds to individual user needs and personalities, providing a personalized experience through natural and tailored interactions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026072667000001_ABST
    Figure 2026072667000001_ABST
Patent Text Reader

Abstract

The system according to this embodiment aims to provide a personalized experience that responds to the individual needs and personalities of users. [Solution] The system according to the embodiment comprises a learning unit, an agency unit, and a clone creation unit. The learning unit learns the user's speech patterns and thought patterns. The agency unit functions as the user's agency based on the information learned by the learning unit. The clone creation unit creates a clone of the user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0004] ,

[0006] , , , ,

[0005] , , , , ,

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a method for controlling a persona chatbot performed by at least one processor, the method including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In the prior art, there is a problem that it is difficult to meet the individual needs and personalities of users and to provide a personalized experience.

[0005] The system according to the embodiment aims to meet the individual needs and personalities of users and provide a personalized experience.

Means for Solving the Problems

[0006] The system according to the embodiment includes a learning unit, a proxy unit, and a clone creation unit . The learning unit learns the user's language usage and thinking pattern. The proxy unit functions as a proxy for the user based on the information learned by the learning unit. The clone creation unit creates a clone of the user. [Effects of the Invention]

[0007] The system according to this embodiment can respond to the individual needs and personalities of users and provide a personalized experience. [Brief explanation of the drawing]

[0008] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Modes for carrying out the invention]

[0009] Hereinafter, an example of an embodiment of the system relating to the technology of this disclosure will be described with reference to the attached drawings.

[0010] First, let's explain the terminology used in the following explanation.

[0011] In the following embodiments, the signed processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Furthermore, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose computing on Graphics Processing Units), APU (Accelerated Processing Unit), or TPU (Tensor Processing Unit).

[0012] In the following embodiments, signed RAM (Random Access Memory) is a memory that temporarily stores information and is used as work memory by the processor.

[0013] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.

[0014] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0015] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B". That is, "A and / or B" means that it may be only A, only B, or a combination of A and B. Also, in this specification, when expressing three or more matters connected by "and / or", the same concept as "A and / or B" is applied.

[0016] [First Embodiment] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0017] As shown in FIG. 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0018] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. Also, the database 24 and the communication I / F 26 are connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0019] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. Also, the reception device 38, the output device 40, and the camera 42 are connected to the bus 52.

[0020] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, and accepts user input. The touch panel 38A accepts user input via touch by detecting contact with an object (e.g., a pen or finger). The microphone 38B accepts user input via voice by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 (see Figure 2) acquires the data indicating the user input.

[0021] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user by outputting the data in a form perceptible to the user (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0022] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0023] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0024] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0025] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.

[0026] In the smart device 14, specific processing is performed by the processor 46. The storage 50 stores a specific processing program 60. The specific processing program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the specific processing program 60 from the storage 50 and executes the read specific processing program 60 on the RAM 48. The specific processing is realized by the processor 46 operating as a control unit 46A according to the specific processing program 60 executed on the RAM 48. The smart device 14 also has a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.

[0027] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device (e.g., a generation server) may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device having the data generation model 58. The data processing device 12 may also be a server device or a terminal device owned by a user (e.g., a mobile phone, robot, home appliance, etc.). Next, an example of processing by the data processing system 10 according to the first embodiment will be described.

[0028] (Example of form 1) The Bot system according to an embodiment of the present invention is a system that learns a user's speech patterns and thought patterns and functions as a proxy for the user over time. This Bot system learns the user's unique speech patterns and thought patterns through interaction with the user and evolves to function as a proxy for the user over time. Technically, it utilizes multimodal AI by making full use of AI APIs, fine-tuning, prompt engineering, etc. This allows the user to enjoy a more personalized and richer communication experience. For example, the Bot system learns the user's unique speech patterns and thought patterns through interaction with the user. Next, based on the learned information, it evolves to function as a proxy for the user. For example, it can not only provide fresh topics for chats with friends and family and liven up conversations, but also automatically answer specific questions. In this way, user satisfaction is improved and deeper relationships can be formed. Furthermore, this Bot system can also create a clone of the user. It can learn the user's way of speaking and thinking and communicate using language that is characteristic of that person. It is also possible for the clone to reply on behalf of the user using the automatic response function. In addition, by using AI APIs, the AI ​​can engage in continuous conversations as if it had memory. This allows users to enjoy more natural conversations. Furthermore, by providing sample statements, prompt engineering and fine-tuning can be performed, allowing the AI ​​to mimic the user's way of speaking and thinking. This bot system can create clones not only of the user but also of other people. For example, you can have conversations and play with a clone of your friend. Thus, this invention realizes a bot system that responds to the individual needs and personalities of users, providing a more personalized communication experience. By learning the user's language and thought patterns, the bot system can function as a proxy and create clones to provide a personalized communication experience.

[0029] The Bot system according to this embodiment comprises a learning unit, an agency unit, and a cloning unit. The learning unit learns the user's speech patterns and thought patterns. For example, the learning unit learns the user's unique speech patterns and thought patterns through dialogue with the user. The learning unit can analyze the user's speech patterns and thought patterns using, for example, a machine learning algorithm. The learning unit can also learn the user's speech patterns and thought patterns based on the user's dialogue history. For example, the learning unit identifies phrases and speech patterns that the user frequently uses and learns based on them. The agency unit functions as a proxy for the user based on the information learned by the learning unit. For example, the agency unit can generate responses by mimicking the user's speech patterns and thought patterns. For example, the agency unit can participate in chats with friends and family on behalf of the user and provide fresh topics. The agency unit can also automatically answer specific questions. For example, the agency unit generates appropriate responses based on the user's past answers. The cloning unit creates a clone of the user. The cloning unit can, for example, learn the user's way of speaking and thinking, and communicate using language that reflects that person's personality. Furthermore, the cloning unit can use an automated response function to have a clone reply on its behalf. For instance, the cloning unit can participate in chats with friends and family on the user's behalf, providing fresh topics of conversation. Thus, the bot system according to this embodiment can learn the user's language and thought patterns, function as a proxy, and create clones to provide a personalized communication experience.

[0030] The learning unit learns the user's language use and thought patterns. Specifically, the learning unit learns the user's unique language use and thought patterns through dialogue with the user. For example, it analyzes in detail the phrases and word choices the user uses daily, sentence structure, and even how they express emotions. The learning unit uses machine learning algorithms to analyze this data and model the user's language use and thought patterns. Specifically, it utilizes natural language processing (NLP) technology to tokenize the user's utterances and generate contextual vectors to understand the context. Furthermore, it uses recurrent neural networks (RNNs) and transformer models to learn the temporal flow and relationships of the user's utterances. Based on the user's dialogue history, the learning unit can continuously learn and update the user's language use and thought patterns. For example, it identifies phrases the user frequently uses and response patterns in specific situations and uses them as a basis for learning. As a result, the learning unit can always reflect the user's latest language use and thought patterns, generating more natural and personalized responses.

[0031] The proxy unit functions as a proxy for the user based on information learned by the learning unit. Specifically, the proxy unit generates responses by mimicking the user's speech patterns and thought patterns. For example, it can participate in chats with friends and family on behalf of the user and provide fresh topics of conversation. The proxy unit refers to past conversation history to generate appropriate responses based on the user's past answers. This allows for responses tailored to the user's style and preferences. The proxy unit uses generative AI to mimic the user's speech patterns and thought patterns. Specifically, the generative AI learns the user's speech patterns and generates new statements based on them. For example, it learns the user's frequently used phrases and word choices and generates natural responses based on them. The proxy unit can also automatically answer specific questions. For example, it generates appropriate responses based on the user's past answers. This allows the proxy unit to provide quick and appropriate responses on behalf of the user. Furthermore, the proxy unit can understand the user's intentions and emotions and generate responses accordingly. For example, if the user expresses gratitude, the proxy unit will generate appropriate words of thanks. This allows the proxy unit to provide natural and personalized communication on behalf of the user.

[0032] The cloning unit creates a clone of the user. Specifically, the cloning unit learns the user's way of speaking and thinking, and can communicate using language that is characteristic of that person. The cloning unit generates a clone of the user based on the user's dialogue history and data collected by the learning unit. For example, the cloning unit identifies phrases and word choices that the user frequently uses and creates a clone based on that. The cloning unit uses generative AI to mimic the user's language and thought patterns. Specifically, the generative AI learns the user's speech patterns and generates new statements based on them. For example, it learns the phrases and word choices that the user frequently uses and generates natural responses based on that. For example, the cloning unit can participate in chats with friends and family on behalf of the user and provide fresh topics. The cloning unit can also use an auto-responder function so that the clone replies on behalf of the user. For example, the cloning unit can participate in chats with friends and family on behalf of the user and provide fresh topics. In this way, the cloning unit can mimic the user's language and thought patterns and provide natural and personalized communication. Furthermore, the cloning unit can continuously learn from and update the user's clone. For example, it can update the clone's vocabulary and thought patterns based on the user's most recent conversation history. This allows the cloning unit to always provide a clone based on the latest information, enabling it to provide natural and personalized communication on behalf of the user.

[0033] The response unit can provide an automated response function. For example, the response unit can have a clone automatically respond on behalf of the user. For example, the response unit can reduce the user's burden by having the clone automatically respond when the user is busy or unable to attend to the user. For example, the response unit can participate in chats with friends and family on behalf of the user and provide fresh topics. The response unit can also automatically answer specific questions. For example, the response unit can generate an appropriate response based on the user's past responses. This allows the response unit to reduce the user's burden by having the clone automatically respond on behalf of the user. Some or all of the above processing in the response unit may be performed using AI, for example, or not using AI. For example, the response unit can generate responses using an AI model for the clone to automatically respond on behalf of the user.

[0034] The conversation continuation unit can engage in continuous conversation. For example, the conversation continuation unit can engage in continuous conversation as if the AI ​​had memory. For example, the conversation continuation unit can maintain the context of the conversation and achieve natural dialogue based on the user's dialogue history. For example, the conversation continuation unit can continue the conversation by reusing topics that the user has previously discussed. The conversation continuation unit can also continue the conversation by referring to the language that the user has previously preferred to use. For example, the conversation continuation unit can continue the conversation while avoiding topics that the user has previously avoided. In this way, the conversation continuation unit can achieve natural dialogue by engaging in continuous conversation as if the AI ​​had memory. Some or all of the above processing in the conversation continuation unit may be performed using AI, for example, or without AI. For example, the conversation continuation unit can continue the conversation using an AI model that maintains the context of the conversation and achieves natural dialogue based on the user's dialogue history.

[0035] The adjustment unit can perform prompt engineering and fine tuning. For example, the adjustment unit can perform prompt engineering and fine tuning based on a sample of speech. For example, the adjustment unit can perform prompt engineering and fine tuning to mimic the user's way of speaking and thinking. For example, the adjustment unit can perform prompt engineering and fine tuning based on phrases and word choices previously used by the user. In this way, the adjustment unit can enable the AI ​​to mimic the user's way of speaking and thinking by performing prompt engineering and fine tuning based on a sample of speech. Some or all of the above processing in the adjustment unit may be performed using AI, for example, or without using AI. For example, the adjustment unit can perform adjustments using an AI model for prompt engineering and fine tuning based on a sample of the user's speech.

[0036] The cloning unit can learn the user's way of speaking and thinking, and communicate using language that is characteristic of that person. For example, the cloning unit can learn the user's way of speaking and thinking, and communicate using language that is characteristic of that person. For example, the cloning unit can learn the user's way of speaking and thinking based on the user's conversation history. For example, the cloning unit can identify phrases and word choices that the user frequently uses and learn based on them. In this way, the cloning unit can achieve more natural and personalized communication by learning the user's way of speaking and thinking. Some or all of the above processes in the cloning unit may be performed using AI, for example, or not using AI. For example, the cloning unit can learn using an AI model to learn the user's way of speaking and thinking based on the user's conversation history.

[0037] The response unit can have a clone reply on behalf of the user. For example, the response unit can have a clone reply on behalf of the user. For example, the response unit can reduce the user's burden by having a clone automatically reply when the user is busy or unavailable. For example, the response unit can participate in chats with friends and family on behalf of the user and provide fresh topics. The response unit can also automatically answer specific questions. For example, the response unit can generate an appropriate response based on the user's past responses. This allows the response unit to reduce the user's burden by having a clone reply on behalf of the user. Some or all of the above processing in the response unit may be performed using AI, for example, or not using AI. For example, the response unit can generate responses using an AI model for a clone to automatically reply on behalf of the user.

[0038] The conversation continuation unit can continue a conversation as if the AI ​​had memory. For example, the conversation continuation unit can maintain the context of the conversation and achieve natural dialogue based on the user's dialogue history. For example, the conversation continuation unit can continue the conversation by reusing topics that the user has previously discussed. It can also continue the conversation by referring to the language that the user has previously preferred to use. For example, the conversation continuation unit can continue the conversation while avoiding topics that the user has previously avoided. In this way, the conversation continuation unit can achieve more natural dialogue by continuing the conversation. Some or all of the above processing in the conversation continuation unit may be performed using AI, for example, or without AI. For example, the conversation continuation unit can continue the conversation using an AI model that maintains the context of the conversation and achieves natural dialogue based on the user's dialogue history.

[0039] The adjustment unit can perform prompt engineering and fine-tuning based on utterance samples. The adjustment unit can perform prompt engineering and fine-tuning, for example, to mimic the user's way of speaking and thinking. The adjustment unit can perform prompt engineering and fine-tuning based on phrases and word choices previously used by the user. In this way, the adjustment unit can enable the AI ​​to mimic the user's way of speaking and thinking by performing prompt engineering and fine-tuning based on utterance samples. Some or all of the above-described processes in the adjustment unit may be performed using AI, for example, or without AI. For example, the adjustment unit can perform adjustments using an AI model for prompt engineering and fine-tuning based on user utterance samples.

[0040] The learning unit can analyze the user's past conversation history and learn responses to specific topics. For example, the learning unit can prioritize learning topics that the user has frequently discussed in the past. For example, the learning unit can focus on learning topics that the user has shown a strong reaction to in the past. For example, the learning unit can avoid learning topics that the user has avoided in the past. This allows the learning unit to learn the user's responses to specific topics by analyzing past conversation history, enabling more personalized responses. Some or all of the processing described above in the learning unit may be performed using AI, for example, or without AI. For example, the learning unit can learn using an AI model to learn responses to specific topics based on the user's past conversation history.

[0041] The learning unit can select the optimal learning timing by considering the user's lifestyle patterns. For example, if the user is active in the morning, the learning unit can perform learning in the morning. For example, if the user is relaxed at night, the learning unit can perform learning in the evening. For example, if the user has free time on weekends, the learning unit can perform intensive learning on weekends. In this way, by considering the user's lifestyle patterns, the learning unit can perform learning at the optimal time, enabling effective learning. Some or all of the above processing in the learning unit may be performed using AI, for example, or without AI. For example, the learning unit can perform learning using an AI model that selects the optimal learning timing based on the user's lifestyle pattern data.

[0042] The learning unit can learn regionally specific language usage by considering the user's geographical background during the learning process. For example, if the user lives in the Kansai region, the learning unit can learn the Kansai dialect. For example, if the user lives overseas, the learning unit can learn the language and dialect of that region. For example, if the user lives in a specific city, the learning unit can learn the slang specific to that city. In this way, the learning unit can learn regionally specific language usage by considering the user's geographical background, enabling more natural communication. Some or all of the above processing in the learning unit may be performed using AI, for example, or without AI. For example, the learning unit can perform learning using an AI model for learning regionally specific language usage based on the user's geographical background data.

[0043] The learning unit can analyze the user's social media activity and learn relevant information during the learning process. For example, the learning unit can learn the topics that the user frequently posts about. For example, the learning unit can learn the content of the accounts that the user follows. For example, the learning unit can learn the content of posts that the user "likes" or shares. This allows the learning unit to learn relevant information by analyzing the user's social media activity and to provide more personalized responses. Some or all of the above processing in the learning unit may be performed using AI, for example, or without AI. For example, the learning unit can perform learning using an AI model that learns relevant information based on the user's social media activity data.

[0044] The proxy unit can refer to the user's past statements and generate the most appropriate response. For example, the proxy unit can reuse phrases the user has used in the past to respond. For example, the proxy unit can refer to the language the user has preferred to use in the past to respond. For example, the proxy unit can avoid topics the user has avoided in the past to respond. This allows the proxy unit to provide more personalized responses by referring to the user's past statements. Some or all of the above processing in the proxy unit may be performed using AI, for example, or not using AI. For example, the proxy unit can use an AI model to generate the most appropriate response based on the user's past statement data to provide a response.

[0045] The proxy unit can respond at the appropriate time, taking into account the user's current situation. For example, if the user is busy, the proxy unit can provide a short, to-the-point response. For example, if the user is relaxed, the proxy unit can provide a detailed response. For example, if the user is in a hurry, the proxy unit can provide a quick response. This allows the proxy unit to respond at a more appropriate time by taking into account the user's current situation. Some or all of the above processing in the proxy unit may be performed using AI, for example, or not using AI. For example, the proxy unit can respond using an AI model that provides a response at the appropriate time based on the user's current situation data.

[0046] The proxy unit can use region-specific expressions, taking into account the user's geographical background. For example, if the user lives in the Kansai region, the proxy unit can respond using the Kansai dialect. For example, if the user lives overseas, the proxy unit can respond using the local language or dialect. For example, if the user lives in a specific city, the proxy unit can respond using slang specific to that city. This allows the proxy unit to use region-specific expressions by taking into account the user's geographical background, enabling more natural communication. Some or all of the processing described above in the proxy unit may be performed using AI, for example, or not. For example, the proxy unit can use an AI model to use region-specific expressions based on the user's geographical background data to provide a response.

[0047] The proxy unit can analyze the user's social media activity and reflect relevant information in its response. For example, the proxy unit can reflect topics that the user frequently posts about in its response. For example, the proxy unit can reflect the content of accounts that the user follows in its response. For example, the proxy unit can reflect the content of posts that the user has liked or shared in its response. In this way, by analyzing the user's social media activity, the proxy unit can reflect relevant information in its response, enabling more personalized responses. Some or all of the above processing in the proxy unit may be performed using AI, for example, or not using AI. For example, the proxy unit can use an AI model to reflect relevant information in the response based on the user's social media activity data to provide a response.

[0048] The cloning unit can generate the optimal clone by referring to the user's past conversation history during the cloning process. For example, the cloning unit can create a clone by reusing phrases the user has used in the past. For example, the cloning unit can create a clone by referring to the user's preferred language in the past. For example, the cloning unit can create a clone that avoids topics the user has avoided in the past. This allows the cloning unit to generate a more personalized clone by referring to the user's past conversation history. Some or all of the above processes in the cloning unit may be performed using AI, for example, or without AI. For example, the cloning unit can create a clone using an AI model that generates the optimal clone based on the user's past conversation history data.

[0049] The cloning unit can customize the behavior of clones by taking into account the user's lifestyle patterns during the cloning process. For example, if the user is active in the morning, the cloning unit can create a clone that is active during the morning hours. For example, if the user is relaxed at night, the cloning unit can create a clone that is active during the evening hours. For example, if the user has free time on weekends, the cloning unit can create a clone that is active intensively on weekends. This allows the cloning unit to create clones that behave more appropriately by taking into account the user's lifestyle patterns. Some or all of the above processing in the cloning unit may be performed using AI, for example, or without AI. For example, the cloning unit can create clones using an AI model to customize the clone's behavior based on the user's lifestyle pattern data.

[0050] The cloning unit can reflect the user's geographical background and regional slang when creating clones. For example, if the user lives in the Kansai region, the cloning unit can create a clone that speaks in the Kansai dialect. For example, if the user lives overseas, the cloning unit can create a clone that speaks in the local language or dialect of that region. For example, if the user lives in a specific city, the cloning unit can create a clone that speaks in the slang specific to that city. This makes it possible for the cloning unit to create clones that reflect regional slang by considering the user's geographical background. Some or all of the above processing in the cloning unit may be performed using AI, for example, or not. For example, the cloning unit can create clones using an AI model that reflects regional slang based on the user's geographical background data.

[0051] The cloning unit can analyze the user's social media activity during cloning and reflect relevant information in the clone. For example, the cloning unit can reflect topics that the user frequently posts about in the clone. For example, the cloning unit can reflect the content of accounts that the user follows in the clone. For example, the cloning unit can reflect the content of posts that the user has liked or shared in the clone. This makes it possible for the cloning unit to create clones that reflect relevant information by analyzing the user's social media activity. Some or all of the above processing in the cloning unit may be performed using AI, for example, or without AI. For example, the cloning unit can create clones using an AI model that reflects relevant information in the clone based on the user's social media activity data.

[0052] The response unit can refer to the user's past statements and generate the most appropriate response. For example, the response unit can reuse phrases the user has used in the past. For example, the response unit can refer to the user's preferred language in the past. For example, the response unit can avoid topics the user has avoided in the past. This allows the response unit to provide more personalized responses by referring to the user's past statements. Some or all of the above processing in the response unit may be performed using AI, for example, or without AI. For example, the response unit can use an AI model to generate the most appropriate response based on the user's past statement data.

[0053] The response unit can respond at an appropriate time, taking into account the user's current situation. For example, if the user is busy, the response unit can provide a short, concise response. For example, if the user is relaxed, the response unit can provide a detailed response. For example, if the user is in a hurry, the response unit can provide a quick response. This allows the response unit to respond at a more appropriate time by taking into account the user's current situation. Some or all of the above processing in the response unit may be performed using AI, for example, or without AI. For example, the response unit can respond using an AI model that provides a response at an appropriate time based on the user's current situation data.

[0054] The response unit can use region-specific expressions, taking into account the user's geographical background. For example, if the user lives in the Kansai region, the response unit can respond using the Kansai dialect. For example, if the user lives overseas, the response unit can respond using the local language or dialect of that region. For example, if the user lives in a specific city, the response unit can respond using slang specific to that city. This allows the response unit to use region-specific expressions by taking into account the user's geographical background, enabling more natural communication. Some or all of the processing described above in the response unit may be performed using AI, for example, or not. For example, the response unit can use an AI model to use region-specific expressions based on the user's geographical background data to provide a response.

[0055] The conversation continuation unit can refer to the user's past dialogue history and generate an optimal conversation flow. For example, the conversation continuation unit can continue the conversation by reusing topics the user has previously discussed. For example, the conversation continuation unit can continue the conversation by referring to the vocabulary the user has previously preferred to use. For example, the conversation continuation unit can continue the conversation while avoiding topics the user has previously avoided. In this way, the conversation continuation unit can provide a more personalized conversation by referring to the user's past dialogue history. Some or all of the above processing in the conversation continuation unit may be performed using AI, for example, or without AI. For example, the conversation continuation unit can continue the conversation using an AI model to generate an optimal conversation flow based on the user's past dialogue history data.

[0056] The conversation continuation unit can continue the conversation at an appropriate time, taking into account the user's current situation. For example, if the user is busy, the conversation continuation unit can continue with a short, to-the-point conversation. For example, if the user is relaxed, the conversation continuation unit can continue with a detailed conversation. For example, if the user is in a hurry, the conversation continuation unit can continue the conversation quickly. In this way, the conversation continuation unit can continue the conversation at a more appropriate time by taking into account the user's current situation. Some or all of the above processing in the conversation continuation unit may be performed using AI, for example, or not using AI. For example, the conversation continuation unit can continue the conversation using an AI model that takes the user's current situation data into account and continues the conversation at an appropriate time.

[0057] The conversation continuation unit can incorporate region-specific topics by considering the user's geographical background. For example, if the user lives in the Kansai region, the conversation continuation unit can incorporate Kansai-related topics to continue the conversation. For example, if the user lives overseas, the conversation continuation unit can incorporate topics of that region to continue the conversation. For example, if the user lives in a specific city, the conversation continuation unit can incorporate topics specific to that city to continue the conversation. In this way, the conversation continuation unit can incorporate region-specific topics into conversations by considering the user's geographical background. Some or all of the above processing in the conversation continuation unit may be performed using AI, for example, or without AI. For example, the conversation continuation unit can continue the conversation using an AI model that incorporates region-specific topics based on the user's geographical background data.

[0058] The adjustment unit can refer to the user's past utterance samples and make optimal adjustments. For example, the adjustment unit can reuse phrases the user has used in the past to make adjustments. For example, the adjustment unit can refer to the wording the user has preferred to use in the past to make adjustments. For example, the adjustment unit can avoid topics the user has avoided in the past to make adjustments. This allows the adjustment unit to make more personalized adjustments by referring to the user's past utterance samples. Some or all of the above processing in the adjustment unit may be performed using AI, for example, or not using AI. For example, the adjustment unit can make adjustments using an AI model that makes optimal adjustments based on the user's past utterance sample data.

[0059] The adjustment unit can make adjustments at the appropriate time, taking into account the user's current situation. For example, if the user is busy, the adjustment unit can make short, concise adjustments. For example, if the user is relaxed, the adjustment unit can make detailed adjustments. For example, if the user is in a hurry, the adjustment unit can make quick adjustments. This allows the adjustment unit to make adjustments at a more appropriate time by taking into account the user's current situation. Some or all of the above processing in the adjustment unit may be performed using AI, for example, or without AI. For example, the adjustment unit can make adjustments using an AI model that makes adjustments at the appropriate time based on the user's current situation data.

[0060] The adjustment unit can take the user's geographical background into consideration and reflect regionally specific language usage. For example, if the user lives in the Kansai region, the adjustment unit can use the Kansai dialect. For example, if the user lives overseas, the adjustment unit can use the language or dialect of that region. For example, if the user lives in a specific city, the adjustment unit can use the slang specific to that city. In this way, the adjustment unit can reflect regionally specific language usage by taking the user's geographical background into consideration. Some or all of the above processing in the adjustment unit may be performed using AI, for example, or not using AI. For example, the adjustment unit can use an AI model to reflect regionally specific language usage based on the user's geographical background data to perform adjustments.

[0061] The system according to the embodiment is not limited to the example described above, and various modifications are possible, for example, as follows.

[0062] The bot system can also include a health management unit that monitors the user's health status. This unit can, for example, monitor the user's heart rate and sleep patterns to understand their health condition. If the user is experiencing stress, the health management unit can suggest relaxing activities. Furthermore, it can manage the user's diet and exercise records to support healthy lifestyle habits. In this way, the health management unit can support the user's health maintenance by monitoring their health status and providing appropriate advice.

[0063] The bot system can also be equipped with a hobby learning unit that learns the user's hobbies and interests. For example, the hobby learning unit can learn about hobbies and interests the user has discussed in the past and provide relevant information based on that. If the user is a movie lover, for instance, the hobby learning unit can provide information on new movies. Similarly, if the user is a music lover, the hobby learning unit can provide information on the latest music trends and concerts. This allows the hobby learning unit to provide more personalized information based on the user's hobbies and interests.

[0064] The bot system can also include a learning management unit to manage the user's learning progress. For example, the learning management unit can record what the user is learning and manage their progress. If the user is learning English, for instance, the learning management unit can record their learning progress and suggest an appropriate learning plan. Furthermore, if the user is studying for an exam, the learning management unit can manage their study schedule leading up to the exam date. In this way, the learning management unit can manage the user's learning progress and support effective learning.

[0065] The bot system can also include a travel planning unit to support the user's travel planning. For example, the travel planning unit can suggest travel plans based on the places the user wants to visit and activities they are interested in. If the user wants to enjoy nature, the travel planning unit can suggest tourist destinations rich in natural beauty. Similarly, if the user is interested in history and culture, the travel planning unit can suggest historical tourist destinations and cultural events. This allows the travel planning unit to propose the optimal travel plan based on the user's interests and preferences.

[0066] The bot system can also be equipped with a task management unit to further improve the user's work efficiency. This unit can, for example, manage the user's schedule and set task priorities. It can also remind the user of other tasks so that they can focus on important projects. Furthermore, it can provide time management advice to help the user work efficiently. In this way, the task management unit can improve the user's work efficiency and reduce stress.

[0067] The following briefly describes the processing flow for example form 1.

[0068] Step 1: The learning unit learns the user's language use and thought patterns. For example, it learns the user's unique language use and thought patterns through dialogue with the user. The learning unit can analyze the user's language use and thought patterns using machine learning algorithms. It also identifies frequently used phrases and language use based on the user's dialogue history and uses that as a basis for learning. Step 2: The proxy unit acts as a proxy for the user based on the information learned by the learning unit. For example, it can generate responses by mimicking the user's speech patterns and thought processes. The proxy unit can participate in chats with friends and family on behalf of the user and provide fresh topics. It can also automatically answer specific questions. For example, it can generate appropriate responses based on the user's past answers. Step 3: The cloning unit creates a clone of the user. For example, it can learn the user's speaking style and way of thinking, and communicate using language that is characteristic of that person. It can also use an automated response function, allowing the clone to reply on the user's behalf. For example, it can participate in chats with friends and family on the user's behalf and provide fresh topics of conversation.

[0069] (Example of form 2) The Bot system according to an embodiment of the present invention is a system that learns a user's speech patterns and thought patterns and functions as a proxy for the user over time. This Bot system learns the user's unique speech patterns and thought patterns through interaction with the user and evolves to function as a proxy for the user over time. Technically, it utilizes multimodal AI by making full use of AI APIs, fine-tuning, prompt engineering, etc. This allows the user to enjoy a more personalized and richer communication experience. For example, the Bot system learns the user's unique speech patterns and thought patterns through interaction with the user. Next, based on the learned information, it evolves to function as a proxy for the user. For example, it can not only provide fresh topics for chats with friends and family and liven up conversations, but also automatically answer specific questions. In this way, user satisfaction is improved and deeper relationships can be formed. Furthermore, this Bot system can also create a clone of the user. It can learn the user's way of speaking and thinking and communicate using language that is characteristic of that person. It is also possible for the clone to reply on behalf of the user using the automatic response function. In addition, by using AI APIs, the AI ​​can engage in continuous conversations as if it had memory. This allows users to enjoy more natural conversations. Furthermore, by providing sample statements, prompt engineering and fine-tuning can be performed, allowing the AI ​​to mimic the user's way of speaking and thinking. This bot system can create clones not only of the user but also of other people. For example, you can have conversations and play with a clone of your friend. Thus, this invention realizes a bot system that responds to the individual needs and personalities of users, providing a more personalized communication experience. By learning the user's language and thought patterns, the bot system can function as a proxy and create clones to provide a personalized communication experience.

[0070] The Bot system according to this embodiment comprises a learning unit, an agency unit, and a cloning unit. The learning unit learns the user's speech patterns and thought patterns. For example, the learning unit learns the user's unique speech patterns and thought patterns through dialogue with the user. The learning unit can analyze the user's speech patterns and thought patterns using, for example, a machine learning algorithm. The learning unit can also learn the user's speech patterns and thought patterns based on the user's dialogue history. For example, the learning unit identifies phrases and speech patterns that the user frequently uses and learns based on them. The agency unit functions as a proxy for the user based on the information learned by the learning unit. For example, the agency unit can generate responses by mimicking the user's speech patterns and thought patterns. For example, the agency unit can participate in chats with friends and family on behalf of the user and provide fresh topics. The agency unit can also automatically answer specific questions. For example, the agency unit generates appropriate responses based on the user's past answers. The cloning unit creates a clone of the user. The cloning unit can, for example, learn the user's way of speaking and thinking, and communicate using language that reflects that person's personality. Furthermore, the cloning unit can use an automated response function to have a clone reply on its behalf. For instance, the cloning unit can participate in chats with friends and family on the user's behalf, providing fresh topics of conversation. Thus, the bot system according to this embodiment can learn the user's language and thought patterns, function as a proxy, and create clones to provide a personalized communication experience.

[0071] The learning unit learns the user's language use and thought patterns. Specifically, the learning unit learns the user's unique language use and thought patterns through dialogue with the user. For example, it analyzes in detail the phrases and word choices the user uses daily, sentence structure, and even how they express emotions. The learning unit uses machine learning algorithms to analyze this data and model the user's language use and thought patterns. Specifically, it utilizes natural language processing (NLP) technology to tokenize the user's utterances and generate contextual vectors to understand the context. Furthermore, it uses recurrent neural networks (RNNs) and transformer models to learn the temporal flow and relationships of the user's utterances. Based on the user's dialogue history, the learning unit can continuously learn and update the user's language use and thought patterns. For example, it identifies phrases the user frequently uses and response patterns in specific situations and uses them as a basis for learning. As a result, the learning unit can always reflect the user's latest language use and thought patterns, generating more natural and personalized responses.

[0072] The proxy unit functions as a proxy for the user based on information learned by the learning unit. Specifically, the proxy unit generates responses by mimicking the user's speech patterns and thought patterns. For example, it can participate in chats with friends and family on behalf of the user and provide fresh topics of conversation. The proxy unit refers to past conversation history to generate appropriate responses based on the user's past answers. This allows for responses tailored to the user's style and preferences. The proxy unit uses generative AI to mimic the user's speech patterns and thought patterns. Specifically, the generative AI learns the user's speech patterns and generates new statements based on them. For example, it learns the user's frequently used phrases and word choices and generates natural responses based on them. The proxy unit can also automatically answer specific questions. For example, it generates appropriate responses based on the user's past answers. This allows the proxy unit to provide quick and appropriate responses on behalf of the user. Furthermore, the proxy unit can understand the user's intentions and emotions and generate responses accordingly. For example, if the user expresses gratitude, the proxy unit will generate appropriate words of thanks. This allows the proxy unit to provide natural and personalized communication on behalf of the user.

[0073] The cloning unit creates a clone of the user. Specifically, the cloning unit learns the user's way of speaking and thinking, and can communicate using language that is characteristic of that person. The cloning unit generates a clone of the user based on the user's dialogue history and data collected by the learning unit. For example, the cloning unit identifies phrases and word choices that the user frequently uses and creates a clone based on that. The cloning unit uses generative AI to mimic the user's language and thought patterns. Specifically, the generative AI learns the user's speech patterns and generates new statements based on them. For example, it learns the phrases and word choices that the user frequently uses and generates natural responses based on that. For example, the cloning unit can participate in chats with friends and family on behalf of the user and provide fresh topics. The cloning unit can also use an auto-responder function so that the clone replies on behalf of the user. For example, the cloning unit can participate in chats with friends and family on behalf of the user and provide fresh topics. In this way, the cloning unit can mimic the user's language and thought patterns and provide natural and personalized communication. Furthermore, the cloning unit can continuously learn from and update the user's clone. For example, it can update the clone's vocabulary and thought patterns based on the user's most recent conversation history. This allows the cloning unit to always provide a clone based on the latest information, enabling it to provide natural and personalized communication on behalf of the user.

[0074] The response unit can provide an automated response function. For example, the response unit can have a clone automatically respond on behalf of the user. For example, the response unit can reduce the user's burden by having the clone automatically respond when the user is busy or unable to attend to the user. For example, the response unit can participate in chats with friends and family on behalf of the user and provide fresh topics. The response unit can also automatically answer specific questions. For example, the response unit can generate an appropriate response based on the user's past responses. This allows the response unit to reduce the user's burden by having the clone automatically respond on behalf of the user. Some or all of the above processing in the response unit may be performed using AI, for example, or not using AI. For example, the response unit can generate responses using an AI model for the clone to automatically respond on behalf of the user.

[0075] The conversation continuation unit can engage in continuous conversation. For example, the conversation continuation unit can engage in continuous conversation as if the AI ​​had memory. For example, the conversation continuation unit can maintain the context of the conversation and achieve natural dialogue based on the user's dialogue history. For example, the conversation continuation unit can continue the conversation by reusing topics that the user has previously discussed. The conversation continuation unit can also continue the conversation by referring to the language that the user has previously preferred to use. For example, the conversation continuation unit can continue the conversation while avoiding topics that the user has previously avoided. In this way, the conversation continuation unit can achieve natural dialogue by engaging in continuous conversation as if the AI ​​had memory. Some or all of the above processing in the conversation continuation unit may be performed using AI, for example, or without AI. For example, the conversation continuation unit can continue the conversation using an AI model that maintains the context of the conversation and achieves natural dialogue based on the user's dialogue history.

[0076] The adjustment unit can perform prompt engineering and fine tuning. For example, the adjustment unit can perform prompt engineering and fine tuning based on a sample of speech. For example, the adjustment unit can perform prompt engineering and fine tuning to mimic the user's way of speaking and thinking. For example, the adjustment unit can perform prompt engineering and fine tuning based on phrases and word choices previously used by the user. In this way, the adjustment unit can enable the AI ​​to mimic the user's way of speaking and thinking by performing prompt engineering and fine tuning based on a sample of speech. Some or all of the above processing in the adjustment unit may be performed using AI, for example, or without using AI. For example, the adjustment unit can perform adjustments using an AI model for prompt engineering and fine tuning based on a sample of the user's speech.

[0077] The cloning unit can learn the user's way of speaking and thinking, and communicate using language that is characteristic of that person. For example, the cloning unit can learn the user's way of speaking and thinking, and communicate using language that is characteristic of that person. For example, the cloning unit can learn the user's way of speaking and thinking based on the user's conversation history. For example, the cloning unit can identify phrases and word choices that the user frequently uses and learn based on them. In this way, the cloning unit can achieve more natural and personalized communication by learning the user's way of speaking and thinking. Some or all of the above processes in the cloning unit may be performed using AI, for example, or not using AI. For example, the cloning unit can learn using an AI model to learn the user's way of speaking and thinking based on the user's conversation history.

[0078] The response unit can have a clone reply on behalf of the user. For example, the response unit can have a clone reply on behalf of the user. For example, the response unit can reduce the user's burden by having a clone automatically reply when the user is busy or unavailable. For example, the response unit can participate in chats with friends and family on behalf of the user and provide fresh topics. The response unit can also automatically answer specific questions. For example, the response unit can generate an appropriate response based on the user's past responses. This allows the response unit to reduce the user's burden by having a clone reply on behalf of the user. Some or all of the above processing in the response unit may be performed using AI, for example, or not using AI. For example, the response unit can generate responses using an AI model for a clone to automatically reply on behalf of the user.

[0079] The conversation continuation unit can continue a conversation as if the AI ​​had memory. For example, the conversation continuation unit can maintain the context of the conversation and achieve natural dialogue based on the user's dialogue history. For example, the conversation continuation unit can continue the conversation by reusing topics that the user has previously discussed. It can also continue the conversation by referring to the language that the user has previously preferred to use. For example, the conversation continuation unit can continue the conversation while avoiding topics that the user has previously avoided. In this way, the conversation continuation unit can achieve more natural dialogue by continuing the conversation. Some or all of the above processing in the conversation continuation unit may be performed using AI, for example, or without AI. For example, the conversation continuation unit can continue the conversation using an AI model that maintains the context of the conversation and achieves natural dialogue based on the user's dialogue history.

[0080] The adjustment unit can perform prompt engineering and fine-tuning based on utterance samples. The adjustment unit can perform prompt engineering and fine-tuning, for example, to mimic the user's way of speaking and thinking. The adjustment unit can perform prompt engineering and fine-tuning based on phrases and word choices previously used by the user. In this way, the adjustment unit can enable the AI ​​to mimic the user's way of speaking and thinking by performing prompt engineering and fine-tuning based on utterance samples. Some or all of the above-described processes in the adjustment unit may be performed using AI, for example, or without AI. For example, the adjustment unit can perform adjustments using an AI model for prompt engineering and fine-tuning based on user utterance samples.

[0081] The learning unit can estimate the user's emotions and adjust learning priorities based on the estimated emotions. For example, if the user is stressed, the learning unit can prioritize learning relaxing topics. For example, if the user is excited, the learning unit can prioritize learning interesting topics. For example, if the user is tired, the learning unit can prioritize learning simple and easy-to-understand content. This allows the learning unit to learn more effectively by adjusting learning priorities based on the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the learning unit may be performed using AI, for example, or without AI. For example, the learning unit can perform learning using an AI model to adjust learning priorities based on user emotion data.

[0082] The learning unit can analyze the user's past conversation history and learn responses to specific topics. For example, the learning unit can prioritize learning topics that the user has frequently discussed in the past. For example, the learning unit can focus on learning topics that the user has shown a strong reaction to in the past. For example, the learning unit can avoid learning topics that the user has avoided in the past. This allows the learning unit to learn the user's responses to specific topics by analyzing past conversation history, enabling more personalized responses. Some or all of the processing described above in the learning unit may be performed using AI, for example, or without AI. For example, the learning unit can learn using an AI model to learn responses to specific topics based on the user's past conversation history.

[0083] The learning unit can select the optimal learning timing by considering the user's lifestyle patterns. For example, if the user is active in the morning, the learning unit can perform learning in the morning. For example, if the user is relaxed at night, the learning unit can perform learning in the evening. For example, if the user has free time on weekends, the learning unit can perform intensive learning on weekends. In this way, by considering the user's lifestyle patterns, the learning unit can perform learning at the optimal time, enabling effective learning. Some or all of the above processing in the learning unit may be performed using AI, for example, or without AI. For example, the learning unit can perform learning using an AI model that selects the optimal learning timing based on the user's lifestyle pattern data.

[0084] The learning unit can estimate the user's emotions and customize the learning content based on the estimated emotions. For example, if the user is sad, the learning unit can learn words of encouragement or cheerful topics. For example, if the user is happy, the learning unit can learn words of empathy and congratulations. For example, if the user is angry, the learning unit can learn words to respond calmly. This allows the learning unit to learn more effectively by customizing the learning content based on the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the learning unit may be performed using AI, for example, or without AI. For example, the learning unit can learn using an AI model to customize the learning content based on the user's emotion data.

[0085] The learning unit can learn regionally specific language usage by considering the user's geographical background during the learning process. For example, if the user lives in the Kansai region, the learning unit can learn the Kansai dialect. For example, if the user lives overseas, the learning unit can learn the language and dialect of that region. For example, if the user lives in a specific city, the learning unit can learn the slang specific to that city. In this way, the learning unit can learn regionally specific language usage by considering the user's geographical background, enabling more natural communication. Some or all of the above processing in the learning unit may be performed using AI, for example, or without AI. For example, the learning unit can perform learning using an AI model for learning regionally specific language usage based on the user's geographical background data.

[0086] The learning unit can analyze the user's social media activity and learn relevant information during the learning process. For example, the learning unit can learn the topics that the user frequently posts about. For example, the learning unit can learn the content of the accounts that the user follows. For example, the learning unit can learn the content of posts that the user "likes" or shares. This allows the learning unit to learn relevant information by analyzing the user's social media activity and to provide more personalized responses. Some or all of the above processing in the learning unit may be performed using AI, for example, or without AI. For example, the learning unit can perform learning using an AI model that learns relevant information based on the user's social media activity data.

[0087] The surrogate unit can estimate the user's emotions and adjust its surrogate expression based on the estimated emotions. For example, if the user is sad, the surrogate unit can respond with gentle language. For example, if the user is happy, the surrogate unit can respond with cheerful language. For example, if the user is angry, the surrogate unit can respond with calm and composed language. This allows the surrogate unit to provide a more appropriate response by adjusting its surrogate expression based on the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the surrogate unit may be performed using AI, for example, or without AI. For example, the surrogate unit can respond using an AI model to adjust its surrogate expression based on the user's emotion data.

[0088] The proxy unit can refer to the user's past statements and generate the most appropriate response. For example, the proxy unit can reuse phrases the user has used in the past to respond. For example, the proxy unit can refer to the language the user has preferred to use in the past to respond. For example, the proxy unit can avoid topics the user has avoided in the past to respond. This allows the proxy unit to provide more personalized responses by referring to the user's past statements. Some or all of the above processing in the proxy unit may be performed using AI, for example, or not using AI. For example, the proxy unit can use an AI model to generate the most appropriate response based on the user's past statement data to provide a response.

[0089] The proxy unit can respond at the appropriate time, taking into account the user's current situation. For example, if the user is busy, the proxy unit can provide a short, to-the-point response. For example, if the user is relaxed, the proxy unit can provide a detailed response. For example, if the user is in a hurry, the proxy unit can provide a quick response. This allows the proxy unit to respond at a more appropriate time by taking into account the user's current situation. Some or all of the above processing in the proxy unit may be performed using AI, for example, or not using AI. For example, the proxy unit can respond using an AI model that provides a response at the appropriate time based on the user's current situation data.

[0090] The proxy unit can estimate the user's emotions and customize the proxy's response based on the estimated emotions. For example, if the user is sad, the proxy unit can provide a response that includes words of encouragement. For example, if the user is happy, the proxy unit can provide a response that includes words of empathy and congratulations. For example, if the user is angry, the proxy unit can provide a response that includes words to calmly address the situation. This allows the proxy unit to provide a more appropriate response by customizing the proxy's response based on the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the proxy unit may be performed using AI, for example, or not using AI. For example, the proxy unit can provide a response using an AI model to customize the proxy's response based on the user's emotion data.

[0091] The proxy unit can use region-specific expressions, taking into account the user's geographical background. For example, if the user lives in the Kansai region, the proxy unit can respond using the Kansai dialect. For example, if the user lives overseas, the proxy unit can respond using the local language or dialect. For example, if the user lives in a specific city, the proxy unit can respond using slang specific to that city. This allows the proxy unit to use region-specific expressions by taking into account the user's geographical background, enabling more natural communication. Some or all of the processing described above in the proxy unit may be performed using AI, for example, or not. For example, the proxy unit can use an AI model to use region-specific expressions based on the user's geographical background data to provide a response.

[0092] The proxy unit can analyze the user's social media activity and reflect relevant information in its response. For example, the proxy unit can reflect topics that the user frequently posts about in its response. For example, the proxy unit can reflect the content of accounts that the user follows in its response. For example, the proxy unit can reflect the content of posts that the user has liked or shared in its response. In this way, by analyzing the user's social media activity, the proxy unit can reflect relevant information in its response, enabling more personalized responses. Some or all of the above processing in the proxy unit may be performed using AI, for example, or not using AI. For example, the proxy unit can use an AI model to reflect relevant information in the response based on the user's social media activity data to provide a response.

[0093] The cloning unit can estimate the user's emotions and adjust the cloning method based on the estimated emotions. For example, if the user is relaxed, the cloning unit can create a clone with relaxed language. For example, if the user is excited, the cloning unit can create a clone with excited language. For example, if the user is sad, the cloning unit can create a clone with gentle language. This allows the cloning unit to create more appropriate clones by adjusting the cloning method based on the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the cloning unit may be performed using AI, for example, or without AI. For example, the cloning unit can create clones using an AI model to adjust the cloning method based on the user's emotion data.

[0094] The cloning unit can generate the optimal clone by referring to the user's past conversation history during the cloning process. For example, the cloning unit can create a clone by reusing phrases the user has used in the past. For example, the cloning unit can create a clone by referring to the user's preferred language in the past. For example, the cloning unit can create a clone that avoids topics the user has avoided in the past. This allows the cloning unit to generate a more personalized clone by referring to the user's past conversation history. Some or all of the above processes in the cloning unit may be performed using AI, for example, or without AI. For example, the cloning unit can create a clone using an AI model that generates the optimal clone based on the user's past conversation history data.

[0095] The cloning unit can customize the behavior of clones by taking into account the user's lifestyle patterns during the cloning process. For example, if the user is active in the morning, the cloning unit can create a clone that is active during the morning hours. For example, if the user is relaxed at night, the cloning unit can create a clone that is active during the evening hours. For example, if the user has free time on weekends, the cloning unit can create a clone that is active intensively on weekends. This allows the cloning unit to create clones that behave more appropriately by taking into account the user's lifestyle patterns. Some or all of the above processing in the cloning unit may be performed using AI, for example, or without AI. For example, the cloning unit can create clones using an AI model to customize the clone's behavior based on the user's lifestyle pattern data.

[0096] The cloning unit can estimate the user's emotions and customize the clone's response based on the estimated emotions. For example, if the user is sad, the cloning unit can create a clone that includes words of encouragement. For example, if the user is happy, the cloning unit can create a clone that includes words of empathy and congratulations. For example, if the user is angry, the cloning unit can create a clone that includes words to respond calmly. This allows the cloning unit to provide more appropriate responses by customizing the clone's response based on the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, an emotion engine or a generative AI. The generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the cloning unit may be performed using AI, for example, or without AI. For example, the cloning unit can create clones using an AI model to customize the clone's response based on the user's emotion data.

[0097] The cloning unit can reflect the user's geographical background and regional slang when creating clones. For example, if the user lives in the Kansai region, the cloning unit can create a clone that speaks in the Kansai dialect. For example, if the user lives overseas, the cloning unit can create a clone that speaks in the local language or dialect of that region. For example, if the user lives in a specific city, the cloning unit can create a clone that speaks in the slang specific to that city. This makes it possible for the cloning unit to create clones that reflect regional slang by considering the user's geographical background. Some or all of the above processing in the cloning unit may be performed using AI, for example, or not. For example, the cloning unit can create clones using an AI model that reflects regional slang based on the user's geographical background data.

[0098] The cloning unit can analyze the user's social media activity during cloning and reflect relevant information in the clone. For example, the cloning unit can reflect topics that the user frequently posts about in the clone. For example, the cloning unit can reflect the content of accounts that the user follows in the clone. For example, the cloning unit can reflect the content of posts that the user has liked or shared in the clone. This makes it possible for the cloning unit to create clones that reflect relevant information by analyzing the user's social media activity. Some or all of the above processing in the cloning unit may be performed using AI, for example, or without AI. For example, the cloning unit can create clones using an AI model that reflects relevant information in the clone based on the user's social media activity data.

[0099] The response unit can estimate the user's emotions and adjust the way it expresses its response based on the estimated emotions. For example, if the user is sad, the response unit can respond using gentle language. For example, if the user is happy, the response unit can respond using cheerful language. For example, if the user is angry, the response unit can respond using calm and composed language. In this way, the response unit can provide a more appropriate response by adjusting the way it expresses its response based on the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the response unit may be performed using AI, for example, or without AI. For example, the response unit can provide a response using an AI model to adjust the way it expresses its response based on the user's emotion data.

[0100] The response unit can refer to the user's past statements and generate the most appropriate response. For example, the response unit can reuse phrases the user has used in the past. For example, the response unit can refer to the user's preferred language in the past. For example, the response unit can avoid topics the user has avoided in the past. This allows the response unit to provide more personalized responses by referring to the user's past statements. Some or all of the above processing in the response unit may be performed using AI, for example, or without AI. For example, the response unit can use an AI model to generate the most appropriate response based on the user's past statement data.

[0101] The response unit can respond at an appropriate time, taking into account the user's current situation. For example, if the user is busy, the response unit can provide a short, concise response. For example, if the user is relaxed, the response unit can provide a detailed response. For example, if the user is in a hurry, the response unit can provide a quick response. This allows the response unit to respond at a more appropriate time by taking into account the user's current situation. Some or all of the above processing in the response unit may be performed using AI, for example, or without AI. For example, the response unit can respond using an AI model that provides a response at an appropriate time based on the user's current situation data.

[0102] The response unit can estimate the user's emotions and customize the content of the response based on the estimated emotions. For example, if the user is sad, the response unit can provide a response that includes words of encouragement. For example, if the user is happy, the response unit can provide a response that includes words of empathy and congratulations. For example, if the user is angry, the response unit can provide a response that includes words to calmly respond. In this way, the response unit can provide a more appropriate response by customizing the content of the response based on the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the response unit may be performed using AI, for example, or without AI. For example, the response unit can provide a response using an AI model to customize the content of the response based on the user's emotion data.

[0103] The response unit can use region-specific expressions, taking into account the user's geographical background. For example, if the user lives in the Kansai region, the response unit can respond using the Kansai dialect. For example, if the user lives overseas, the response unit can respond using the local language or dialect of that region. For example, if the user lives in a specific city, the response unit can respond using slang specific to that city. This allows the response unit to use region-specific expressions by taking into account the user's geographical background, enabling more natural communication. Some or all of the processing described above in the response unit may be performed using AI, for example, or not. For example, the response unit can use an AI model to use region-specific expressions based on the user's geographical background data to provide a response.

[0104] The conversation continuation unit can estimate the user's emotions and adjust how it continues the conversation based on those emotions. For example, if the user is sad, the conversation continuation unit can continue the conversation using gentle language. For example, if the user is happy, the conversation continuation unit can continue the conversation using cheerful language. For example, if the user is angry, the conversation continuation unit can continue the conversation using calm and composed language. In this way, the conversation continuation unit can have a more appropriate conversation by adjusting how it continues the conversation based on the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the conversation continuation unit may be performed using AI, for example, or without AI. For example, the conversation continuation unit can continue the conversation using an AI model to adjust how it continues the conversation based on the user's emotion data.

[0105] The conversation continuation unit can refer to the user's past dialogue history and generate an optimal conversation flow. For example, the conversation continuation unit can continue the conversation by reusing topics the user has previously discussed. For example, the conversation continuation unit can continue the conversation by referring to the vocabulary the user has previously preferred to use. For example, the conversation continuation unit can continue the conversation while avoiding topics the user has previously avoided. In this way, the conversation continuation unit can provide a more personalized conversation by referring to the user's past dialogue history. Some or all of the above processing in the conversation continuation unit may be performed using AI, for example, or without AI. For example, the conversation continuation unit can continue the conversation using an AI model to generate an optimal conversation flow based on the user's past dialogue history data.

[0106] The conversation continuation unit can continue the conversation at an appropriate time, taking into account the user's current situation. For example, if the user is busy, the conversation continuation unit can continue with a short, to-the-point conversation. For example, if the user is relaxed, the conversation continuation unit can continue with a detailed conversation. For example, if the user is in a hurry, the conversation continuation unit can continue the conversation quickly. In this way, the conversation continuation unit can continue the conversation at a more appropriate time by taking into account the user's current situation. Some or all of the above processing in the conversation continuation unit may be performed using AI, for example, or not using AI. For example, the conversation continuation unit can continue the conversation using an AI model that takes the user's current situation data into account and continues the conversation at an appropriate time.

[0107] The conversation continuation unit can estimate the user's emotions and customize the conversation content based on the estimated emotions. For example, if the user is sad, the conversation continuation unit can continue the conversation with words of encouragement. For example, if the user is happy, the conversation continuation unit can continue the conversation with words of empathy and congratulations. For example, if the user is angry, the conversation continuation unit can continue the conversation with words to respond calmly. In this way, the conversation continuation unit can have a more appropriate conversation by customizing the conversation content based on the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI is a text generation AI (e.g., LLM) or a multimodal generation AI, but is not limited to such examples. Some or all of the above processing in the conversation continuation unit may be performed using AI, for example, or without AI. For example, the conversation continuation unit can continue the conversation using an AI model for customizing the conversation content based on the user's emotion data.

[0108] The conversation continuation unit can incorporate region-specific topics by considering the user's geographical background. For example, if the user lives in the Kansai region, the conversation continuation unit can incorporate Kansai-related topics to continue the conversation. For example, if the user lives overseas, the conversation continuation unit can incorporate topics of that region to continue the conversation. For example, if the user lives in a specific city, the conversation continuation unit can incorporate topics specific to that city to continue the conversation. In this way, the conversation continuation unit can incorporate region-specific topics into conversations by considering the user's geographical background. Some or all of the above processing in the conversation continuation unit may be performed using AI, for example, or without AI. For example, the conversation continuation unit can continue the conversation using an AI model that incorporates region-specific topics based on the user's geographical background data.

[0109] The adjustment unit can estimate the user's emotions and adjust the prompt engineering and fine-tuning methods based on the estimated user emotions. For example, if the user is relaxed, the adjustment unit can use prompts with relaxing content. For example, if the user is excited, the adjustment unit can use prompts with exciting content. For example, if the user is sad, the adjustment unit can use prompts with gentle content. This allows the adjustment unit to make more appropriate adjustments by adjusting the prompt engineering and fine-tuning methods based on the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the adjustment unit may be performed using AI, for example, or without AI. For example, the adjustment unit can make adjustments using an AI model for adjusting prompt engineering and fine-tuning methods based on user emotion data.

[0110] The adjustment unit can refer to the user's past utterance samples and make optimal adjustments. For example, the adjustment unit can reuse phrases the user has used in the past to make adjustments. For example, the adjustment unit can refer to the wording the user has preferred to use in the past to make adjustments. For example, the adjustment unit can avoid topics the user has avoided in the past to make adjustments. This allows the adjustment unit to make more personalized adjustments by referring to the user's past utterance samples. Some or all of the above processing in the adjustment unit may be performed using AI, for example, or not using AI. For example, the adjustment unit can make adjustments using an AI model that makes optimal adjustments based on the user's past utterance sample data.

[0111] The adjustment unit can make adjustments at the appropriate time, taking into account the user's current situation. For example, if the user is busy, the adjustment unit can make short, concise adjustments. For example, if the user is relaxed, the adjustment unit can make detailed adjustments. For example, if the user is in a hurry, the adjustment unit can make quick adjustments. This allows the adjustment unit to make adjustments at a more appropriate time by taking into account the user's current situation. Some or all of the above processing in the adjustment unit may be performed using AI, for example, or without AI. For example, the adjustment unit can make adjustments using an AI model that makes adjustments at the appropriate time based on the user's current situation data.

[0112] The adjustment unit can estimate the user's emotions and customize the adjustment content based on the estimated user emotions. For example, if the user is sad, the adjustment unit can make gentle adjustments. For example, if the user is happy, the adjustment unit can make cheerful adjustments. For example, if the user is angry, the adjustment unit can make calm adjustments. In this way, the adjustment unit can make more appropriate adjustments by customizing the adjustment content based on the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the adjustment unit may be performed using AI, for example, or without AI. For example, the adjustment unit can make adjustments using an AI model to customize the adjustment content based on the user's emotion data.

[0113] The adjustment unit can take the user's geographical background into consideration and reflect regionally specific language usage. For example, if the user lives in the Kansai region, the adjustment unit can use the Kansai dialect. For example, if the user lives overseas, the adjustment unit can use the language or dialect of that region. For example, if the user lives in a specific city, the adjustment unit can use the slang specific to that city. In this way, the adjustment unit can reflect regionally specific language usage by taking the user's geographical background into consideration. Some or all of the above processing in the adjustment unit may be performed using AI, for example, or not using AI. For example, the adjustment unit can use an AI model to reflect regionally specific language usage based on the user's geographical background data to perform adjustments.

[0114] The system according to the embodiment is not limited to the example described above, and various modifications are possible, for example, as follows.

[0115] The bot system can also include a health management unit that monitors the user's health status. This unit can, for example, monitor the user's heart rate and sleep patterns to understand their health condition. If the user is experiencing stress, the health management unit can suggest relaxing activities. Furthermore, it can manage the user's diet and exercise records to support healthy lifestyle habits. In this way, the health management unit can support the user's health maintenance by monitoring their health status and providing appropriate advice.

[0116] The bot system can also be equipped with a hobby learning unit that learns the user's hobbies and interests. For example, the hobby learning unit can learn about hobbies and interests the user has discussed in the past and provide relevant information based on that. If the user is a movie lover, for instance, the hobby learning unit can provide information on new movies. Similarly, if the user is a music lover, the hobby learning unit can provide information on the latest music trends and concerts. This allows the hobby learning unit to provide more personalized information based on the user's hobbies and interests.

[0117] The bot system can also include a music recommendation unit that estimates the user's emotions and recommends music based on those emotions. For example, if the user is sad, the music recommendation unit can recommend relaxing music. If the user is happy, the music recommendation unit can recommend upbeat music. Also, if the user is stressed, the music recommendation unit can recommend relaxing music. In this way, the music recommendation unit can improve the user's mood by recommending appropriate music based on the user's emotions.

[0118] The bot system can also include a learning management unit to manage the user's learning progress. For example, the learning management unit can record what the user is learning and manage their progress. If the user is learning English, for instance, the learning management unit can record their learning progress and suggest an appropriate learning plan. Furthermore, if the user is studying for an exam, the learning management unit can manage their study schedule leading up to the exam date. In this way, the learning management unit can manage the user's learning progress and support effective learning.

[0119] The bot system can also include a relaxation suggestion unit that estimates the user's emotions and proposes relaxation methods based on those emotions. For example, if the user is feeling stressed, the relaxation suggestion unit can suggest relaxation methods such as deep breathing or meditation. If the user is feeling tired, the relaxation suggestion unit can suggest relaxing music or aromatherapy. Furthermore, if the user is feeling anxious, the relaxation suggestion unit can suggest relaxing activities. In this way, the relaxation suggestion unit can reduce the user's stress by suggesting appropriate relaxation methods based on the user's emotions.

[0120] The bot system can also include a meal suggestion unit that estimates the user's emotions and suggests meals based on those emotions. For example, if the user is tired, the meal suggestion unit can suggest a nutritious meal. If the user is stressed, the meal suggestion unit can suggest a relaxing meal. Furthermore, if the user is happy, the meal suggestion unit can suggest a special meal. In this way, the meal suggestion unit can support the user's health by suggesting appropriate meals based on the user's emotions.

[0121] The bot system can also include an exercise suggestion unit that estimates the user's emotions and suggests exercises based on those emotions. For example, if the user is feeling stressed, the exercise suggestion unit can suggest relaxing yoga or stretching. If the user is feeling energetic, the exercise suggestion unit can suggest running or dancing. If the user is tired, the exercise suggestion unit can suggest light walking. In this way, the exercise suggestion unit can support the user's health by suggesting appropriate exercises based on the user's emotions.

[0122] The bot system can also include a travel planning unit to support the user's travel planning. For example, the travel planning unit can suggest travel plans based on the places the user wants to visit and activities they are interested in. If the user wants to enjoy nature, the travel planning unit can suggest tourist destinations rich in natural beauty. Similarly, if the user is interested in history and culture, the travel planning unit can suggest historical tourist destinations and cultural events. This allows the travel planning unit to propose the optimal travel plan based on the user's interests and preferences.

[0123] The bot system can also include a reading suggestion unit that estimates the user's emotions and suggests books based on those emotions. For example, if the user wants to relax, the reading suggestion unit can suggest a relaxing book. If the user is seeking excitement, the reading suggestion unit can suggest a thrilling book. Furthermore, if the user wants to learn, the reading suggestion unit can suggest an educational book. In this way, the reading suggestion unit can improve the user's reading experience by suggesting appropriate books based on the user's emotions.

[0124] The bot system can also be equipped with a task management unit to further improve the user's work efficiency. This unit can, for example, manage the user's schedule and set task priorities. It can also remind the user of other tasks so that they can focus on important projects. Furthermore, it can provide time management advice to help the user work efficiently. In this way, the task management unit can improve the user's work efficiency and reduce stress.

[0125] The following briefly describes the processing flow for example form 2.

[0126] Step 1: The learning unit learns the user's language use and thought patterns. For example, it learns the user's unique language use and thought patterns through dialogue with the user. The learning unit can analyze the user's language use and thought patterns using machine learning algorithms. It also identifies frequently used phrases and language use based on the user's dialogue history and uses that as a basis for learning. Step 2: The proxy unit acts as a proxy for the user based on the information learned by the learning unit. For example, it can generate responses by mimicking the user's speech patterns and thought processes. The proxy unit can participate in chats with friends and family on behalf of the user and provide fresh topics. It can also automatically answer specific questions. For example, it can generate appropriate responses based on the user's past answers. Step 3: The cloning unit creates a clone of the user. For example, it can learn the user's speaking style and way of thinking, and communicate using language that is characteristic of that person. It can also use an automated response function, allowing the clone to reply on the user's behalf. For example, it can participate in chats with friends and family on the user's behalf and provide fresh topics of conversation.

[0127] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0128] Data generation model 58 is a form of so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> Examples of generative AI include text generation AI, image generation AI, and multimodal generation AI. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats from audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVMs), k-means clustering, convolutional neural networks (CNNs), recurrent neural networks (RNNs), generative adversarial networks (GANs), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI ​​may be an AI agent. Furthermore, when the processing of each of the above parts is performed by the AI, the processing may be performed by the AI ​​in part or in whole, but is not limited to this example.Furthermore, processing performed by AI, including generative AI, may be replaced with rule-based processing, and rule-based processing may be replaced with processing performed by AI, including generative AI.

[0129] Furthermore, the processing performed by the data processing system 10 described above is carried out by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart device 14, but it may also be carried out by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart device 14. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the smart device 14 or an external device, and the smart device 14 acquires or collects information necessary for processing from the data processing device 12 or an external device.

[0130] Each of the multiple elements described above, including the learning unit, proxy unit, cloning unit, response unit, conversation continuation unit, and adjustment unit, is implemented in at least one of the smart device 14 and the data processing unit 12. For example, the learning unit is implemented by the processor 46 of the smart device 14 and the specific processing unit 290 of the data processing unit 12. The proxy unit is implemented by the control unit 46A of the smart device 14 and the specific processing unit 290 of the data processing unit 12. The cloning unit is implemented by the processor 46 of the smart device 14 and the specific processing unit 290 of the data processing unit 12. The response unit is implemented by the control unit 46A of the smart device 14 and the specific processing unit 290 of the data processing unit 12. The conversation continuation unit is implemented by the processor 46 of the smart device 14 and the specific processing unit 290 of the data processing unit 12. The adjustment unit is implemented by the control unit 46A of the smart device 14 and the specific processing unit 290 of the data processing unit 12. The correspondence between each part and the device or control unit is not limited to the examples described above, and various modifications are possible.

[0131] [Second Embodiment] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0132] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0133] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.

[0134] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0135] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0136] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).

[0137] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0138] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing by the processor 28. The storage 32 stores the specific processing program 56.

[0139] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0140] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.

[0141] In the smart glasses 214, specific processing is performed by the processor 46. The storage 50 stores a specific processing program 60. The processor 46 reads the specific processing program 60 from the storage 50 and executes the read specific processing program 60 on the RAM 48. The specific processing is realized by the processor 46 acting as a control unit 46A according to the specific processing program 60 executed on the RAM 48. The smart glasses 214 also have a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.

[0142] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).

[0143] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0144] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI ​​may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI ​​in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.

[0145] The data processing system 210 according to the second embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 210 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart glasses 214, but it may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart glasses 214. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the smart glasses 214 or an external device, and the smart glasses 214 acquires or collects information necessary for processing from the data processing device 12 or an external device.

[0146] Each of the multiple elements described above, including the learning unit, proxy unit, cloning unit, response unit, conversation continuation unit, and adjustment unit, is implemented, for example, in at least one of the smart glasses 214 and the data processing unit 12. For example, the learning unit is implemented by the processor 46 of the smart glasses 214 and the specific processing unit 290 of the data processing unit 12. The proxy unit is implemented, for example, by the control unit 46A of the smart glasses 214 and the specific processing unit 290 of the data processing unit 12. The cloning unit is implemented, for example, by the processor 46 of the smart glasses 214 and the specific processing unit 290 of the data processing unit 12. The response unit is implemented, for example, by the control unit 46A of the smart glasses 214 and the specific processing unit 290 of the data processing unit 12. The conversation continuation unit is implemented, for example, by the processor 46 of the smart glasses 214 and the specific processing unit 290 of the data processing unit 12. The adjustment unit is implemented, for example, by the control unit 46A of the smart glasses 214 and the specific processing unit 290 of the data processing unit 12. The correspondence between each part and the device or control unit is not limited to the examples described above, and various modifications are possible.

[0147] [Third Embodiment] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0148] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0149] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.

[0150] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0151] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0152] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).

[0153] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0154] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0155] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0156] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.

[0157] In the headset terminal 314, specific processing is performed by the processor 46. The storage 50 stores a specific program 60. The processor 46 reads the specific program 60 from the storage 50 and executes the read specific program 60 on the RAM 48. The specific processing is realized by the processor 46 acting as a control unit 46A according to the specific program 60 executed on the RAM 48. The headset terminal 314 also has a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.

[0158] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).

[0159] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0160] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI ​​may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI ​​in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.

[0161] The data processing system 310 according to the third embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 310 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the headset terminal 314, but may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the headset terminal 314. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the headset terminal 314 or an external device, and the headset terminal 314 acquires or collects information necessary for processing from the data processing device 12 or an external device.

[0162] Each of the multiple elements described above, including the learning unit, proxy unit, clone creation unit, response unit, conversation continuation unit, and adjustment unit, is implemented in at least one of the headset terminal 314 and the data processing unit 12. For example, the learning unit is implemented by the processor 46 of the headset terminal 314 and the specific processing unit 290 of the data processing unit 12. The proxy unit is implemented by the control unit 46A of the headset terminal 314 and the specific processing unit 290 of the data processing unit 12. The clone creation unit is implemented by the processor 46 of the headset terminal 314 and the specific processing unit 290 of the data processing unit 12. The response unit is implemented by the control unit 46A of the headset terminal 314 and the specific processing unit 290 of the data processing unit 12. The conversation continuation unit is implemented by the processor 46 of the headset terminal 314 and the specific processing unit 290 of the data processing unit 12. The adjustment unit is implemented by the control unit 46A of the headset terminal 314 and the specific processing unit 290 of the data processing unit 12. The correspondence between each part and the device or control unit is not limited to the examples described above, and various modifications are possible.

[0163] [Fourth Embodiment] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0164] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0165] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.

[0166] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0167] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0168] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS image sensor or CCD image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).

[0169] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0170] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. The robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0171] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0172] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0173] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.

[0174] In robot 414, specific processing is performed by processor 46. A specific program 60 is stored in storage 50. Processor 46 reads the specific program 60 from storage 50 and executes it on RAM 48. The specific processing is achieved by processor 46 acting as a control unit 46A according to the specific program 60 executed on RAM 48. Robot 414 also has data generation model 58 and emotion identification model 59, similar to those of the robot, and can perform processing similar to that of the specific processing unit 290 using these models.

[0175] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).

[0176] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0177] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI ​​may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI ​​in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.

[0178] The data processing system 410 according to the fourth embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 410 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the robot 414, but it may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the robot 414. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the robot 414 or an external device, and the robot 414 acquires or collects information necessary for processing from the data processing device 12 or an external device.

[0179] Each of the multiple elements described above, including the learning unit, proxy unit, cloning unit, response unit, conversation continuation unit, and adjustment unit, is implemented, for example, in at least one of the robot 414 and the data processing unit 12. For example, the learning unit is implemented by the processor 46 of the robot 414 and the specific processing unit 290 of the data processing unit 12. The proxy unit is implemented, for example, by the control unit 46A of the robot 414 and the specific processing unit 290 of the data processing unit 12. The cloning unit is implemented, for example, by the processor 46 of the robot 414 and the specific processing unit 290 of the data processing unit 12. The response unit is implemented, for example, by the control unit 46A of the robot 414 and the specific processing unit 290 of the data processing unit 12. The conversation continuation unit is implemented, for example, by the processor 46 of the robot 414 and the specific processing unit 290 of the data processing unit 12. The adjustment unit is implemented, for example, by the control unit 46A of the robot 414 and the specific processing unit 290 of the data processing unit 12. The correspondence between each part and the device or control unit is not limited to the examples described above, and various modifications are possible.

[0180] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0181] Figure 9 shows the emotion map 400, in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0182] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0183] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0184] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, and motorcycles, emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated based, for example, on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0185] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0186] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0187] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing method for the specific process may be used, which includes computer 22 and multiple other computers.

[0188] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0189] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0190] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0191] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0192] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0193] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0194] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0195] Furthermore, although the above-described examples were divided into four embodiments, some or all of these embodiments may be combined. Also, the smart device 14, smart glasses 214, headset terminal 314, and robot 414 are just examples, and they may be combined, or other devices may be used. Also, although the above-described examples were divided into two embodiments, Embodiment 1 and Embodiment 2, these may be combined.

[0196] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and other things that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0197] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.

[0198] (Note 1) A learning unit that learns the user's language use and thought patterns, An agency unit that functions as a proxy for the user based on the information learned by the learning unit, It includes a cloning unit that creates a clone of the user. A system characterized by the following features. (Note 2) It includes a response unit that provides an automatic response function. The system described in Appendix 1, characterized by the features described herein. (Note 3) It is equipped with a conversation continuation unit for continuous conversation. The system described in Appendix 1, characterized by the features described herein. (Note 4) It features an adjustment unit for prompt engineering and fine-tuning. The system described in Appendix 1, characterized by the features described herein. (Note 5) The aforementioned cloning unit is It learns the user's way of speaking and thinking, and communicates using language that is uniquely theirs. The system described in Appendix 1, characterized by the features described herein. (Note 6) The response unit is A clone replies on behalf of the user. The system described in Appendix 2, characterized by the features described herein. (Note 7) The aforementioned conversation continuation section is, AI engages in continuous conversation as if it possessed memory. The system described in Appendix 3, characterized by the features described herein. (Note 8) The adjustment unit is, Prompt engineering and fine-tuning are performed based on sample statements. The system described in Appendix 4, characterized by the features described herein. (Note 9) The aforementioned learning unit, It estimates the user's emotions and adjusts the learning priorities based on the estimated user emotions. The system described in Appendix 1, characterized by the features described herein. (Note 10) The aforementioned learning unit, Analyze the user's past conversation history and learn their responses to specific topics. The system described in Appendix 1, characterized by the features described herein. (Note 11) The aforementioned learning unit, During the learning process, the optimal learning timing is selected by considering the user's lifestyle patterns. The system described in Appendix 1, characterized by the features described herein. (Note 12) The aforementioned learning unit, It estimates the user's emotions and customizes the learning content based on the estimated user emotions. The system described in Appendix 1, characterized by the features described herein. (Note 13) The aforementioned learning unit, During training, the system learns regionally specific language usage by taking into account the user's geographical background. The system described in Appendix 1, characterized by the features described herein. (Note 14) The aforementioned learning unit, During learning, the system analyzes users' social media activity and learns relevant information. The system described in Appendix 1, characterized by the features described herein. (Note 15) The aforementioned agency unit is It estimates the user's emotions and adjusts the surrogate expression based on the estimated user emotions. The system described in Appendix 1, characterized by the features described herein. (Note 16) The aforementioned agency unit is It references the user's past statements and generates the most appropriate response. The system described in Appendix 1, characterized by the features described herein. (Note 17) The aforementioned agency unit is Respond at the appropriate time, taking into account the user's current situation. The system described in Appendix 1, characterized by the features described herein. (Note 18) The aforementioned agency unit is It estimates the user's emotions and customizes the surrogate response based on the estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 19) The aforementioned agency unit is Use region-specific language, taking into account the user's geographical background. The system described in Appendix 1, characterized by the features described herein. (Note 20) The aforementioned agency unit is Analyze users' social media activity and reflect relevant information in responses. The system described in Appendix 1, characterized by the features described herein. (Note 21) The aforementioned cloning unit is It estimates the user's emotions and adjusts the cloning process based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 22) The aforementioned cloning unit is When creating a clone, the system references the user's past interaction history to generate the optimal clone. The system described in Appendix 1, characterized by the features described herein. (Note 23) The aforementioned cloning unit is When creating a clone, the clone's behavior is customized by taking into account the user's lifestyle patterns. The system described in Appendix 1, characterized by the features described herein. (Note 24) The aforementioned cloning unit is It estimates the user's emotions and customizes the clone's responses based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 25) The aforementioned cloning unit is When creating a clone, the system takes the user's geographical background into account and reflects regionally specific language usage. The system described in Appendix 1, characterized by the features described herein. (Note 26) The aforementioned cloning unit is During cloning, the system analyzes the user's social media activity and incorporates relevant information into the clone. The system described in Appendix 1, characterized by the features described herein. (Note 27) The response unit is It estimates the user's emotions and adjusts the way responses are expressed based on those estimated emotions. The system described in Appendix 2, characterized by the features described herein. (Note 28) The response unit is It references the user's past statements and generates the most appropriate response. The system described in Appendix 2, characterized by the features described herein. (Note 29) The response unit is Respond at the appropriate time, taking into account the user's current situation. The system described in Appendix 2, characterized by the features described herein. (Note 30) The response unit is It estimates the user's emotions and customizes the response based on those emotions. The system described in Appendix 2, characterized by the features described herein. (Note 31) The response unit is Use region-specific language, taking into account the user's geographical background. The system described in Appendix 2, characterized by the features described herein. (Note 32) The aforementioned conversation continuation section is, It estimates the user's emotions and adjusts how the conversation continues based on those estimated emotions. The system described in Appendix 3, characterized by the features described herein. (Note 33) The aforementioned conversation continuation section is, It references the user's past conversation history to generate the optimal flow of conversation. The system described in Appendix 3, characterized by the features described herein. (Note 34) The aforementioned conversation continuation section is, Considering the user's current situation, continue the conversation at the appropriate time. The system described in Appendix 3, characterized by the features described herein. (Note 35) The aforementioned conversation continuation section is, It estimates the user's emotions and customizes the conversation content based on those estimated emotions. The system described in Appendix 3, characterized by the features described herein. (Note 36) The aforementioned conversation continuation section is, Take into account the users' geographical backgrounds and incorporate topics specific to their region. The system described in Appendix 3, characterized by the features described herein. (Note 37) The adjustment unit is, It estimates the user's emotions and adjusts prompt engineering and fine-tuning methods based on the estimated user emotions. The system described in Appendix 4, characterized by the features described herein. (Note 38) The adjustment unit is, Refer to the user's past statements and make optimal adjustments. The system described in Appendix 4, characterized by the features described herein. (Note 39) The adjustment unit is, Adjustments will be made at the appropriate time, taking into account the user's current situation. The system described in Appendix 4, characterized by the features described herein. (Note 40) The adjustment unit is, It estimates the user's emotions and customizes the adjustments based on those estimated emotions. The system described in Appendix 4, characterized by the features described herein. (Note 41) The adjustment unit is, Take into account the user's geographical background and reflect regionally specific language usage. The system described in Appendix 4, characterized by the features described herein. [Explanation of Symbols]

[0199] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots

Claims

1. A learning unit that learns the user's language use and thought patterns, An agency unit that functions as a proxy for the user based on the information learned by the learning unit, It includes a cloning unit that creates a clone of the user. A system characterized by the following features.

2. It includes a response unit that provides an automatic response function. The system according to feature 1.

3. It is equipped with a conversation continuation unit for continuous conversation. The system according to feature 1.

4. It features an adjustment unit for prompt engineering and fine-tuning. The system according to feature 1.

5. The aforementioned cloning unit is It learns the user's way of speaking and thinking, and communicates using language that is uniquely theirs. The system according to feature 1.

6. The response unit is A clone replies on behalf of the user. The system according to feature 2.

7. The aforementioned conversation continuation section is, AI engages in continuous conversation as if it possessed memory. The system according to claim 3.

8. The adjustment unit is, Prompt engineering and fine-tuning are performed based on sample statements. The system according to feature 4.

9. The aforementioned learning unit, It estimates the user's emotions and adjusts the learning priorities based on the estimated user emotions. The system according to feature 1.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A