system
The system effectively reconstructs users' memories and personalities using life log data from messenger apps and internet search tools, generating avatars for continuous communication by training AI models and employing a subscription model for enhanced accuracy.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2024-09-19
- Publication Date
- 2026-04-14
AI Technical Summary
Existing systems fail to efficiently reconstruct users' memories and personalities using life log data and characteristic data, and lack means to maintain communication with acquaintances, friends, and family when the user is absent or after death, without a sustainable billing model to improve accuracy.
A system that collects life log data and characteristic data from messenger apps and internet search tools, preprocesses this data, trains AI models to recreate users' memories and personalities, generates avatars in the cloud, and maintains communication through a subscription model, enhancing accuracy over time.
Enables efficient reconstruction of users' memories and personalities, allowing continuous communication with loved ones even after death, with improved accuracy through a subscription-based learning mechanism.
Smart Images

Figure 0007846180000001 
Figure 0007846180000002 
Figure 0007846180000003
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, the method including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] The death of a human is an inevitable reality, and as a result, people will lose communication with their loved ones forever. Also, even during their lifetime, due to busyness or absence, communication with important people may be restricted.
Means for Solving the Problems
[0005] This invention links life log data and characteristic data such as text and voiceprints from messenger apps, internet search tools, etc., and uses this as training data for AI to recreate the user's memories and personality. It then generates an avatar of the person in the cloud, which responds to casual conversations and consultations with acquaintances, friends, and family members when the person is absent or busy during their lifetime, and as an eternal presence after death. Furthermore, it provides a system with a monthly subscription fee, where the accuracy of the recreation increases the longer the subscription is continued. This makes it possible for people to maintain communication with important people in their lives, both during and after death. [Brief explanation of the drawing]
[0006] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Embodiment 1 of Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1 of Form Example 1. [Figure 13] This is a sequence diagram showing the processing flow of the data processing system in Example 2 of Embodiment 2. [Figure 14] This is a sequence diagram showing the processing flow of the data processing system in Application Example 2 of Form Example 2. [Figure 15] This is a sequence diagram showing the processing flow of the data processing system in Embodiment 3 of Example 3. [Figure 16] This is a sequence diagram showing the processing flow of the data processing system in Application Example 3 of Form Example 3. [Figure 17] This is a sequence diagram showing the processing flow of the data processing system in Example 1 of the Form 1 when an emotion engine is combined. [Figure 18] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1 of Form Example 1 when an emotion engine is combined. [Figure 19] This is a sequence diagram showing the processing flow of the data processing system in Example 2 of the Form 2 when an emotion engine is combined. [Figure 20] This is a sequence diagram showing the processing flow of the data processing system in Application Example 2 of Form Example 2 when an emotion engine is combined. [Figure 21] This is a sequence diagram showing the processing flow of the data processing system in Example 3 of the Form 3 when an emotion engine is combined. [Figure 22] This is a sequence diagram showing the processing flow of the data processing system in Application Example 3 of Form Example 3 when an emotion engine is combined. [Modes for carrying out the invention]
[0007] Hereinafter, an example of an embodiment of the system relating to the technology of this disclosure will be described with reference to the attached drawings.
[0008] ru.
[0009] First, the terms used in the following description will be explained.
[0010] In the following embodiments, the labeled processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. As an example of an arithmetic unit, there are a CPU (Central Processing Unit),
[0011] GPU (Graphics Processing Unit), GPGPU (General-Purpose computing on Graphics Processing Units), APU (Accelerated Processing Unit), or TPU (TENSOR PROCESSING UNIT (registered trademark)), etc.
[0012] In the following embodiments, the labeled RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor. [[ID=__16]]
[0013] In the following embodiments, the labeled storage is one or more non-volatile storage devices that store various programs and various parameters, etc. As an example of a non-volatile storage device, there are a flash memory (SSD (Solid State Drive)), a magnetic disk (e.g.,
[0014] hard disk), or magnetic tape, etc.
[0015] In the following embodiments, the labeled communication I / F (Interface) is a communication processor and
[0016] A communication interface (I / F) is an interface that includes antennas and other components. The I / F manages communication between multiple computers. Examples of communication standards applicable to an I / F include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0017] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0018] [First Embodiment]
[0019] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0020] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0021] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0022] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0023] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0024] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0025] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0026] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0027] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0028] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0029] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0030] Next, the identification process performed by the identification processing unit 290 of the data processing device 12 will be described.
[0031] "Example of form 1"
[0032] One embodiment of the present invention involves collecting life log data and characteristic data such as text and voiceprints from messenger apps and internet search tools that users use on a daily basis.
[0033] This data reflects the user's behavioral patterns, hobbies, interests, and personality. By using this data as training data to train the AI, it can recreate the user's memories and personality.
[0034] "Example of form 2"
[0035] Next, a virtual avatar of the user is created in the cloud. This avatar is there to maintain communication with acquaintances, friends, and family even when the user is absent, busy, or after the user's death. Specifically, the avatar mimics the user's speech patterns and thought processes, recreating how the user normally interacts with others.
[0036] "Example of form 3"
[0037] Furthermore, the system of this invention employs a monthly subscription model. The longer a user continues to subscribe, the more data the system collects, and the more the AI learns. This increases the accuracy of the avatar's reproduction, making it possible to more accurately reproduce the user's memories and personality.
[0038] The following describes the processing flow for each example of the form.
[0039] "Example of form 1"
[0040] Step 1: Collect life log data and characteristic data such as text and voiceprints from messenger apps and internet search tools that the user uses on a daily basis.
[0041] Step 2: The collected data is used as training data to train the AI. Through this training, the AI understands the user's behavior patterns, hobbies, interests, personality, etc.
[0042] Step 3: As the AI learns, its ability to recreate the user's memories and personality improves.
[0043] "Example of form 2"
[0044] Step 1: Create a virtual avatar of the user in the cloud. This avatar will maintain communication with acquaintances, friends, and family even when the user is absent, busy, or has passed away.
[0045] Step 2: The avatar mimics the user's speech patterns and thought processes, recreating how the user typically interacts with others.
[0046] "Example of form 3"
[0047] Step 1: The system of the present invention employs a monthly subscription model.
[0048] Step 2: The longer a user continues to pay, the more data the system collects and the more the AI learns.
[0049] Step 3: This increases the accuracy of the avatar's reproduction, making it possible to more accurately recreate the user's memories and personality.
[0050] (Example 1)
[0051] Next, we will describe Example 1 of Form Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0052] In modern society, there is a need for technology that collects data reflecting users' behavioral patterns, interests, and personalities, and uses that data to reconstruct their memories and personalities. However, existing systems do not integrate the steps of data collection, preprocessing, storage, learning, evaluation, and generation, making it difficult to efficiently reconstruct users' memories and personalities. Furthermore, there is a lack of means to maintain communication with acquaintances, friends, and family even when the user is absent or after death. In addition, a sustainable billing model to improve the accuracy of the reconstruction has not been established.
[0053] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0054] In this invention, the server includes means for collecting life log data and characteristic data such as text and voiceprints from messenger apps, internet search tools, etc.; means for preprocessing the collected data and converting it into a format suitable for training an AI model; means for storing the preprocessed data in a database; means for inputting the stored data as training data into an AI model and performing training to reproduce the user's memories and personality; means for evaluating the AI model after training is complete and confirming its performance; means for saving models with good evaluations; and means for generating an avatar of the user on the cloud. This makes it possible to efficiently collect and process data that reflects the user's behavior patterns, interests, and personality, and to reproduce the user's memories and personality with high accuracy. Furthermore, it is possible to maintain communication with acquaintances, friends, and family even when the user is absent or after death, and the accuracy of reproduction can be improved through a continuous billing model.
[0055] A "messenger app" is a software application that allows users to send and receive text messages, voice messages, images, videos, and other content.
[0056] An "internet search tool" is a software application that allows users to search for information on the internet.
[0057] "Life log data" refers to data that records a user's actions and activities in their daily life.
[0058] "Feature data" refers to data that reflects an individual's characteristics, such as a user's written text or voiceprint.
[0059] "Preprocessing" refers to the process of converting collected data into a format suitable for training an AI model.
[0060] A "database" is a system for efficiently storing, managing, and retrieving data.
[0061] "Training data" refers to data with correct labels that is used to train an AI model.
[0062] An "AI model" is a mathematical model that uses artificial intelligence technology to analyze data and perform predictions and classifications.
[0063] "Learning" is the process by which an AI model discovers patterns and rules based on training data.
[0064] "Evaluation" is the process of checking the performance of an AI model after it has completed its training.
[0065] "Cloud" refers to a collection of computer resources and services provided via the internet.
[0066] An "incarnation" is a digital entity that recreates the user's memories and personality.
[0067] A "billing model" is a business model that collects fees for the use of a service.
[0068] Modes for carrying out the invention
[0069] This invention is a system that collects life log data and characteristic data such as text and voiceprints from messenger apps and internet search tools that users use on a daily basis, and uses this data to reconstruct the user's memories and personality. A specific embodiment of this system is described below.
[0070] Data collection
[0071] The server collects life log data and characteristic data such as text and voiceprints from messenger apps and internet search tools that users use on a daily basis. This collection uses common software such as messenger apps and internet search tools. The collected data reflects the user's behavioral patterns, hobbies, interests, personality, etc.
[0072] Data preprocessing
[0073] The server preprocesses the collected data and converts it into a format suitable for training the AI model. Specifically, this involves tokenizing text data and converting audio data into spectrograms. This is done using Python libraries such as NLTK and Librosa. For example, NLTK is used for tokenizing text data, and Librosa is used for converting audio data into spectrograms.
[0074] Data storage
[0075] The server stores the pre-processed data in a database. This uses database systems such as MySQL® or MongoDB. Specifically, it establishes a database connection and inserts the pre-processed data into the appropriate tables and collections.
[0076] AI model training
[0077] The server inputs pre-processed data as training data into an AI model and performs training to reproduce the user's memories and personality. For this training, generative AI models such as GPT-4® or BERT are used. Specifically, the server performs model initialization, data batch processing, and training.
[0078] Model evaluation and saving
[0079] The server evaluates the trained AI models and verifies their performance. Evaluation metrics such as accuracy and recall are used. Models that perform well are saved to the file system or cloud storage.
[0080] Cloud-based avatar generation
[0081] The server uses a well-rated model to generate an avatar of the user in the cloud. This avatar is used to recreate the user's memories and personality, and to maintain communication with acquaintances, friends, and family.
[0082] Specific example
[0083] Example 1: Data collection from messenger apps
[0084] The server collects messages that users send and receive using common messenger apps. The collected messages reflect the user's conversation patterns and interests. By preprocessing this message data and training an AI model, the system can replicate the user's conversational style.
[0085] Specific example 2: Data collection from internet search tools
[0086] The server collects search queries that users make using common internet search tools. These collected search queries reflect the user's interests and preferences. By preprocessing this search data and training an AI model, the system can reproduce the user's interests.
[0087] Example of a prompt
[0088] "Create a program that collects messages sent by users through common messenger apps and trains an AI model to replicate the user's conversational style."
[0089] "Create a program that collects search queries that users make using common internet search tools and trains an AI model to replicate those users' interests."
[0090] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0091] Step 1: Data Collection
[0092] The server collects life log data and characteristic data such as text and voiceprints from messenger apps and internet search tools that users use on a daily basis. It receives log data from messenger apps and internet search tools as input and generates the collected raw data as output. Specifically, the server uses APIs and scraping techniques to acquire data and collects it using a secure communication protocol (e.g., HTTPS).
[0093] Step 2: Data Preprocessing
[0094] The server preprocesses the collected data and converts it into a format suitable for training the AI model. It receives raw data as input and generates preprocessed data as output. Specifically, the server performs the following operations:
[0095] Text data tokenization: Use Python's NLTK library to split text into words and phrases.
[0096] Spectrogram conversion of audio data: Convert audio data to a frequency spectrum using the Librosa library in Python.
[0097] Step 3: Save data
[0098] The server saves the pre-processed data to the database. It receives pre-processed data as input and generates data stored in the database as output. Specifically, the server performs the following operations:
[0099] Establishing a database connection: Connect to MySQL or MongoDB.
[0100] Data insertion: Insert pre-processed data into the appropriate tables or collections.
[0101] Step 4: Training the AI model
[0102] The server inputs pre-processed data as training data into the AI model and performs training to reproduce the user's memories and personality. It receives pre-processed data stored in a database as input and generates a trained AI model as output. Specifically, the server performs the following operations:
[0103] Model initialization: Initialize generative AI models such as GPT-4 and BERT.
[0104] Data batch processing: Pre-processed data is divided into batches and input into the model.
[0105] Running the training: Input data into the model and run the training.
[0106] Step 5: Evaluate and save the model.
[0107] The server evaluates the trained AI model and verifies its performance. It receives the trained AI model and evaluation data as input, and generates the evaluation results and a saved model as output. Specifically, the server performs the following operations:
[0108] Model evaluation: Calculate metrics such as accuracy and recall.
[0109] Model saving: Save models with good performance to the file system or cloud storage.
[0110] Step 6: Creating an avatar in the cloud
[0111] The server generates an avatar of the user in the cloud using a well-rated model. It receives a stored AI model as input and generates an avatar in the cloud as output. Specifically, the server performs the following operations:
[0112] Model Deployment: Deploy the AI model to the cloud environment.
[0113] Avatar Generation: Use the deployed model to generate an avatar that replicates the user's memories and personality.
[0114] (Application Example 1)
[0115] Next, we will describe Application Example 1 of Form Example 1. In the following description, the data processing device 12 will be referred to as a "server," and the smart device 14 will be referred to as a "terminal."
[0116] In modern society, there is a growing need to provide personalized content to users by leveraging their life log data and characteristic data. However, conventional systems have struggled to recreate users' memories and personalities and recommend personalized content. Furthermore, there has been a lack of means to maintain communication with acquaintances and family when users are absent, busy, or even after death. A new system is needed to solve these problems.
[0117] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0118] This invention includes a server that provides means for linking life log data and characteristic data such as text and voiceprints from messenger apps, internet search tools, etc., means for an AI to recreate the user's memories and personality using these as training data, means for generating an avatar of the user on the cloud, and means for automatically recommending content based on the user's preferences and interests based on the user's life log data and characteristic data. This makes it possible to recreate the user's memories and personality and provide personalized content. Furthermore, it enables the user to maintain communication with acquaintances and family even when they are absent, busy, or even after death.
[0119] A "messenger app" is an application that allows users to send and receive text messages, voice messages, images, videos, and other content.
[0120] An "internet search tool" is software or a web service that allows users to search for information on the internet.
[0121] "Life log data" refers to records of a user's actions and activities in their daily life.
[0122] "Feature data" refers to data used to identify individual users, such as user text or voiceprints.
[0123] "Training data" refers to a dataset used for AI to learn.
[0124] "AI" stands for artificial intelligence, which is a technology in which machines imitate human intelligence.
[0125] "Recreating the user's memories and personality" means that the AI imitates the user's thought and behavioral patterns based on the user's past actions and statements.
[0126] "Creating a digital avatar of oneself on the cloud" means creating a digital alter ego of a user using cloud computing technology.
[0127] "Automatically recommending content" means that AI suggests appropriate information and entertainment based on the user's preferences and interests.
[0128] The system for implementing this invention collects user life log data and characteristic data, reconstructs the user's memories and personality based on this data, and recommends personalized content. Specific embodiments of this system are described below.
[0129] System Configuration
[0130] hardware
[0131] Server: A server used for data collection, processing, storage, and training and inference of AI models.
[0132] Device: A device used by a user, such as a smartphone or computer.
[0133] software
[0134] Messenger app: An application that allows users to send and receive text messages, voice messages, images, videos, and other content.
[0135] Internet search tool: Software or web service used by users to search for information on the internet.
[0136] AI Model: A generative AI model that recreates the user's memories and personality based on the user's life log data and characteristic data, and recommends personalized content.
[0137] Data collection and processing
[0138] The server collects user lifelog and characteristic data from messenger apps and internet search tools. This includes the user's message history, search history, and voice data. The collected data is preprocessed and then used as training data for AI models.
[0139] AI model training and inference
[0140] The server trains an AI model based on the collected data. Specifically, text data is digitized using TfidfVectorizer, and user profiles are created using clustering algorithms (e.g., KMeans). Audio data is converted to text using speech recognition technology and processed similarly.
[0141] Content Recommendation
[0142] The server uses a pre-trained AI model to automatically recommend content based on the user's preferences and interests. This includes entertainment content such as movies, music, and articles. Recommended content is notified to the user's device.
[0143] Specific example
[0144] For example, if user A frequently talks about "movies" on a messenger app, the server collects that data and uses it to train an AI model. If user A's search history includes many searches for "latest movie reviews" and "movie trailers," the server creates a profile of user A based on this data. As a result, user A is recommended the latest movie review articles and movie trailer videos.
[0145] Example of a prompt
[0146] Develop an application that recommends content tailored to users' interests and preferences based on data collected from their messenger apps and search history. Specifically, implement a function that analyzes themes users frequently discuss and keywords they search for, and then automatically recommends content such as movies, music, and articles based on that analysis.
[0147] In this way, it is possible to build a system that utilizes users' life log data and characteristic data to deliver personalized content.
[0148] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0149] Step 1:
[0150] The server collects user lifelog and characteristic data from messenger apps and internet search tools. Specifically, it retrieves user message history, search history, voice data, etc. This data is collected using APIs and database queries.
[0151] Input: Raw data from messenger apps and search tools
[0152] Output: Collected lifelog data and feature data
[0153] Step 2:
[0154] The server preprocesses the collected data. Specifically, text data is digitized using TfidfVectorizer, and speech data is converted to text using speech recognition technology. This converts the data into a format suitable for the AI model.
[0155] Input: Collected life log data and feature data
[0156] Output: Preprocessed data
[0157] Step 3:
[0158] The server trains the AI model based on pre-processed data. Specifically, it creates user profiles using clustering algorithms (e.g., KMeans). This allows user behavior patterns and interests to be reflected in the model.
[0159] Input: Preprocessed data
[0160] Output: Trained AI model
[0161] Step 4:
[0162] The server uses a pre-trained AI model to automatically recommend content based on the user's preferences and interests. Specifically, it selects entertainment content such as movies, music, and articles based on the user's profile.
[0163] Input: Pre-trained AI model, user profile
[0164] Output: Recommended content
[0165] Step 5:
[0166] The server notifies the user's device of the recommended content. Specifically, it presents the content to the user using push notifications or in-app messages.
[0167] Input: Recommended content
[0168] Output: Content displayed on the user's device
[0169] In this way, a system is built that utilizes users' life log data and characteristic data to deliver personalized content.
[0170] (Example 2)
[0171] Next, we will describe Example 2 of Form Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0172] In modern society, there is a need to maintain communication with acquaintances, friends, and family even when users are absent, busy, or have passed away. However, conventional technologies have struggled to accurately mimic users' speech patterns and thought processes to achieve natural dialogue. Furthermore, there has been a lack of effective means to utilize users' past messages and dialogue history to train generative AI models. This has resulted in the problem of user avatars generating unnatural responses.
[0173] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0174] In this invention, the server includes means for linking life log data and characteristic data such as text and voiceprints from messenger apps, internet search tools, etc., means for a generative AI model to reproduce the user's memories and personality using these as training data, means for generating an avatar of the user on the cloud, means for collecting the user's past messages and dialogue history and storing it in a database, means for preprocessing the collected data to clean the text, tokenize it, and extract important phrases and patterns, means for training the generative AI model using the preprocessed data, means for generating an avatar of the user using the trained generative AI model, means for receiving messages from acquaintances and family, inputting them as prompts into the generative AI model, generating appropriate replies, and means for sending the generated replies to acquaintances and family. This makes it possible to accurately imitate the user's speech patterns and thought patterns and realize natural dialogue.
[0175] A "messenger app" is software that allows users to send and receive text messages, voice messages, images, videos, and other content.
[0176] An "internet search tool" is software that allows users to search for and retrieve information from the internet.
[0177] "Life log data" refers to data related to a user's daily life, including activity history, location information, and health data.
[0178] "Feature data" refers to data that represents the individual characteristics of a user, and includes things like writing style and voiceprints.
[0179] A "generative AI model" is a model that uses artificial intelligence technology to imitate a user's speech patterns and thought processes.
[0180] "Cloud" refers to a collection of computer resources and services provided via the internet.
[0181] An "avatar" is a virtual entity that mimics the user's speech patterns and thought processes, and engages in dialogue on their behalf.
[0182] A "database" is a system for efficiently storing, searching, and managing data.
[0183] "Preprocessing" refers to the process of converting data into a format suitable for analysis and model training.
[0184] "Tokenization" is the process of dividing text data into smaller units such as words and phrases.
[0185] A "prompt" is text input to a generative AI model that contains instructions for the model to generate an appropriate response.
[0186] This invention is a system for maintaining communication with acquaintances, friends, and family even when the user is absent, busy, or has passed away. This system links life log data and characteristic data such as text and voiceprints from messenger apps, internet search tools, etc., and uses this as training data for a generated AI model to recreate the user's memories and personality. A specific embodiment of this system is described below.
[0187] 1. Collection of user data
[0188] The server collects the user's past messages and conversation history. This includes emails, chat logs, and social media posts. The server stores this data in a database. For example, the server accesses email accounts and downloads past emails. It also uses chat application APIs to retrieve the user's chat logs.
[0189] 2. Data preprocessing
[0190] The server preprocesses the collected data. This includes text cleaning, tokenization, and extraction of important phrases and patterns. Specifically, the server removes unnecessary HTML tags and special characters from the text and divides the text into words and phrases (tokenization). Furthermore, it extracts frequently occurring phrases and patterns and stores them in a database.
[0191] 3. Training the Generative AI Model
[0192] The server trains a generative AI model using pre-processed data. For example, OpenAI's GPT-4 is used here. The server inputs the pre-processed data into the generative AI model and trains it. Once trained, the model becomes capable of mimicking the user's speech patterns and thought processes.
[0193] 4. The creation of an incarnation
[0194] The server uses a trained generative AI model to generate an avatar of the user. This avatar mimics the user's speech patterns and thought processes, recreating how the user typically interacts. The generated avatar is stored in a database.
[0195] 5. Receiving messages and generating replies
[0196] The device receives messages from acquaintances and family. The device inputs these messages as prompts into a generation AI model, which then generates an appropriate reply. For example, if the device receives the message "What are your plans for tomorrow?" from an acquaintance, it inputs this message into the generation AI model as a prompt. An example of a prompt would be: "User name: Taro Yamada, User's previous message: 'I'm busy today, so please check my plans for tomorrow,' Prompt: 'Generate a reply to confirm Taro Yamada's plans for tomorrow when he is busy.'"
[0197] 6. Send a reply
[0198] The device sends the generated reply to acquaintances and family. For example, it receives a reply from the generating AI model, "Tomorrow's schedule is a meeting at 10am and a client meeting at 2pm," and sends this to an acquaintance.
[0199] In this way, the server and terminal work together to create a user avatar, making it possible to maintain communication with acquaintances, friends, and family even when the user is absent, busy, or has passed away.
[0200] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0201] Step 1: Collecting User Data
[0202] The server collects the user's past messages and conversation history. Inputs include emails, chat logs, and social media posts. The server stores this data in a database. Specifically, the server accesses email accounts and downloads past emails. It also uses chat application APIs to retrieve the user's chat logs. The output is the collected data stored in the database.
[0203] Step 2: Data preprocessing
[0204] The server preprocesses the collected data. The input includes the data collected in step 1. The server cleans, tokenizes, and extracts important phrases and patterns from the text. Specifically, the server removes unnecessary HTML tags and special characters from the text and divides the text into words and phrases (tokenization). Furthermore, it extracts frequently occurring phrases and patterns and stores them in the database. The output is the preprocessed data.
[0205] Step 3: Training the Generative AI Model
[0206] The server trains a generative AI model using preprocessed data. The input includes the data preprocessed in step 2. The server inputs the preprocessed data into the generative AI model and trains the model. Specifically, the server inputs data into the generative AI model (e.g., GPT-4) and adjusts the model's parameters. The output is a fully trained generative AI model.
[0207] Step 4: Creation of the Incarnation
[0208] The server generates an avatar of the user using a trained generative AI model. The input includes the generative AI model trained in step 3. The server inputs the user's characteristics into the generative AI model and generates an avatar. Specifically, the server inputs the user's speech patterns and thought patterns into the model and generates an avatar. As output, the generated avatar is saved in the database.
[0209] Step 5: Receiving messages and generating replies
[0210] The device receives messages from acquaintances and family. The input includes messages from acquaintances and family. The device inputs these messages as prompts into a generating AI model, which then generates an appropriate reply. Specifically, the device uses the messaging app's API to retrieve the received message and inputs it as a prompt into the generating AI model. The output is the reply from the generating AI model.
[0211] Step 6: Send a reply
[0212] The device sends the generated reply to acquaintances and family. The input includes the reply generated in step 5. Specifically, the device uses the messaging app's API to send the generated reply. The output is the reply sent to acquaintances and family.
[0213] In this way, the server and terminal work together to create a user avatar, making it possible to maintain communication with acquaintances, friends, and family even when the user is absent, busy, or has passed away.
[0214] (Application Example 2)
[0215] Next, we will describe application example 2 of form example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".
[0216] In modern society, it is often difficult to provide customer support when users are absent or busy. Furthermore, after a user's death, there are limited means of maintaining communication with acquaintances, friends, and family. Additionally, while virtual stores require the ability to mimic users' speech patterns and thought processes, the technology to achieve this is lacking.
[0217] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for linking life log data and characteristic data such as text and voiceprints from messenger apps, internet search tools, etc., means for the AI to reproduce the user's memories and personality using these as training data, means for generating an avatar of the user on the cloud, means for the avatar of the user to handle customer interactions in a virtual store, means for generating answers to customer questions using a generated AI model, means for mimicking the user's speech patterns and thought patterns using prompt sentences, and means for installation on smartphones and head-mounted displays. This makes it possible to handle customer interactions even when the user is absent or busy, and to maintain communication with acquaintances, friends, and family even after the user's death. Furthermore, it enables customer interactions in a virtual store that mimic the user's speech patterns and thought patterns.
[0218] A "messenger app" is software that allows users to send and receive text messages, voice messages, images, videos, and other content.
[0219] An "internet search tool" is software used to search for information on the internet and provide it to users.
[0220] "Life log data" refers to data related to a user's daily life, including activity history, location information, and health data.
[0221] "Feature data" refers to data used to identify an individual, such as a user's written text or voiceprint.
[0222] "Training data" refers to data used to train machine learning models.
[0223] "AI" is an abbreviation for artificial intelligence, which is a technology in which computers imitate human intelligence to learn and reason.
[0224] "Cloud" refers to a collection of computer resources and services provided via the internet.
[0225] An "avatar" is a virtual entity created by mimicking the user's speech patterns and thought processes.
[0226] A "virtual store" is a virtual store that exists on the internet, where users can purchase goods and services.
[0227] "Customer service" refers to responding to and addressing customer questions and requests.
[0228] A "generative AI model" is an artificial intelligence model that generates text and speech by mimicking a user's speech patterns and thought processes.
[0229] A "prompt" is text input to a generative AI model, instructing it on how to respond.
[0230] A "smartphone" is a mobile phone that is capable of connecting to the internet and running applications.
[0231] A "head-mounted display" is a display device worn on the head that provides virtual reality and augmented reality experiences.
[0232] The system for implementing this invention collects user lifelog data and characteristic data, and generates a user avatar on the cloud based on this data. Specific embodiments of this system are described below.
[0233] First, the server collects user lifelog data and characteristic data from messenger apps, internet search tools, and other sources. Lifelog data includes the user's daily activity history, location information, and health data. Characteristic data includes data used to identify the individual, such as the user's text and voiceprint.
[0234] Next, the server uses this data as training data to train a generative AI model. This generative AI model is used to mimic the user's speech patterns and thought patterns. Specifically, it uses prompt sentences to mimic the user's speech patterns and thought patterns.
[0235] The server generates a virtual avatar of the user in the cloud. This avatar exists to maintain communication with acquaintances, friends, and family even when the user is absent, busy, or has passed away. The avatar can handle customer interactions in a virtual store. For example, if a customer asks, "What are the features of this product?", the avatar uses a generative AI model to generate an appropriate answer.
[0236] This system will be implemented as an application installed on smartphones and head-mounted displays. It can improve the customer experience by allowing an avatar to handle customer interactions even when the user is absent or busy.
[0237] As a concrete example, by using the following prompt, the avatar can mimic the user's speech patterns and thought processes to provide customer service.
[0238] Example of a prompt:
[0239] You are a virtual store clerk. Please answer the following questions by mimicking the user's language and thought patterns.
[0240] Question: {question}
[0241] In this way, even when the user is absent or busy, the avatar can handle customer interactions, and communication with acquaintances, friends, and family can be maintained even after the user's death. Furthermore, it becomes possible to provide customer service in virtual stores that mimics the user's speech patterns and thought processes.
[0242] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0243] Step 1:
[0244] The server collects user lifelog data and characteristic data from messenger apps and internet search tools. Inputs include user activity history, location information, health data, text, and voiceprints. This data is collected and stored in a database. The output is the collected lifelog data and characteristic data.
[0245] Step 2:
[0246] The server preprocesses the collected lifelog data and feature data. The data collected in step 1 is used as input. Data processing such as data cleaning, normalization, and feature extraction is performed to convert the data into a format suitable for training the generative AI model. The preprocessed data is obtained as output.
[0247] Step 3:
[0248] The server uses the pre-processed data as training data to train a generative AI model. The data pre-processed in step 2 is used as input. The generative AI model analyzes the data and extracts patterns to learn the user's language use and thought patterns. The trained generative AI model is obtained as output.
[0249] Step 4:
[0250] The server generates a user avatar in the cloud. The generative AI model trained in step 3 is used as input. Based on the generative AI model, a virtual entity is generated that mimics the user's speech patterns and thought patterns. The output is the avatar generated in the cloud.
[0251] Step 5:
[0252] The terminal runs an application that allows a user's avatar to interact with customers in a virtual store. The input includes the avatar generated in the cloud and questions from the customer. The terminal sends the customer's questions to the cloud and generates appropriate answers using a generative AI model. The output is the answer provided to the customer.
[0253] Step 6:
[0254] The server uses prompts to mimic the user's vocabulary and thought patterns. The input includes customer questions and prompts. The generative AI model uses the prompts to mimic the user's vocabulary and thought patterns and generates appropriate responses. The output is the mimicked response.
[0255] Step 7:
[0256] The device displays the customer's response through an application installed on a smartphone or head-mounted display. The input includes the response generated in step 6. The device displays the response to the customer, improving the customer experience. The output is the response displayed to the customer.
[0257] (Example 3)
[0258] Next, we will describe Embodiment 3 of Embodiment Example 3. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0259] In modern society, there is a need to digitize and permanently store and reproduce users' memories and personalities. However, conventional systems have difficulty accurately reproducing users' memories and personalities, and they lack mechanisms to improve the accuracy of reproduction through long-term use by the user. Therefore, a system is needed that can reproduce users' memories and personalities with high accuracy and further improve accuracy through long-term use.
[0260] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 3 is realized by the following means.
[0261] In this invention, the server includes means for linking life log data and characteristic data such as text and voiceprints from messenger apps, internet search tools, etc., means for the AI to reproduce the user's memories and personality using these as training data, means for generating an avatar of the user on the cloud, means for collecting data provided by the user and storing it in a database, means for preprocessing the collected data and converting it into a format suitable for the generative AI model, means for training the generative AI model using the preprocessed data, means for receiving prompt sentences from the user and generating a response using the generative AI model, and means for providing the generated response to the user. This makes it possible to reproduce the user's memories and personality with high accuracy and to further improve accuracy with long-term use.
[0262] A "messenger app" is software that allows users to send and receive text messages and voice messages.
[0263] An "internet search tool" is software that allows users to search for information on the internet.
[0264] "Life log data" refers to data related to a user's daily life, including behavioral history and activity records.
[0265] "Feature data" refers to data that indicates the individual characteristics of a user, and includes things like text and voiceprints.
[0266] "Training data" refers to data used to train a generative AI model, and is a dataset that has been assigned correct labels.
[0267] A "generative AI model" is an artificial intelligence model used to recreate a user's memories and personality.
[0268] "Cloud" refers to a collection of computer resources and services provided via the internet.
[0269] An "incarnation" is a digital entity that recreates the user's memories and personality.
[0270] A "database" is a system for systematically storing and managing collected data.
[0271] "Preprocessing" refers to the process of converting collected data into a format suitable for generating AI models.
[0272] A "prompt message" is a sentence of a question or request that a user enters into the system.
[0273] "Response" refers to the answer generated by the generative AI model for the prompt text.
[0274] This invention is a system that digitizes a user's memory and personality and generates an avatar on the cloud. The system collects life log data and feature data from messenger apps, internet search tools, etc., and uses this as teacher data for the generative AI model. Specific embodiments of this system will be described below.
[0275] First, the user provides life log data such as text messages, voice data, and image data through a messenger app or an internet search tool. The terminal sends this data to the server, and the server saves the received data in the database.
[0276] Next, the server preprocesses the collected data. Specifically, text data is tokenized using natural language processing (NLP) technology, voice data is converted to text using voice recognition software, and image data undergoes feature extraction using image recognition technology. Machine learning frameworks such as TENSORFLOW (registered trademark) and PyTorch are used for these preprocessings.
[0277] The preprocessed data is used for the training of the generative AI model. The server uses this data to train the generative AI model and makes the model learn patterns for reproducing the user's memory and personality. As the training progresses, the reproducibility of the model improves.
[0278] When the user inputs a question or request as a prompt text to the system, the terminal sends this prompt text to the server. The server uses the generative AI model to generate an appropriate response to the prompt text. The generated response is sent to the terminal and provided to the user.
[0279] As a specific example, when a user asks "What is my favorite movie?", the server uses a generative AI model to search for information about "favorite movies" from the user's past data and generates a response such as "Your favorite movie is Inception". Also, when the user asks "Where was the last travel destination I went to?", the server generates a response such as "You went to Hawaii last summer".
[0280] This system adopts a monthly billing system. The longer the user continues to be billed, the more data is collected and the learning of the AI progresses. As a result, the reproducibility of the avatar increases, and it becomes possible to more accurately reproduce the user's memory and personality. The flow of the specific process in Example 3 will be described using FIG. 15.
[0281] Step 1:
[0282] The user provides life log data such as text messages, voice data, and image data through a messenger app or an internet search tool. The terminal sends this data to the server. The input is the life log data from the user, and the output is the data sent to the server.
[0283] Step 2:
[0284] The server saves the received data in a database. Specifically, it stores text messages, voice data, and image data in the corresponding database tables. The input is the data sent from the terminal, and the output is the data saved in the database.
[0285] Step 3:
[0286] The server preprocesses the collected data. Text data is tokenized using natural language processing (NLP) techniques, and important keywords are extracted. Audio data is converted to text using speech recognition software. Image data is used for feature extraction using image recognition techniques. The input is raw data stored in a database, and the output is preprocessed data.
[0287] Step 4:
[0288] The server trains a generative AI model using preprocessed data. Specifically, it uses machine learning frameworks such as TensorFlow and PyTorch to train the model to reproduce patterns that replicate the user's memories and personality. The input is preprocessed data, and the output is the trained generative AI model.
[0289] Step 5:
[0290] The user inputs questions or requests to the system as prompts. The terminal sends these prompts to the server. The input is the prompt from the user, and the output is the prompt sent to the server.
[0291] Step 6:
[0292] The server inputs the received prompt into a generative AI model and generates an appropriate response. The generative AI model generates the most appropriate answer based on the user's past data. The input is the prompt and the trained generative AI model, and the output is the generated response.
[0293] Step 7:
[0294] The server sends the generated response to the terminal. The terminal displays this response to the user. The input is the generated response, and the output is the response provided to the user.
[0295] As a concrete example, if a user asks, "What's my favorite movie?", the server uses a generative AI model to search the user's past data for information about their "favorite movie" and generates a response such as, "Your favorite movie is Inception." Similarly, if a user asks, "Where was the last place I traveled to?", the server generates a response such as, "You went to Hawaii last summer."
[0296] (Application Example 3)
[0297] Next, we will describe application example 3 of form example 3. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".
[0298] Conventional systems had low accuracy in recreating user memories and personalities, making it difficult to recommend individually customized content. Furthermore, the lack of mechanisms to improve system accuracy through long-term user engagement meant that improving the user experience was a challenge.
[0299] The identification processing performed by the identification processing unit 290 of the data processing device 12 in Application Example 3 is realized by the following means. In this invention, the server includes means for linking life log data and characteristic data such as text and voiceprints from messenger apps, internet search tools, etc., means for the AI to reproduce the user's memories and personality using these as training data, means for generating an avatar of the person on the cloud, and means for collecting user data, generating an avatar using an AI model, and recommending individually customized content. This makes it possible to reproduce the user's memories and personality with high accuracy and recommend individually customized content.
[0300] A "messenger app" is software that allows users to send and receive text messages, voice messages, images, videos, and other content.
[0301] An "Internet search tool" is software for users to search for information on the Internet.
[0302] "Life log data" refers to data related to a user's daily life, including behavior history, location information, health data, etc.
[0303] "Characteristic data" refers to data for identifying an individual, such as a user's text or voiceprint.
[0304] "Teacher data" refers to data used to train a machine learning model.
[0305] "AI" is an abbreviation for artificial intelligence, which is a technology for computers to imitate human intelligence.
[0306] "To reproduce the user's memory and personality" means to imitate the user's past experiences and personality based on the collected data.
[0307] "Cloud" refers to a collection of computer resources and services provided through the Internet.
[0308] "Avatar" refers to a virtual existence that reproduces the user's memory and personality.
[0309] "Individually customized content" refers to information and entertainment specially selected based on the user's preferences and interests.
[0310] "AI model" refers to a mathematical model that uses artificial intelligence algorithms to analyze data and perform predictions and classifications.
[0311] "Monthly billing system" refers to a fee system where users can use services by paying a certain fee every month.
[0312] The system for implementing this invention is configured as follows: The server has means for linking life log data and characteristic data such as text and voiceprints from messenger apps, internet search tools, etc. This makes it possible to collect data about the user's daily life and data for identifying the individual.
[0313] Next, the server has the means for the AI to recreate the user's memories and personality using the collected data as training data. Specifically, it uses an artificial intelligence (AI) model to mimic the user's past experiences and personality based on the collected life log data and feature data. This AI model is implemented using the OpenAI API.
[0314] Furthermore, the server has the means to generate an avatar of the user in the cloud. The generated avatar is a virtual entity that reproduces the user's memories and personality, and is managed in the cloud.
[0315] Furthermore, the server has the means to collect user data, generate avatars using AI models, and recommend individually customized content. Specifically, it collects data such as the user's viewing history, preferences, and feedback, and uses this data to recommend the most suitable content to the user using an AI model. This recommendation provides information and entertainment specially selected based on the user's preferences and interests.
[0316] As users continue to pay for the service over a long period, the system collects more data, and the AI learns more effectively. This improves the accuracy of the avatar's reproduction, allowing it to more accurately recreate the user's memories and personality.
[0317] As a concrete example, consider a case where a user prefers "action movies" and "comedy movies," liked "movie1," and disliked "movie2." Based on this data, an avatar is generated, and an example of a prompt message recommending content is as follows.
[0318] Example of a prompt:
[0319] Based on the user's data: {"user_id": "user123", "viewing_history": ["movie1", "movie2"], "preferences": ["action", "comedy"], "feedback": ["liked movie1", "disliked movie2"]}, generate an avatar that recreates the user's memories and personality.
[0320] By inputting this prompt into the OpenAI API, an avatar is generated that replicates the user's memories and personality. Based on this generated avatar, it becomes possible to recommend content that is most suitable for the user.
[0321] The flow of the specific processing in Application Example 3 will be explained using Figure 16.
[0322] Step 1:
[0323] The server collects life log data and characteristic data from messenger apps and internet search tools. Specifically, it retrieves user text messages, voice messages, search history, location information, etc. The input is various user data, and the output is the collected life log data and characteristic data.
[0324] Step 2:
[0325] The server inputs the collected lifelog data and feature data into the AI model as training data. Specifically, it preprocesses this data and converts it into a format that the AI model can learn from. The input is the preprocessed data, and the output is the training data used to train the AI model.
[0326] Step 3:
[0327] The server uses an AI model to generate an avatar that replicates the user's memories and personality. Specifically, it uses the OpenAI API to generate prompt statements based on the user's data and inputs them into the AI model. The input is the prompt statements, and the output is the generated avatar.
[0328] Step 4:
[0329] The server stores the generated avatars in the cloud. Specifically, it uses a cloud storage service to securely store the avatar data. The input is the generated avatar, and the output is the avatar data stored in the cloud.
[0330] Step 5:
[0331] The server collects data such as the user's viewing history, preferences, and feedback, and uses an AI model to recommend individually customized content. Specifically, it generates prompt messages based on the user's data and inputs them into the AI model. The input is the user's viewing history and preference data, and the output is the recommended content.
[0332] Step 6:
[0333] Users receive and view or use recommended content. Specifically, they use smartphones or other devices to view and enjoy the recommended content. The input is the recommended content, and the output is the user's viewing and usage history.
[0334] Step 7:
[0335] The server collects more data and continues training the AI model as users continue to pay for the service over a long period. Specifically, it periodically collects user usage data through a monthly subscription system and uses it to retrain the AI model. The input is the continuously collected user data, and the output is the retrained AI model.
[0336] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0337] "Example of form 1"
[0338] One embodiment of the present invention is a system that incorporates an emotion engine. This system recognizes the user's emotions and adjusts the avatar's response accordingly. Specifically, when the user is feeling happy, the avatar speaks in a cheerful tone. Conversely, when the user is sad, the avatar speaks in a calm tone. In this way, the emotion engine grasps the user's emotional state in real time and adjusts the avatar's response accordingly.
[0339] "Example of form 2"
[0340] Furthermore, the emotion engine controls the avatar's actions according to the user's emotional state. For example, when the user is angry, the avatar will choose to apologize. When the user is happy, the avatar will choose to offer congratulations. In this way, the emotion engine understands the user's emotional state and controls the avatar's actions accordingly.
[0341] "Example of form 3"
[0342] Furthermore, the emotion engine controls the avatar's facial expressions according to the user's emotional state. For example, when the user is happy, the avatar will smile. Conversely, when the user is sad, the avatar will make a sad face. In this way, the emotion engine understands the user's emotional state and controls the avatar's facial expressions accordingly.
[0343] The following describes the processing flow for each example of the form.
[0344] "Example of form 1"
[0345] Step 1: The emotion engine understands the user's emotional state in real time.
[0346] Step 2: The emotional engine adjusts the avatar's response according to the emotional state it has identified.
[0347] Step 3: The avatar delivers a pre-tuned response to the user.
[0348] "Example of form 2"
[0349] Step 1: The emotion engine understands the user's emotional state in real time.
[0350] Step 2: The emotional engine controls the avatar's actions according to the emotional state it has identified.
[0351] Step 3: The avatar performs controlled actions towards the user.
[0352] "Example of form 3"
[0353] Step 1: The emotion engine understands the user's emotional state in real time.
[0354] Step 2: The emotional engine controls the avatar's facial expressions according to the emotional state it has identified.
[0355] Step 3: The avatar displays controlled facial expressions to the user.
[0356] (Example 1)
[0357] Next, we will describe Example 1 of Form Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0358] In modern society, there is a need for systems that can accurately understand users' behavioral patterns and emotional states and engage in appropriate dialogue accordingly. However, conventional systems have struggled to recognize users' emotions in real time and generate appropriate responses. Furthermore, data collection and analysis to recreate users' memories and personalities have been insufficient, resulting in a failure to improve user satisfaction.
[0359] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0360] In this invention, the server includes means for linking life log data and characteristic data such as text and voiceprints from messenger apps, internet search tools, etc., means for the AI to reproduce the user's memories and personality using these as training data, means for generating an avatar of the user on the cloud, means for recognizing the user's emotions in real time using an emotion engine, and means for adjusting the tone and dialogue content of the avatar according to the user's emotional state. This makes it possible to conduct dialogue in real time according to the user's emotional state and reproduce the user's memories and personality with high accuracy.
[0361] A "messenger app" is a software application that allows users to send and receive text messages, voice messages, images, videos, and other content.
[0362] An "internet search tool" is a software application that allows users to search for information on the internet.
[0363] "Life log data" refers to data related to a user's daily life, including behavioral history, location information, and activity records.
[0364] "Text" refers to text data entered by the user, including messages, comments, and posts.
[0365] A "voiceprint" is characteristic data extracted from a user's voice data, and includes voice characteristics that can be used to identify an individual.
[0366] "Feature data" refers to data that shows specific patterns or characteristics extracted from user behavior and statements.
[0367] "Training data" refers to a dataset used to train a machine learning model, and includes input data and corresponding correct answer data.
[0368] "AI" is an abbreviation for artificial intelligence, which is a technology in which computers imitate human intelligence to learn, reason, recognize, and make judgments.
[0369] "Recreating the user's memories and personality" means using collected data to recreate the user's past actions and statements, and to mimic the user's personality and preferences.
[0370] "Cloud" refers to a collection of computer resources and services provided via the internet, serving as infrastructure for data storage and processing.
[0371] An "avatar" is a virtual entity that interacts with the user on their behalf, and is generated based on the user's memories and personality.
[0372] An "emotion engine" is a software component that analyzes and recognizes emotions from a user's text or voice.
[0373] "Recognizing in real time" means instantly analyzing emotions in response to user input and providing the results.
[0374] "Tone" refers to the tone or atmosphere of voice or text in a conversation, and is an element used to convey emotions and intentions.
[0375] "Adjusting the dialogue content" means changing the content and expression of the messages delivered by the avatar according to the user's emotional state.
[0376] This invention is a system that collects life log data and characteristic data such as text and voiceprints from messenger apps and internet search tools that users use daily, and uses this data to reconstruct the user's memories and personality. Furthermore, by combining it with an emotion engine, it enables dialogue that responds to the user's emotional state.
[0377] Data collection
[0378] When users use messenger apps (e.g., WhatsApp) or internet search tools (e.g., Google® Search), life log data (e.g., chat history, search history) and characteristic data such as text and voiceprints are collected from these applications. Scripts that retrieve data via APIs are used for data collection.
[0379] Data preprocessing
[0380] The server receives the collected data and filters out unnecessary information. For example, it removes unwanted advertising messages from chat history and extracts only the user's messages. The Python pandas library is used for this process.
[0381] Data analysis and feature extraction
[0382] The server analyzes the pre-processed data to extract user behavior patterns, hobbies, interests, personality traits, and other characteristics. For example, natural language processing (NLP) techniques are used to extract emotions and topics from user statements. Libraries such as NLTK and spaCy are used for this analysis.
[0383] AI model training
[0384] The server uses the extracted feature data as training data to train a generative AI model (e.g., GPT-4). This allows the model to acquire the ability to reproduce the user's memories and personality. Frameworks such as TensorFlow and PyTorch are used for training.
[0385] emotion recognition
[0386] The server uses an emotion engine (e.g., IBM Watson®'s emotion analysis API) to recognize emotions in real time from the user's text and voice. For example, if a user sends the message "I had a great time today!", the emotion engine recognizes the emotion as "happy".
[0387] Adjusting the reaction of the avatar
[0388] The device adjusts the avatar's tone and dialogue based on information from the emotion engine. For example, if the user is happy, the avatar will respond in a cheerful tone, saying, "That's wonderful! What happened?" Conversely, if the user is sad, the avatar will respond in a calm tone, saying, "That must have been tough. Is there anything I can do to help?"
[0389] Specific examples and prompt statements
[0390] Specific example
[0391] If a user sends the message "I had so much fun today!" via a messenger app, the server receives this message and recognizes the emotion of "fun" through its emotion engine. Based on this information, the device's avatar responds in a cheerful tone, "That's wonderful! What happened?"
[0392] Example of a prompt
[0393] "Design a system that analyzes the emotions of users from messages sent via messenger apps and has their avatar respond in an appropriate tone. Specifically, if the user is happy, the response should be in a cheerful tone; if they are sad, the response should be in a calm tone."
[0394] In this way, a system is built that enables dialogue tailored to the user's emotional state.
[0395] The flow of the specific processing in Example 1 will be explained using Figure 17.
[0396] Step 1:
[0397] When users use messenger apps or internet search tools, life log data and characteristic data such as text and voiceprints are collected from these applications.
[0398] Input: User's chat history, search history, voice data
[0399] Output: Collected life log data, text data, voiceprint data
[0400] Specific operation: Execute a script that retrieves data via the API and saves it to the database.
[0401] Step 2:
[0402] The server receives the collected data and filters out unnecessary information.
[0403] Input: Collected life log data, text data, voiceprint data
[0404] Output: Preprocessed data (data with unnecessary information removed)
[0405] Specific actions: Use the Python pandas library to create a dataframe and delete unnecessary rows and columns.
[0406] Step 3:
[0407] The server analyzes the pre-processed data to extract user behavior patterns and characteristics such as hobbies, interests, and personality.
[0408] Input: Preprocessed data
[0409] Output: User characteristic data (behavioral patterns, hobbies, interests, personality, etc.)
[0410] Specific operation: Uses natural language processing (NLP) techniques to extract sentiment and topics from text data. Libraries such as NLTK and spaCy are used.
[0411] Step 4:
[0412] The server uses the extracted feature data as training data to train a generative AI model.
[0413] Input: User characteristic data
[0414] Output: Trained generative AI model
[0415] Specific operation: Train a generative AI model (e.g., GPT-4) using frameworks such as TensorFlow or PyTorch.
[0416] Step 5:
[0417] The server uses an emotion engine to recognize emotions in real time from the user's text and voice.
[0418] Input: User's text data, audio data
[0419] Output: Recognized emotion data
[0420] Specific operation: Call an emotion engine (e.g., IBM Watson's Sentiment Analysis API) to analyze emotions.
[0421] Step 6:
[0422] The device adjusts the avatar's tone and dialogue based on information from the emotion engine.
[0423] Input: Recognized emotion data
[0424] Output: Adjusted dialogue
[0425] Specific actions: Based on emotional data, the avatar generates dialogue and responds to the user. For example, if the user is happy, the avatar will respond in a cheerful tone, "That's wonderful! What happened?"
[0426] (Application Example 1)
[0427] Next, we will describe Application Example 1 of Form Example 1. In the following description, the data processing device 12 will be referred to as a "server," and the smart device 14 will be referred to as a "terminal."
[0428] Traditional content distribution services have faced challenges in providing personalized content recommendations that adequately reflect users' emotional states and individual hobbies and interests. Furthermore, it has been difficult to understand users' emotional states in real time and provide appropriate content accordingly.
[0429] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means. In this invention, the server includes means for linking life log data and characteristic data such as text and voiceprints from messenger apps, internet search tools, etc., means for the AI to reproduce the user's memories and personality using these as training data, means for generating an avatar of the person on the cloud, means including an emotion engine that grasps the user's emotional state in real time, and means for recommending optimal content according to the user's emotional state. This makes it possible to provide personalized content recommendations that are tailored to the user's emotional state and individual hobbies and interests.
[0430] A "messenger app" is an application that allows users to send and receive text messages, voice messages, images, videos, and other content.
[0431] An "internet search tool" is software or a web service that allows users to search for information on the internet.
[0432] "Life log data" refers to records of a user's actions and activities in their daily life.
[0433] "Feature data" refers to data that represents individual characteristics of a user, such as their written text or voiceprint.
[0434] "Training data" refers to a dataset used to train machine learning models.
[0435] "AI" is an abbreviation for artificial intelligence, which is a technology in which computers imitate human intelligence.
[0436] "Means for recreating a user's memories and personality" refers to technologies that recreate a user's past behavior and personality based on collected data.
[0437] "Cloud" refers to a collection of computer resources and services provided via the internet.
[0438] An "avatar" is a digital avatar that functions as a representation of the user.
[0439] An "emotion engine" is a technology that recognizes a user's emotional state in real time and generates a response accordingly.
[0440] "Content recommendation" is a technology that suggests the most suitable content based on the user's tastes, interests, and emotional state.
[0441] The system for implementing this invention is configured as follows: First, it collects life log data and characteristic data such as text and voiceprints from messenger apps and internet search tools installed on the user's smartphone or personal computer. This data reflects the user's behavioral patterns, hobbies, interests, personality, etc.
[0442] The server uses collected lifelog data and feature data as training data to recreate the user's memories and personality using a generative AI model. This generative AI model is implemented using programming languages such as Python and runs on the cloud. A digital avatar (incarnation) of the user is generated on the cloud, enabling conversations that reflect the user's memories and personality.
[0443] Furthermore, it incorporates an emotion engine that understands the user's emotional state in real time. This emotion engine analyzes the user's text and voice messages to recognize their emotional state. For example, if a user types "I'm a little tired today," the emotion engine detects the user's fatigue.
[0444] The server recommends the most suitable content to the user based on the emotional state obtained from the emotion engine. Content recommendations are based on the user's behavior patterns and emotional state, suggesting movies, music, and other content that will help the user relax.
[0445] For example, if a user types "I'm a little tired today," the emotion engine detects the user's fatigue and recommends relaxing movies or music. This process is achieved by inputting the following prompt into the generating AI model:
[0446] "Analyze the user's emotional state and recommend relaxing content. The user's input is 'I'm a little tired today.'"
[0447] In this way, personalized content recommendations tailored to the user's emotional state and individual hobbies and interests become possible.
[0448] The flow of a specific process in Application Example 1 will be explained using Figure 18.
[0449] Step 1:
[0450] Users use messenger apps and internet search tools installed on their smartphones or personal computers to send everyday messages and perform searches. This generates life log data and characteristic data such as text and voiceprints. The input is the user's messages and search queries, and the output is life log data and characteristic data.
[0451] Step 2:
[0452] The device sends the generated lifelog data and feature data to the server. The input is lifelog data and feature data, and the output is the data sent to the server. Specifically, the operation involves data collection and transmission.
[0453] Step 3:
[0454] The server inputs the received lifelog data and feature data into a generated AI model as training data to recreate the user's memories and personality. The input is lifelog data and feature data, and the output is a model that recreates the user's memories and personality. Specifically, the process involves data preprocessing, model training, and model generation.
[0455] Step 4:
[0456] The server generates a digital avatar (incarnation) of the user in the cloud. The input is a model that reproduces the user's memories and personality, and the output is the digital avatar in the cloud. Specifically, the server generates the avatar and places it in the cloud.
[0457] Step 5:
[0458] When a user enters a message, the terminal sends that message to the server. The input is the user's message, and the output is the message sent to the server. Specifically, the action involves sending a message.
[0459] Step 6:
[0460] The server inputs received messages into the emotion engine, allowing it to understand the user's emotional state in real time. The input is the user's message, and the output is the user's emotional state. Specifically, the system performs message analysis and emotional state recognition.
[0461] Step 7:
[0462] The server recommends optimal content based on the emotional state obtained from the emotion engine. The input is the user's emotional state, and the output is the recommended content. Specifically, the process involves evaluating the emotional state and selecting content.
[0463] Step 8:
[0464] The server sends recommended content to the terminal and presents it to the user. The input is the recommended content, and the output is the content presented to the user. Specifically, the operation involves sending and displaying content.
[0465] In this way, personalized content recommendations tailored to the user's emotional state and individual hobbies and interests are realized.
[0466] (Example 2)
[0467] Next, we will describe Example 2 of Form Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0468] In modern society, users are often too busy to maintain communication with acquaintances, friends, and family. Furthermore, there is a need for those left behind to continue communicating with a deceased person after their passing. However, conventional technology has struggled to accurately reproduce a user's speech patterns and thought processes, and to engage in appropriate dialogue tailored to their emotional state.
[0469] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0470] In this invention, the server includes means for linking life log data and characteristic data such as text and voiceprints from messenger apps, internet search tools, etc., means for artificial intelligence to recreate the user's memories and personality using these as training data, means for generating an avatar of the user on the cloud, means for analyzing the user's emotional state, means for generating dialogue with the avatar based on the emotional analysis results, and means for executing the generated dialogue. This makes it possible for the user to maintain communication with acquaintances, friends, and family even when they are absent, busy, or have passed away.
[0471] A "messenger app" is software that allows users to send and receive text messages, voice messages, images, videos, and other content.
[0472] An "internet search tool" is software that allows users to search for information on the internet and obtain the necessary data.
[0473] "Life log data" refers to data related to a user's daily life, including activity history, location information, and health data.
[0474] "Feature data" refers to data that represents the individual characteristics of a user, including writing style and voiceprints.
[0475] "Artificial intelligence" is a technology in which computers imitate human intelligence and perform learning, reasoning, and self-correction.
[0476] "Cloud" refers to a collection of computer resources and services provided via the internet.
[0477] An "avatar" is a virtual entity that mimics the user's speech patterns and thought processes, and engages in dialogue on their behalf.
[0478] "Emotional state" refers to the state of a user's current emotions, including joy, anger, sadness, etc.
[0479] "Emotion analysis" is the process of reading and analyzing emotions from a user's text and voice data.
[0480] "Dialogue generation" is the process of generating appropriate responses based on the user's emotional state and past dialogue history.
[0481] "Dialogue execution" is the process of actually sending the generated dialogue on behalf of the user.
[0482] This invention is a system that generates a user's avatar on the cloud, enabling communication with acquaintances, friends, and family even when the user is absent, busy, or has passed away. Specific embodiments of this system are described below.
[0483] First, the server collects user data. Specifically, it obtains life log data and characteristic data such as text and voiceprints from messenger apps and internet search tools. This data is used as training data to learn the user's language use and thought patterns.
[0484] Next, the server analyzes the collected data and learns the user's language use and thought patterns. Specifically, it analyzes the data using natural language processing techniques and learns user characteristics using generative AI models (for example, Google Cloud Natural Language API or OpenAI's GPT-4).
[0485] Subsequently, the server generates a user avatar based on the learned model. This avatar mimics the user's speech patterns and thought processes, and interacts on the user's behalf. The avatar is located in the cloud and sends and receives messages on the user's behalf.
[0486] Furthermore, the server uses an emotion engine to analyze the user's emotional state. Specifically, it analyzes the user's text and voice data to understand their emotional state in real time. For example, it uses Microsoft® Azure® Text Analytics for Sentiment Analysis to perform emotion analysis.
[0487] Based on the emotion analysis results, the server generates dialogue for the avatar. For example, if the user is angry, it will choose words of apology; if they are happy, it will choose words of congratulations. The generated dialogue will be appropriate to the user's emotional state.
[0488] Finally, the device executes the generated conversation and sends the message on behalf of the user. This allows the user to maintain communication with acquaintances, friends, and family even when they are away or busy.
[0489] For example, if a user is too busy to reply, their avatar will send a message on their behalf. For instance, if a user receives a message from a friend asking "What are your plans for this weekend?", the avatar will reply "I'm planning to spend this weekend with my family" based on the user's past conversation history. Also, if a user is angry, the emotion engine will detect that emotion and the avatar will choose an apology. For example, if a user sends an angry message saying "How did this happen!", the avatar will reply "I'm sorry, I'll take care of it right away."
[0490] Examples of prompt messages include the following:
[0491] "Based on the user's past message history, generate an appropriate reply to the following message: 'What are your plans for this weekend?'"
[0492] "Generate an appropriate response for when a user is angry. User message: 'How did this happen!'"
[0493] In this way, the server and terminal work together to generate a user avatar and use an emotion engine to engage in appropriate conversations that correspond to the user's emotional state.
[0494] The flow of the specific processing in Example 2 will be explained using Figure 19.
[0495] Step 1:
[0496] The server collects user data. Specifically, it obtains life log data and characteristic data such as text and voiceprints from messenger apps and internet search tools. User account information and access permissions are required as input, and the collected data is obtained as output. For example, email data is collected using the Gmail API.
[0497] Step 2:
[0498] The server analyzes the collected data and learns the user's language use and thought patterns. Specifically, it uses natural language processing techniques to analyze the data and employs generative AI models (e.g., Google Cloud Natural Language API or OpenAI's GPT-4). The collected data is required as input, and the output is a model that captures the user's characteristics. For example, it analyzes past message history to extract patterns in the user's language use.
[0499] Step 3:
[0500] The server generates a user avatar based on a trained model. This avatar mimics the user's speech patterns and thought processes, and interacts on their behalf. The trained model is required as input, and the avatar, deployed on the cloud, is generated as output. For example, a user avatar is created and deployed on the cloud.
[0501] Step 4:
[0502] The server analyzes the user's emotional state using an emotion engine. Specifically, it analyzes the user's text and voice data to understand their emotional state in real time. New messages and voice data from the user are required as input, and the emotion analysis results are obtained as output. For example, emotion analysis is performed using Microsoft Azure's Text Analytics for Sentiment Analysis.
[0503] Step 5:
[0504] The server generates dialogue for the avatar based on the sentiment analysis results. For example, if the user is angry, it will choose words of apology, and if the user is happy, it will choose words of congratulations. The input requires the sentiment analysis results and a model that captures the user's characteristics, and the output is the generated dialogue. For example, if the user sends an angry message such as "How did this happen!", the avatar will reply, "I'm sorry, I'll take care of it right away."
[0505] Step 6:
[0506] The device executes the generated dialogue and sends the message on behalf of the user. The generated dialogue is required as input, and the sent message is obtained as output. For example, if the user receives a message from a friend asking "What are your plans for this weekend?", the avatar will reply "I'm planning to spend this weekend with my family."
[0507] In this way, the server and terminal work together to generate a user avatar and use an emotion engine to engage in appropriate conversations that correspond to the user's emotional state.
[0508] (Application Example 2)
[0509] Next, we will describe application example 2 of form example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".
[0510] In modern society, there are problems such as users being too busy to fully utilize content distribution services, or being unable to communicate appropriately when users are absent. There is also the challenge of providing appropriate responses tailored to users' emotional states. Furthermore, there is a lack of means to maintain communication with acquaintances, friends, and family after a user's death.
[0511] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0512] This invention includes a server that provides means for linking life log data and characteristic data such as text and voiceprints from messenger apps, internet search tools, etc., means for an AI to recreate the user's memories and personality using these as training data, means for generating an avatar of the user on the cloud, means for recommending content based on the user's viewing history and ratings, means for posting comments that mimic the user's speech patterns and thought patterns, and means for controlling the avatar's actions according to the user's emotional state. As a result, even if the user is busy, they can make full use of the content distribution service, appropriate communication can be maintained even when the user is absent, and appropriate responses can be made according to the user's emotional state. Furthermore, even after the user's death, communication with acquaintances, friends, and family can be maintained.
[0513] A "messenger app" is software that allows users to send and receive text messages, voice messages, images, videos, and other content.
[0514] An "internet search tool" is software that allows users to search for information on the internet.
[0515] "Life log data" refers to data related to a user's daily life, including activity history, location information, and health data.
[0516] "Feature data" refers to data used to identify an individual, such as a user's written text or voiceprint.
[0517] "Training data" refers to data used to train machine learning models, and consists of pairs of input data and corresponding correct answer data.
[0518] "AI" is an abbreviation for artificial intelligence, which is a technology in which computers imitate human intelligence to learn and reason.
[0519] "Means of recreating a user's memories and personality" refers to technologies that mimic a user's thought patterns and language use based on their past actions and statements.
[0520] "A means of creating a person's avatar on the cloud" refers to a technology that creates a digital alter ego of a user on a server accessible via the internet.
[0521] "Viewing history" refers to a record of content that a user has viewed in the past.
[0522] "Ratings" refer to records of evaluations and feedback given by users to content they have viewed.
[0523] "A means of recommending content" is a technology that suggests new content that users might be interested in, based on their viewing history and ratings.
[0524] "Methods for posting comments by mimicking language use and thought patterns" refers to technologies that generate and post comments that a user is likely to make, based on their past statements.
[0525] "Means of controlling the avatar's actions according to emotional state" refers to technology that analyzes the user's emotions and adjusts the avatar's actions and statements based on the results.
[0526] The embodiments for carrying out this invention will now be described. First, the entire system is built on the cloud and operates in conjunction with the user's terminal (smartphone or head-mounted display). The system consists of the following main components.
[0527] 1. Data Collection
[0528] The server collects life log data and characteristic data (such as text and voiceprints) from messenger apps and internet search tools. This data includes the user's activity history and statements.
[0529] 2. Incarnation generation
[0530] The server uses the collected data as training data to train a generative AI model (e.g., GPT-4). This model is used to recreate the user's memories and personality, specifically mimicking the user's speech patterns and thought processes.
[0531] 3. Avatar on the Cloud
[0532] The server creates a virtual avatar of the user in the cloud. This avatar communicates on the user's behalf when the user is absent or busy. It also maintains communication with acquaintances, friends, and family even after the user's death.
[0533] 4. Content Recommendation
[0534] The server uses a generative AI model to recommend new content based on the user's viewing history and ratings. For example, it might suggest the next movie a user should watch based on the genre of movies they've recently watched.
[0535] Example of a prompt
[0536] "The user recently watched an action movie. Please recommend three more movies to watch next."
[0537] 5. Post a comment
[0538] The server mimics the user's speech patterns and thought processes, posting comments on their behalf when they are absent or busy. This helps maintain the user's presence.
[0539] Example of a prompt
[0540] "Generate comments that users would post after watching 'Inception'."
[0541] 6. Emotional Response
[0542] The server uses an emotion engine (e.g., Emotion API) to analyze the user's emotional state. Based on the analysis, it adjusts the avatar's actions and statements. For example, it posts encouraging words when the user is sad and congratulatory words when the user is happy.
[0543] In this way, even busy users can fully utilize content delivery services, and appropriate communication can be maintained even when the user is absent. Furthermore, it becomes possible to respond appropriately to the user's emotional state, and communication with acquaintances, friends, and family can be maintained even after the user's death.
[0544] The flow of a specific process in Application Example 2 will be explained using Figure 20.
[0545] Step 1:
[0546] The server collects life log data and characteristic data (such as text and voiceprints) from messenger apps and internet search tools. The input is the user's behavior history and statements, and the output is the storage of this data in the cloud. Specifically, the data is retrieved via an API and stored in an AWS® S3 bucket.
[0547] Step 2:
[0548] The server uses the collected data as training data to train a generative AI model (e.g., GPT-4). The input is the collected lifelog data and feature data, and the output is a model designed to reproduce the user's memories and personality. Specifically, the process involves preprocessing the data and feeding it to the AI model for learning.
[0549] Step 3:
[0550] The server generates a user avatar in the cloud. The input is a trained AI model, and the output is a digital replica of the user. Specifically, it deploys the AI model to a server in the cloud and provides an API to generate the user avatar.
[0551] Step 4:
[0552] The server recommends new content using a generative AI model based on the user's viewing history and ratings. The input is the user's viewing history and ratings, and the output is the recommended content. Specifically, it analyzes the viewing history and ratings, inputs prompts into the AI model, and generates recommendation results.
[0553] Step 5:
[0554] The server mimics the user's speech patterns and thought processes to post comments on their behalf when they are absent or busy. The input is the user's past statements, and the output is a generated comment. Specifically, it analyzes past statements, inputs prompts into an AI model, and generates comments.
[0555] Step 6:
[0556] The server analyzes the user's emotional state using an emotion engine (e.g., the Emotion API). The input is the user's current emotional data, and the output is the analysis result. Specifically, it acquires emotional data in real time and performs analysis using the Emotion API.
[0557] Step 7:
[0558] The server adjusts the avatar's actions and statements based on the analysis results. The input is the analysis results of the emotion engine, and the output is the adjusted actions and statements of the avatar. Specifically, it inputs prompt sentences into the AI model based on the analysis results, and generates appropriate actions and statements.
[0559] (Example 3)
[0560] Next, we will describe Embodiment 3 of Embodiment Example 3. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0561] Conventional systems had low accuracy in recreating users' memories and personalities, and struggled to respond in real time to users' emotional states. Furthermore, insufficient long-term data collection and analysis prevented them from accurately reflecting users' emotional states. As a result, the user experience was limited, and the usefulness of the system was reduced.
[0562] The identification processing performed by the identification processing unit 290 of the data processing device 12 in Example 3 is realized by the following means. In this invention, the server includes means for linking life log data and characteristic data such as text and voiceprints from messenger apps, internet search tools, etc., means for the AI to reproduce the user's memories and personality using these as training data, means for generating an avatar of the person on the cloud, means for detecting the user's emotional state in real time, means for controlling the avatar's facial expression according to the user's emotional state, and means for increasing the accuracy of reproduction the longer the monthly subscription is continued. This makes it possible to reproduce the user's memories and personality with high accuracy and to respond in real time according to the emotional state.
[0563] A "messenger app" is software that allows users to send and receive text messages, voice messages, images, videos, and other content in real time.
[0564] An "internet search tool" is software that allows users to search for information on the internet and obtain the necessary data.
[0565] "Life log data" refers to data that includes a user's daily life history and activity records.
[0566] "Feature data" refers to unique data used to identify an individual, such as a user's text or voiceprint.
[0567] "Training data" refers to pairs of known inputs and outputs used to train a machine learning model.
[0568] "AI" is an abbreviation for artificial intelligence, a technology in which computers imitate human intelligence to learn, reason, and self-correct.
[0569] "Cloud" refers to a collection of computer resources and services provided via the internet.
[0570] An "avatar" is a virtual entity that recreates the user's memories and personality, and acts as a proxy for the user, engaging in dialogue and other actions.
[0571] "Emotional state" refers to the state of a user's current emotions and mood.
[0572] "Real-time" refers to the immediate collection, analysis, and response of data.
[0573] A "monthly subscription" is a pricing system where users pay a fixed fee each month to use a service.
[0574] "Reproducibility" refers to the degree to which a system can accurately reproduce a user's memories and personality.
[0575] Modes for carrying out the invention
[0576] This invention is a system that accurately reproduces a user's memories and personality and responds in real time according to their emotional state. Specific embodiments of this system are described below.
[0577] 1. Program generation
[0578] The server collects user data and generates a program that recreates the user's memories and personality using a generative AI model. This program also includes an emotion engine that controls the avatar's facial expressions according to the user's emotional state.
[0579] 2. Program Processing Description
[0580] The server collects more data as users continue to pay monthly subscriptions. This data includes the user's behavioral history, emotional state, and conversation content. The collected data is analyzed using generative AI models (e.g., OpenAI's GPT-4) and used to reconstruct the user's memory and personality.
[0581] The device uses hardware such as a camera and microphone to detect the user's emotional state in real time. For example, the camera captures the user's facial expressions, and the microphone analyzes the tone of the user's voice. This data is sent to an emotion engine to determine the user's emotional state.
[0582] The emotion engine controls the avatar's facial expressions based on the user's emotional state. For example, when the user is happy, the emotion engine makes the avatar smile. Conversely, when the user is sad, the emotion engine makes the avatar make a sad expression.
[0583] 3. Specific Examples and Examples of Prompt Statements
[0584] As a concrete example, consider a case where a user continues to pay for a subscription for one year. If a user continues to pay for a year, the system collects a vast amount of data about the user's behavior history and emotional state. Based on this data, a generative AI model can reproduce the user's memory and personality with very high accuracy.
[0585] Example of a prompt:
[0586] Based on data from a user who has continued paying for a year, please generate a program to recreate the user's memories and personality. Additionally, configure the emotion engine so that the avatar smiles when the user is happy and makes a sad expression when the user is sad.
[0587] In this way, the server collects data and generates a program that recreates the user's memories and personality using a generative AI model. The terminal detects the user's emotional state, and the emotion engine controls the avatar's facial expressions. As the user continues to pay for the service over a long period, the system collects more data, and the accuracy of the avatar's recreation improves. The flow of specific processing in Example 3 will be explained using Figure 21.
[0588] Program processing flow
[0589] Step 1: Collecting user data
[0590] Users log into the system and engage in everyday activities and interactions. The device uses a camera and microphone to capture the user's facial expressions and voice tone in real time. It also collects text data and activity history entered by the user.
[0591] Input: User facial expression data, voice tone, text data, behavioral history
[0592] Output: Collected user data
[0593] Specific actions:
[0594] The device's camera captures the user's face and collects facial expression data.
[0595] The device's microphone records the user's voice and analyzes their tone and emotions.
[0596] The system collects text data entered by users in chat.
[0597] Step 2: Sending data to the server
[0598] The device periodically sends the collected data to the server. The data is encrypted and transferred securely.
[0599] Input: Collected user data
[0600] Output: Data sent to the server
[0601] Specific actions:
[0602] The device encrypts the data it collects and sends it to the server.
[0603] The server saves the received data to the database.
[0604] Step 3: Data analysis on the server
[0605] The server analyzes the received data to identify user behavior patterns and emotional states. Data mining techniques and machine learning algorithms are used for this analysis.
[0606] Input: Data sent to the server
[0607] Output: Analysis results (user behavior patterns, emotional state)
[0608] Specific actions:
[0609] The server retrieves data from the database and applies an analysis algorithm.
[0610] The server identifies the user's behavioral patterns and emotional state, and saves the results.
[0611] Step 4: Recreating the user's memories and personality using a generative AI model.
[0612] The server uses a generated AI model (e.g., OpenAI's GPT-4) based on the analysis results to recreate the user's memories and personality. The generated model then simulates interactions and actions with the user.
[0613] Input: Analysis results
[0614] Output: Reconstructed user memories and personality
[0615] Specific actions:
[0616] The server inputs the analysis results into an AI model that recreates the user's memories and personality.
[0617] The generated model simulates user interactions and actions.
[0618] Step 5: Detecting the user's emotional state using the emotion engine
[0619] The device uses an emotion engine to detect the user's emotional state in real time. The emotion engine analyzes data from the camera and microphone to determine the user's emotional state.
[0620] Input: Real-time data from cameras and microphones
[0621] Output: User's emotional state
[0622] Specific actions:
[0623] The device's emotion engine analyzes facial expression data from the camera to determine the user's emotional state.
[0624] The device's emotion engine analyzes the tone of voice from the microphone to determine the user's emotional state.
[0625] Step 6: Controlling the avatar's facial expressions
[0626] After the emotion engine determines the user's emotional state, the device controls the avatar's facial expressions. For example, when the user is happy, the avatar will smile. Conversely, when the user is sad, the avatar will make a sad face.
[0627] Input: User's emotional state
[0628] Output: The avatar's expression
[0629] Specific actions:
[0630] The device controls the avatar's facial expressions based on the results of the emotion engine's judgment.
[0631] The avatar creates appropriate facial expressions according to the user's emotional state.
[0632] (Application Example 3)
[0633] Next, we will describe application example 3 of form example 3. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".
[0634] Traditional avatar systems focus on recreating the user's memories and personality, but they have limitations in providing real-time responses based on the user's emotional state and delivering personalized content. Furthermore, the lack of means to analyze the user's emotional state could potentially degrade the quality of the user experience.
[0635] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 3 is realized by the following means. In this invention, the server includes means for linking life log data and characteristic data such as text and voiceprints from messenger apps, internet search tools, etc., means for the AI to reproduce the user's memories and personality using these as training data, means for generating an avatar of the person on the cloud, means for analyzing the user's emotional state in real time, and means for providing personalized content according to the analyzed emotional state. This makes it possible to provide real-time responses and personalized content according to the user's emotional state.
[0636] A "messenger app" is software that allows users to send and receive text messages and voice messages.
[0637] An "internet search tool" is software used to search for information on the internet.
[0638] "Life log data" refers to data related to a user's daily life, including activity history and location information.
[0639] "Feature data" refers to data that indicates the individual characteristics of a user, and includes things like text and voiceprints.
[0640] "Training data" refers to data used to train machine learning models.
[0641] "AI" is an abbreviation for artificial intelligence, which is a technology in which computers imitate human intelligence.
[0642] "Recreating the user's memories and personality" means imitating the user's character and thought patterns based on their past actions and statements.
[0643] "Creating a virtual avatar of oneself on the cloud" means creating a virtual alter ego of a user on a server accessible via the internet.
[0644] "Analyzing the user's emotional state in real time" means instantly determining the user's emotions at any given moment based on their facial expressions, voice, and other factors.
[0645] "Providing personalized content" means offering the most suitable information and entertainment based on the user's individual preferences and emotional state.
[0646] The system for carrying out this invention uses the following hardware and software to provide personalized content that responds to the user's emotional state.
[0647] hardware
[0648] 1. Camera: Used to capture the user's face. This utilizes the camera built into the smartphone or head-mounted display.
[0649] 2. Server: Use a cloud server for data processing and storage.
[0650] software
[0651] 1. OpenCV: This is a library for face detection. It analyzes video captured from a camera to detect faces.
[0652] 2. Keras: A deep learning library for implementing emotion recognition models. It predicts a user's emotions from detected facial images.
[0653] 3. Requests: This is a library for communicating with content delivery servers. It is used to retrieve content that matches the user's emotional state.
[0654] Data processing and data calculation
[0655] 1. Data Capture: Capture video from the camera in real time and perform face detection. Use OpenCV to detect faces, convert the detected faces to grayscale, and adjust their size.
[0656] 2. Emotion Prediction: The adjusted facial image is input into Keras's emotion recognition model to predict the user's emotion. There are seven types of emotions: anger, disgust, fear, happiness, sadness, surprise, and neutral expression.
[0657] 3. Content Delivery: Based on predicted sentiment, the Requests library is used to send a request to the content delivery server and retrieve appropriate content. The retrieved content is then displayed to the user.
[0658] Specific example
[0659] For example, if a user is using their smartphone and the camera captures their face and recognizes their emotion as "happy," the content delivery server might offer them a "comedy movie" or a "fun music playlist."
[0660] Example of a prompt
[0661] "Develop an application that analyzes user emotions in real time and provides personalized content based on those emotions. Use a Keras model for emotion recognition and the Requests library for communication with the content delivery server."
[0662] The flow of the specific processing in Application Example 3 will be explained using Figure 22.
[0663] Step 1:
[0664] The device captures the user's face in real time using its camera. The input is video data from the camera, and the output is the captured image frame. Specifically, the device activates the camera and continuously acquires video.
[0665] Step 2:
[0666] The device detects faces from captured image frames using OpenCV. The input is the captured image frame, and the output is the region of the detected face. Specifically, the device converts the image to grayscale and applies a face detection algorithm to determine the location of the face.
[0667] Step 3:
[0668] The device converts the detected face region to grayscale and resizes it to 48x48 pixels. The input is the detected face region, and the output is the resized face image. Specifically, the device normalizes the face image and converts it to a format suitable for machine learning models.
[0669] Step 4:
[0670] The device predicts emotions from resized face images using a Keras emotion recognition model. The input is a resized face image, and the output is the predicted emotion label. Specifically, the device inputs the face image into the model, calculates the probability distribution of emotions, and selects the emotion with the highest probability.
[0671] Step 5:
[0672] The device uses the Requests library to send a request to the content delivery server based on the predicted sentiment label. The input is the predicted sentiment label, and the output is the content retrieved from the content delivery server. Specifically, the device sends the sentiment label to the server in JSON format and requests appropriate content.
[0673] Step 6:
[0674] The server selects personalized content based on the received sentiment label and sends it back to the device. The input is the sentiment label, and the output is the selected content. Specifically, the server searches the database for content corresponding to the sentiment label and sends it to the device.
[0675] Step 7:
[0676] The terminal displays content received from the server to the user. The input is the content received from the server, and the output is the content displayed to the user. Specifically, the terminal displays the received content on the screen, making it viewable or usable by the user.
[0677] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0678] Data generation model 58 is a form of so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0679] Other examples of generative AI include Gemini® (registered trademark) (Internet search). <url: https: gemini.google.com ?hl="ja">) are some examples.
[0680] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0681] [Second Embodiment]
[0682] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0683] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0684] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0685] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0686] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0687] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0688] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0689] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0690] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0691] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0692] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0693] Next, the identification process performed by the identification processing unit 290 of the data processing device 12 will be described.
[0694] "Example of form 1"
[0695] One embodiment of the present invention involves collecting life log data and characteristic data such as text and voiceprints from messenger apps and internet search tools that users use daily. This data reflects the user's behavioral patterns, hobbies, interests, and personality. By using this data as training data to train an AI, the user's memories and personality can be reproduced.
[0696] "Example of form 2"
[0697] Next, a virtual avatar of the user is created in the cloud. This avatar is there to maintain communication with acquaintances, friends, and family even when the user is absent, busy, or after the user's death. Specifically, the avatar mimics the user's speech patterns and thought processes, recreating how the user normally interacts with others.
[0698] "Example of form 3"
[0699] Furthermore, the system of this invention employs a monthly subscription model. The longer a user continues to subscribe, the more data the system collects, and the more the AI learns. This increases the accuracy of the avatar's reproduction, making it possible to more accurately reproduce the user's memories and personality.
[0700] The following describes the processing flow for each example of the form.
[0701] "Example of form 1"
[0702] Step 1: Collect life log data and characteristic data such as text and voiceprints from messenger apps and internet search tools that the user uses on a daily basis.
[0703] Step 2: The collected data is used as training data to train the AI. Through this training, the AI understands the user's behavior patterns, hobbies, interests, personality, etc.
[0704] Step 3: As the AI learns, its ability to recreate the user's memories and personality improves.
[0705] "Example of form 2"
[0706] Step 1: Create a virtual avatar of the user in the cloud. This avatar will maintain communication with acquaintances, friends, and family even when the user is absent, busy, or has passed away.
[0707] Step 2: The avatar mimics the user's speech patterns and thought processes, recreating how the user typically interacts with others.
[0708] "Example of form 3"
[0709] Step 1: The system of the present invention employs a monthly subscription model.
[0710] Step 2: The longer a user continues to pay, the more data the system collects and the more the AI learns.
[0711] Step 3: This increases the accuracy of the avatar's reproduction, making it possible to more accurately recreate the user's memories and personality.
[0712] (Example 1)
[0713] Next, we will describe Example 1 of Form Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0714] In modern society, there is a need for technology that collects data reflecting users' behavioral patterns, interests, and personalities, and uses that data to reconstruct their memories and personalities. However, existing systems do not integrate the steps of data collection, preprocessing, storage, learning, evaluation, and generation, making it difficult to efficiently reconstruct users' memories and personalities. Furthermore, there is a lack of means to maintain communication with acquaintances, friends, and family even when the user is absent or after death. In addition, a sustainable billing model to improve the accuracy of the reconstruction has not been established.
[0715] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0716] In this invention, the server includes means for collecting life log data and characteristic data such as text and voiceprints from messenger apps, internet search tools, etc.; means for preprocessing the collected data and converting it into a format suitable for training an AI model; means for storing the preprocessed data in a database; means for inputting the stored data as training data into an AI model and performing training to reproduce the user's memories and personality; means for evaluating the AI model after training is complete and confirming its performance; means for saving models with good evaluations; and means for generating an avatar of the user on the cloud. This makes it possible to efficiently collect and process data that reflects the user's behavior patterns, interests, and personality, and to reproduce the user's memories and personality with high accuracy. Furthermore, it is possible to maintain communication with acquaintances, friends, and family even when the user is absent or after death, and the accuracy of reproduction can be improved through a continuous billing model.
[0717] A "messenger app" is a software application that allows users to send and receive text messages, voice messages, images, videos, and other content.
[0718] An "internet search tool" is a software application that allows users to search for information on the internet.
[0719] "Life log data" refers to data that records a user's actions and activities in their daily life.
[0720] "Feature data" refers to data that reflects an individual's characteristics, such as a user's written text or voiceprint.
[0721] "Preprocessing" refers to the process of converting collected data into a format suitable for training an AI model.
[0722] A "database" is a system for efficiently storing, managing, and retrieving data.
[0723] "Training data" refers to data with correct labels that is used to train an AI model.
[0724] An "AI model" is a mathematical model that uses artificial intelligence technology to analyze data and perform predictions and classifications.
[0725] "Learning" is the process by which an AI model discovers patterns and rules based on training data.
[0726] "Evaluation" is the process of checking the performance of an AI model after it has completed its training.
[0727] "Cloud" refers to a collection of computer resources and services provided via the internet.
[0728] An "incarnation" is a digital entity that recreates the user's memories and personality.
[0729] A "billing model" is a business model that collects fees for the use of a service.
[0730] Modes for carrying out the invention
[0731] This invention is a system that collects life log data and characteristic data such as text and voiceprints from messenger apps and internet search tools that users use on a daily basis, and uses this data to reconstruct the user's memories and personality. A specific embodiment of this system is described below.
[0732] Data collection
[0733] The server collects life log data and characteristic data such as text and voiceprints from messenger apps and internet search tools that users use on a daily basis. This collection uses common software such as messenger apps and internet search tools. The collected data reflects the user's behavioral patterns, hobbies, interests, personality, etc.
[0734] Data preprocessing
[0735] The server preprocesses the collected data and converts it into a format suitable for training the AI model. Specifically, this involves tokenizing text data and converting audio data into spectrograms. This is done using Python libraries such as NLTK and Librosa. For example, NLTK is used for tokenizing text data, and Librosa is used for converting audio data into spectrograms.
[0736] Data storage
[0737] The server stores the pre-processed data in a database. This uses a database system such as MySQL or MongoDB. Specifically, it establishes a database connection and inserts the pre-processed data into the appropriate tables or collections.
[0738] AI model training
[0739] The server inputs pre-processed data as training data into an AI model and performs training to reproduce the user's memories and personality. For this training, generative AI models such as GPT-4 or BERT are used. Specifically, the server performs model initialization, data batch processing, and training.
[0740] Model evaluation and saving
[0741] The server evaluates the trained AI models and verifies their performance. Evaluation metrics such as accuracy and recall are used. Models that perform well are saved to the file system or cloud storage.
[0742] Cloud-based avatar generation
[0743] The server uses a well-rated model to generate an avatar of the user in the cloud. This avatar is used to recreate the user's memories and personality, and to maintain communication with acquaintances, friends, and family.
[0744] Specific example
[0745] Example 1: Data collection from messenger apps
[0746] The server collects messages that users send and receive using common messenger apps. The collected messages reflect the user's conversation patterns and interests. By preprocessing this message data and training an AI model, the system can replicate the user's conversational style.
[0747] Specific example 2: Data collection from internet search tools
[0748] The server collects search queries that users make using common internet search tools. These collected search queries reflect the user's interests and preferences. By preprocessing this search data and training an AI model, the system can reproduce the user's interests.
[0749] Example of a prompt
[0750] "Create a program that collects messages sent by users through common messenger apps and trains an AI model to replicate the user's conversational style."
[0751] "Create a program that collects search queries that users make using common internet search tools and trains an AI model to replicate those users' interests."
[0752] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0753] Step 1: Data Collection
[0754] The server collects life log data and characteristic data such as text and voiceprints from messenger apps and internet search tools that users use on a daily basis. It receives log data from messenger apps and internet search tools as input and generates the collected raw data as output. Specifically, the server uses APIs and scraping techniques to acquire data and collects it using a secure communication protocol (e.g., HTTPS).
[0755] Step 2: Data Preprocessing
[0756] The server preprocesses the collected data and converts it into a format suitable for training the AI model. It receives raw data as input and generates preprocessed data as output. Specifically, the server performs the following operations:
[0757] Text data tokenization: Use Python's NLTK library to split text into words and phrases.
[0758] Spectrogram conversion of audio data: Convert audio data to a frequency spectrum using the Librosa library in Python.
[0759] Step 3: Save data
[0760] The server saves the pre-processed data to the database. It receives pre-processed data as input and generates data stored in the database as output. Specifically, the server performs the following operations:
[0761] Establishing a database connection: Connect to MySQL or MongoDB.
[0762] Data insertion: Insert pre-processed data into the appropriate tables or collections.
[0763] Step 4: Training the AI model
[0764] The server inputs pre-processed data as training data into the AI model and performs training to reproduce the user's memories and personality. It receives pre-processed data stored in a database as input and generates a trained AI model as output. Specifically, the server performs the following operations:
[0765] Model initialization: Initialize generative AI models such as GPT-4 and BERT.
[0766] Data batch processing: Pre-processed data is divided into batches and input into the model.
[0767] Running the training: Input data into the model and run the training.
[0768] Step 5: Evaluate and save the model.
[0769] The server evaluates the trained AI model and verifies its performance. It receives the trained AI model and evaluation data as input, and generates the evaluation results and a saved model as output. Specifically, the server performs the following operations:
[0770] Model evaluation: Calculate metrics such as accuracy and recall.
[0771] Model saving: Save models with good performance to the file system or cloud storage.
[0772] Step 6: Creating an avatar in the cloud
[0773] The server generates an avatar of the user in the cloud using a well-rated model. It receives a stored AI model as input and generates an avatar in the cloud as output. Specifically, the server performs the following operations:
[0774] Model Deployment: Deploy the AI model to the cloud environment.
[0775] Avatar Generation: Use the deployed model to generate an avatar that replicates the user's memories and personality.
[0776] (Application Example 1)
[0777] Next, we will describe Application Example 1 of Form Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0778] In modern society, there is a growing need to provide personalized content to users by leveraging their life log data and characteristic data. However, conventional systems have struggled to recreate users' memories and personalities and recommend personalized content. Furthermore, there has been a lack of means to maintain communication with acquaintances and family when users are absent, busy, or even after death. A new system is needed to solve these problems.
[0779] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0780] This invention includes a server that provides means for linking life log data and characteristic data such as text and voiceprints from messenger apps, internet search tools, etc., means for an AI to recreate the user's memories and personality using these as training data, means for generating an avatar of the user on the cloud, and means for automatically recommending content based on the user's preferences and interests based on the user's life log data and characteristic data. This makes it possible to recreate the user's memories and personality and provide personalized content. Furthermore, it enables the user to maintain communication with acquaintances and family even when they are absent, busy, or even after death.
[0781] A "messenger app" is an application that allows users to send and receive text messages, voice messages, images, videos, and other content.
[0782] An "internet search tool" is software or a web service that allows users to search for information on the internet.
[0783] "Life log data" refers to records of a user's actions and activities in their daily life.
[0784] "Feature data" refers to data used to identify individual users, such as user text or voiceprints.
[0785] "Training data" refers to a dataset used for AI to learn.
[0786] "AI" stands for artificial intelligence, which is a technology in which machines imitate human intelligence.
[0787] "Recreating the user's memories and personality" means that the AI imitates the user's thought and behavioral patterns based on the user's past actions and statements.
[0788] "Creating a digital avatar of oneself on the cloud" means creating a digital alter ego of a user using cloud computing technology.
[0789] "Automatically recommending content" means that AI suggests appropriate information and entertainment based on the user's preferences and interests.
[0790] The system for implementing this invention collects user life log data and characteristic data, reconstructs the user's memories and personality based on this data, and recommends personalized content. Specific embodiments of this system are described below.
[0791] System Configuration
[0792] hardware
[0793] Server: A server used for data collection, processing, storage, and training and inference of AI models.
[0794] Device: A device used by a user, such as a smartphone or computer.
[0795] software
[0796] Messenger app: An application that allows users to send and receive text messages, voice messages, images, videos, and other content.
[0797] Internet search tool: Software or web service used by users to search for information on the internet.
[0798] AI Model: A generative AI model that recreates the user's memories and personality based on the user's life log data and characteristic data, and recommends personalized content.
[0799] Data collection and processing
[0800] The server collects user lifelog and characteristic data from messenger apps and internet search tools. This includes the user's message history, search history, and voice data. The collected data is preprocessed and then used as training data for AI models.
[0801] AI model training and inference
[0802] The server trains an AI model based on the collected data. Specifically, text data is digitized using TfidfVectorizer, and user profiles are created using clustering algorithms (e.g., KMeans). Audio data is converted to text using speech recognition technology and processed similarly.
[0803] Content Recommendation
[0804] The server uses a pre-trained AI model to automatically recommend content based on the user's preferences and interests. This includes entertainment content such as movies, music, and articles. Recommended content is notified to the user's device.
[0805] Specific example
[0806] For example, if user A frequently talks about "movies" on a messenger app, the server collects that data and uses it to train an AI model. If user A's search history includes many searches for "latest movie reviews" and "movie trailers," the server creates a profile of user A based on this data. As a result, user A is recommended the latest movie review articles and movie trailer videos.
[0807] Example of a prompt
[0808] Develop an application that recommends content tailored to users' interests and preferences based on data collected from their messenger apps and search history. Specifically, implement a function that analyzes themes users frequently discuss and keywords they search for, and then automatically recommends content such as movies, music, and articles based on that analysis.
[0809] In this way, it is possible to build a system that utilizes users' life log data and characteristic data to deliver personalized content.
[0810] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0811] Step 1:
[0812] The server collects user lifelog and characteristic data from messenger apps and internet search tools. Specifically, it retrieves user message history, search history, voice data, etc. This data is collected using APIs and database queries.
[0813] Input: Raw data from messenger apps and search tools
[0814] Output: Collected lifelog data and feature data
[0815] Step 2:
[0816] The server preprocesses the collected data. Specifically, text data is digitized using TfidfVectorizer, and speech data is converted to text using speech recognition technology. This converts the data into a format suitable for the AI model.
[0817] Input: Collected life log data and feature data
[0818] Output: Preprocessed data
[0819] Step 3:
[0820] The server trains the AI model based on pre-processed data. Specifically, it creates user profiles using clustering algorithms (e.g., KMeans). This allows user behavior patterns and interests to be reflected in the model.
[0821] Input: Preprocessed data
[0822] Output: Trained AI model
[0823] Step 4:
[0824] The server uses a pre-trained AI model to automatically recommend content based on the user's preferences and interests. Specifically, it selects entertainment content such as movies, music, and articles based on the user's profile.
[0825] Input: Pre-trained AI model, user profile
[0826] Output: Recommended content
[0827] Step 5:
[0828] The server notifies the user's device of the recommended content. Specifically, it presents the content to the user using push notifications or in-app messages.
[0829] Input: Recommended content
[0830] Output: Content displayed on the user's device
[0831] In this way, a system is built that utilizes users' life log data and characteristic data to deliver personalized content.
[0832] (Example 2)
[0833] Next, we will describe Example 2 of Form Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0834] In modern society, there is a need to maintain communication with acquaintances, friends, and family even when users are absent, busy, or have passed away. However, conventional technologies have struggled to accurately mimic users' speech patterns and thought processes to achieve natural dialogue. Furthermore, there has been a lack of effective means to utilize users' past messages and dialogue history to train generative AI models. This has resulted in the problem of user avatars generating unnatural responses.
[0835] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0836] In this invention, the server includes means for linking life log data and characteristic data such as text and voiceprints from messenger apps, internet search tools, etc., means for a generative AI model to reproduce the user's memories and personality using these as training data, means for generating an avatar of the user on the cloud, means for collecting the user's past messages and dialogue history and storing it in a database, means for preprocessing the collected data to clean the text, tokenize it, and extract important phrases and patterns, means for training the generative AI model using the preprocessed data, means for generating an avatar of the user using the trained generative AI model, means for receiving messages from acquaintances and family, inputting them as prompts into the generative AI model, generating appropriate replies, and means for sending the generated replies to acquaintances and family. This makes it possible to accurately imitate the user's speech patterns and thought patterns and realize natural dialogue.
[0837] A "messenger app" is software that allows users to send and receive text messages, voice messages, images, videos, and other content.
[0838] An "internet search tool" is software that allows users to search for and retrieve information from the internet.
[0839] "Life log data" refers to data related to a user's daily life, including activity history, location information, and health data.
[0840] "Feature data" refers to data that represents the individual characteristics of a user, and includes things like writing style and voiceprints.
[0841] A "generative AI model" is a model that uses artificial intelligence technology to imitate a user's speech patterns and thought processes.
[0842] "Cloud" refers to a collection of computer resources and services provided via the internet.
[0843] An "avatar" is a virtual entity that mimics the user's speech patterns and thought processes, and engages in dialogue on their behalf.
[0844] A "database" is a system for efficiently storing, searching, and managing data.
[0845] "Preprocessing" refers to the process of converting data into a format suitable for analysis and model training.
[0846] "Tokenization" is the process of dividing text data into smaller units such as words and phrases.
[0847] A "prompt" is text input to a generative AI model that contains instructions for the model to generate an appropriate response.
[0848] This invention is a system for maintaining communication with acquaintances, friends, and family even when the user is absent, busy, or has passed away. This system links life log data and characteristic data such as text and voiceprints from messenger apps, internet search tools, etc., and uses this as training data for a generated AI model to recreate the user's memories and personality. A specific embodiment of this system is described below.
[0849] 1. Collection of user data
[0850] The server collects the user's past messages and conversation history. This includes emails, chat logs, and social media posts. The server stores this data in a database. For example, the server accesses email accounts and downloads past emails. It also uses chat application APIs to retrieve the user's chat logs.
[0851] 2. Data preprocessing
[0852] The server preprocesses the collected data. This includes text cleaning, tokenization, and extraction of important phrases and patterns. Specifically, the server removes unnecessary HTML tags and special characters from the text and divides the text into words and phrases (tokenization). Furthermore, it extracts frequently occurring phrases and patterns and stores them in a database.
[0853] 3. Training the Generative AI Model
[0854] The server trains a generative AI model using preprocessed data. For example, OpenAI's GPT-4 is used here. The server inputs the preprocessed data into the generative AI model and trains it. Once trained, the model becomes capable of mimicking the user's speech patterns and thought processes.
[0855] 4. The creation of an incarnation
[0856] The server uses a trained generative AI model to generate an avatar of the user. This avatar mimics the user's speech patterns and thought processes, recreating how the user typically interacts. The generated avatar is stored in a database.
[0857] 5. Receiving messages and generating replies
[0858] The device receives messages from acquaintances and family. The device inputs these messages as prompts into a generation AI model, which then generates an appropriate reply. For example, if the device receives the message "What are your plans for tomorrow?" from an acquaintance, it inputs this message into the generation AI model as a prompt. An example of a prompt would be: "User name: Taro Yamada, User's previous message: 'I'm busy today, so please check my plans for tomorrow,' Prompt: 'Generate a reply to confirm Taro Yamada's plans for tomorrow when he is busy.'"
[0859] 6. Send a reply
[0860] The device sends the generated reply to acquaintances and family. For example, it receives a reply from the generating AI model, "Tomorrow's schedule is a meeting at 10am and a client meeting at 2pm," and sends this to an acquaintance.
[0861] In this way, the server and terminal work together to create a user avatar, making it possible to maintain communication with acquaintances, friends, and family even when the user is absent, busy, or has passed away.
[0862] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0863] Step 1: Collecting User Data
[0864] The server collects the user's past messages and conversation history. Inputs include emails, chat logs, and social media posts. The server stores this data in a database. Specifically, the server accesses email accounts and downloads past emails. It also uses chat application APIs to retrieve the user's chat logs. The output is the collected data stored in the database.
[0865] Step 2: Data preprocessing
[0866] The server preprocesses the collected data. The input includes the data collected in step 1. The server cleans, tokenizes, and extracts important phrases and patterns from the text. Specifically, the server removes unnecessary HTML tags and special characters from the text and divides the text into words and phrases (tokenization). Furthermore, it extracts frequently occurring phrases and patterns and stores them in the database. The output is the preprocessed data.
[0867] Step 3: Training the Generative AI Model
[0868] The server trains a generative AI model using preprocessed data. The input includes the data preprocessed in step 2. The server inputs the preprocessed data into the generative AI model and trains the model. Specifically, the server inputs data into the generative AI model (e.g., GPT-4) and adjusts the model's parameters. The output is a fully trained generative AI model.
[0869] Step 4: Creation of the Incarnation
[0870] The server generates an avatar of the user using a trained generative AI model. The input includes the generative AI model trained in step 3. The server inputs the user's characteristics into the generative AI model and generates an avatar. Specifically, the server inputs the user's speech patterns and thought patterns into the model and generates an avatar. As output, the generated avatar is saved in the database.
[0871] Step 5: Receiving messages and generating replies
[0872] The device receives messages from acquaintances and family. The input includes messages from acquaintances and family. The device inputs these messages as prompts into a generating AI model, which then generates an appropriate reply. Specifically, the device uses the messaging app's API to retrieve the received message and inputs it as a prompt into the generating AI model. The output is the reply from the generating AI model.
[0873] Step 6: Send a reply
[0874] The device sends the generated reply to acquaintances and family. The input includes the reply generated in step 5. Specifically, the device uses the messaging app's API to send the generated reply. The output is the reply sent to acquaintances and family.
[0875] In this way, the server and terminal work together to create a user avatar, making it possible to maintain communication with acquaintances, friends, and family even when the user is absent, busy, or has passed away.
[0876] (Application Example 2)
[0877] Next, we will describe application example 2 of form example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 as the "terminal".
[0878] In modern society, it is often difficult to provide customer support when users are absent or busy. Furthermore, after a user's death, there are limited means of maintaining communication with acquaintances, friends, and family. Additionally, while virtual stores require the ability to mimic users' speech patterns and thought processes, the technology to achieve this is lacking.
[0879] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for linking life log data and characteristic data such as text and voiceprints from messenger apps, internet search tools, etc., means for the AI to reproduce the user's memories and personality using these as training data, means for generating an avatar of the user on the cloud, means for the avatar of the user to handle customer interactions in a virtual store, means for generating answers to customer questions using a generated AI model, means for mimicking the user's speech patterns and thought patterns using prompt sentences, and means for installation on smartphones and head-mounted displays. This makes it possible to handle customer interactions even when the user is absent or busy, and to maintain communication with acquaintances, friends, and family even after the user's death. Furthermore, it enables customer interactions in a virtual store that mimic the user's speech patterns and thought patterns.
[0880] A "messenger app" is software that allows users to send and receive text messages, voice messages, images, videos, and other content.
[0881] An "internet search tool" is software used to search for information on the internet and provide it to users.
[0882] "Life log data" refers to data related to a user's daily life, including activity history, location information, and health data.
[0883] "Feature data" refers to data used to identify an individual, such as a user's written text or voiceprint.
[0884] "Training data" refers to data used to train machine learning models.
[0885] "AI" is an abbreviation for artificial intelligence, which is a technology in which computers imitate human intelligence to learn and reason.
[0886] "Cloud" refers to a collection of computer resources and services provided via the internet.
[0887] An "avatar" is a virtual entity created by mimicking the user's speech patterns and thought processes.
[0888] A "virtual store" is a virtual store that exists on the internet, where users can purchase goods and services.
[0889] "Customer service" refers to responding to and addressing customer questions and requests.
[0890] A "generative AI model" is an artificial intelligence model that generates text and speech by mimicking a user's speech patterns and thought processes.
[0891] A "prompt" is text input to a generative AI model, instructing it on how to respond.
[0892] A "smartphone" is a mobile phone that is capable of connecting to the internet and running applications.
[0893] A "head-mounted display" is a display device worn on the head that provides virtual reality and augmented reality experiences.
[0894] The system for implementing this invention collects user lifelog data and characteristic data, and generates a user avatar on the cloud based on this data. Specific embodiments of this system are described below.
[0895] First, the server collects user lifelog data and characteristic data from messenger apps, internet search tools, and other sources. Lifelog data includes the user's daily activity history, location information, and health data. Characteristic data includes data used to identify the individual, such as the user's text and voiceprint.
[0896] Next, the server uses this data as training data to train a generative AI model. This generative AI model is used to mimic the user's speech patterns and thought patterns. Specifically, it uses prompt sentences to mimic the user's speech patterns and thought patterns.
[0897] The server generates a virtual avatar of the user in the cloud. This avatar exists to maintain communication with acquaintances, friends, and family even when the user is absent, busy, or has passed away. The avatar can handle customer interactions in a virtual store. For example, if a customer asks, "What are the features of this product?", the avatar uses a generative AI model to generate an appropriate answer.
[0898] This system will be implemented as an application installed on smartphones and head-mounted displays. It can improve the customer experience by allowing an avatar to handle customer interactions even when the user is absent or busy.
[0899] As a concrete example, by using the following prompt, the avatar can mimic the user's speech patterns and thought processes to provide customer service.
[0900] Example of a prompt:
[0901] You are a virtual store clerk. Please answer the following questions by mimicking the user's language and thought patterns.
[0902] Question: {question}
[0903] In this way, even when the user is absent or busy, the avatar can handle customer interactions, and communication with acquaintances, friends, and family can be maintained even after the user's death. Furthermore, it becomes possible to provide customer service in virtual stores that mimics the user's speech patterns and thought processes.
[0904] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0905] Step 1:
[0906] The server collects user lifelog data and characteristic data from messenger apps and internet search tools. Inputs include user activity history, location information, health data, text, and voiceprints. This data is collected and stored in a database. The output is the collected lifelog data and characteristic data.
[0907] Step 2:
[0908] The server preprocesses the collected lifelog data and feature data. The data collected in step 1 is used as input. Data processing such as data cleaning, normalization, and feature extraction is performed to convert the data into a format suitable for training the generative AI model. The preprocessed data is obtained as output.
[0909] Step 3:
[0910] The server uses the pre-processed data as training data to train a generative AI model. The data pre-processed in step 2 is used as input. The generative AI model analyzes the data and extracts patterns to learn the user's language use and thought patterns. The trained generative AI model is obtained as output.
[0911] Step 4:
[0912] The server generates a user avatar in the cloud. The generative AI model trained in step 3 is used as input. Based on the generative AI model, a virtual entity is generated that mimics the user's speech patterns and thought patterns. The output is the avatar generated in the cloud.
[0913] Step 5:
[0914] The terminal runs an application that allows a user's avatar to interact with customers in a virtual store. The input includes the avatar generated in the cloud and questions from the customer. The terminal sends the customer's questions to the cloud and generates appropriate answers using a generative AI model. The output is the answer provided to the customer.
[0915] Step 6:
[0916] The server uses prompts to mimic the user's vocabulary and thought patterns. The input includes customer questions and prompts. The generative AI model uses the prompts to mimic the user's vocabulary and thought patterns and generates appropriate responses. The output is the mimicked response.
[0917] Step 7:
[0918] The device displays the customer's response through an application installed on a smartphone or head-mounted display. The input includes the response generated in step 6. The device displays the response to the customer, improving the customer experience. The output is the response displayed to the customer.
[0919] (Example 3)
[0920] Next, we will describe Embodiment 3 of Embodiment Example 3. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0921] In modern society, there is a need to digitize and permanently store and reproduce users' memories and personalities. However, conventional systems have difficulty accurately reproducing users' memories and personalities, and they lack mechanisms to improve the accuracy of reproduction through long-term use by the user. Therefore, a system is needed that can reproduce users' memories and personalities with high accuracy and further improve accuracy through long-term use.
[0922] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 3 is realized by the following means.
[0923] In this invention, the server includes means for linking life log data and characteristic data such as text and voiceprints from messenger apps, internet search tools, etc., means for the AI to reproduce the user's memories and personality using these as training data, means for generating an avatar of the user on the cloud, means for collecting data provided by the user and storing it in a database, means for preprocessing the collected data and converting it into a format suitable for the generative AI model, means for training the generative AI model using the preprocessed data, means for receiving prompt sentences from the user and generating a response using the generative AI model, and means for providing the generated response to the user. This makes it possible to reproduce the user's memories and personality with high accuracy and to further improve accuracy with long-term use.
[0924] A "messenger app" is software that allows users to send and receive text messages and voice messages.
[0925] An "internet search tool" is software that allows users to search for information on the internet.
[0926] "Life log data" refers to data related to a user's daily life, including behavioral history and activity records.
[0927] "Feature data" refers to data that indicates the individual characteristics of a user, and includes things like text and voiceprints.
[0928] "Training data" refers to data used to train a generative AI model, and is a dataset that has been assigned correct labels.
[0929] A "generative AI model" is an artificial intelligence model used to recreate a user's memories and personality.
[0930] "Cloud" refers to a collection of computer resources and services provided via the internet.
[0931] An "incarnation" is a digital entity that recreates the user's memories and personality.
[0932] A "database" is a system for systematically storing and managing collected data.
[0933] "Preprocessing" refers to the process of converting collected data into a format suitable for generating AI models.
[0934] A "prompt message" is a sentence of a question or request that a user enters into the system.
[0935] A "response" is the answer that a generative AI model generates in response to a prompt.
[0936] This invention is a system that digitizes a user's memories and personality and generates an avatar on the cloud. The system collects life log data and feature data from messenger apps, internet search tools, etc., and uses this as training data for the generated AI model. A specific embodiment of this system is described below.
[0937] First, users provide life log data such as text messages, voice data, and image data through messenger apps or internet search tools. The device sends this data to the server, and the server stores the received data in a database.
[0938] Next, the server preprocesses the collected data. Specifically, text data is tokenized using natural language processing (NLP) techniques, and audio data is converted to text using speech recognition software. Image data is used for feature extraction using image recognition techniques. Machine learning frameworks such as TensorFlow and PyTorch are used for these preprocessing steps.
[0939] The pre-processed data is used to train the generative AI model. The server uses this data to train the generative AI model, teaching it patterns to reproduce the user's memories and personality. As training progresses, the model's accuracy improves.
[0940] When a user inputs a question or request as a prompt to the system, the terminal sends this prompt to the server. The server uses a generative AI model to generate an appropriate response to the prompt. The generated response is sent to the terminal and provided to the user.
[0941] For example, if a user asks, "What's my favorite movie?", the server uses a generative AI model to search the user's past data for information about their favorite movie and generates a response such as, "Your favorite movie is Inception." Also, if a user asks, "Where was the last place I traveled to?", the server generates a response such as, "You went to Hawaii last summer."
[0942] This system employs a monthly subscription model, and the longer a user continues to subscribe, the more data is collected and the more the AI learns. This increases the accuracy of the avatar's reproduction, making it possible to more accurately reproduce the user's memories and personality. The flow of the specific processing in Example 3 will be explained using Figure 15.
[0943] Step 1:
[0944] Users provide lifelog data such as text messages, voice data, and image data through messenger apps and internet search tools. The device sends this data to the server. The input is the lifelog data from the user, and the output is the data sent to the server.
[0945] Step 2:
[0946] The server stores the received data in a database. Specifically, it stores text messages, audio data, and image data in corresponding database tables. The input is the data sent from the terminal, and the output is the data stored in the database.
[0947] Step 3:
[0948] The server preprocesses the collected data. Text data is tokenized using natural language processing (NLP) techniques, and important keywords are extracted. Audio data is converted to text using speech recognition software. Image data is used for feature extraction using image recognition techniques. The input is raw data stored in a database, and the output is preprocessed data.
[0949] Step 4:
[0950] The server trains a generative AI model using preprocessed data. Specifically, it uses machine learning frameworks such as TensorFlow and PyTorch to train the model to reproduce patterns that replicate the user's memories and personality. The input is preprocessed data, and the output is the trained generative AI model.
[0951] Step 5:
[0952] The user inputs questions or requests to the system as prompts. The terminal sends these prompts to the server. The input is the prompt from the user, and the output is the prompt sent to the server.
[0953] Step 6:
[0954] The server inputs the received prompt into a generative AI model and generates an appropriate response. The generative AI model generates the most appropriate answer based on the user's past data. The input is the prompt and the trained generative AI model, and the output is the generated response.
[0955] Step 7:
[0956] The server sends the generated response to the terminal. The terminal displays this response to the user. The input is the generated response, and the output is the response provided to the user.
[0957] As a concrete example, if a user asks, "What's my favorite movie?", the server uses a generative AI model to search the user's past data for information about their "favorite movie" and generates a response such as, "Your favorite movie is Inception." Similarly, if a user asks, "Where was the last place I traveled to?", the server generates a response such as, "You went to Hawaii last summer."
[0958] (Application Example 3)
[0959] Next, we will describe application example 3 of form example 3. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 as the "terminal".
[0960] Conventional systems had low accuracy in recreating user memories and personalities, making it difficult to recommend individually customized content. Furthermore, the lack of mechanisms to improve system accuracy through long-term user engagement meant that improving the user experience was a challenge.
[0961] The identification processing performed by the identification processing unit 290 of the data processing device 12 in Application Example 3 is realized by the following means. In this invention, the server includes means for linking life log data and characteristic data such as text and voiceprints from messenger apps, internet search tools, etc., means for the AI to reproduce the user's memories and personality using these as training data, means for generating an avatar of the person on the cloud, and means for collecting user data, generating an avatar using an AI model, and recommending individually customized content. This makes it possible to reproduce the user's memories and personality with high accuracy and recommend individually customized content.
[0962] A "messenger app" is software that allows users to send and receive text messages, voice messages, images, videos, and other content.
[0963] An "internet search tool" is software that allows users to search for information on the internet.
[0964] "Life log data" refers to data related to a user's daily life, including activity history, location information, and health data.
[0965] "Feature data" refers to data used to identify an individual, such as a user's written text or voiceprint.
[0966] "Training data" refers to data used to train machine learning models.
[0967] "AI" is an abbreviation for artificial intelligence, which is a technology in which computers imitate human intelligence.
[0968] "Recreating the user's memories and personality" means imitating the user's past experiences and character based on collected data.
[0969] "Cloud" refers to a collection of computer resources and services provided via the internet.
[0970] An "incarnation" is a virtual entity that recreates the user's memories and personality.
[0971] "Personally customized content" refers to information and entertainment specially selected based on the user's preferences and interests.
[0972] An "AI model" is a mathematical model that uses artificial intelligence algorithms to analyze data and perform predictions and classifications.
[0973] A "monthly subscription system" is a pricing model in which users can access a service by paying a fixed fee each month.
[0974] The system for implementing this invention is configured as follows: The server has means for linking life log data and characteristic data such as text and voiceprints from messenger apps, internet search tools, etc. This makes it possible to collect data about the user's daily life and data for identifying the individual.
[0975] Next, the server has the means for the AI to recreate the user's memories and personality using the collected data as training data. Specifically, it uses an artificial intelligence (AI) model to mimic the user's past experiences and personality based on the collected life log data and feature data. This AI model is implemented using the OpenAI API.
[0976] Furthermore, the server has the means to generate an avatar of the user in the cloud. The generated avatar is a virtual entity that reproduces the user's memories and personality, and is managed in the cloud.
[0977] Furthermore, the server has the means to collect user data, generate avatars using AI models, and recommend individually customized content. Specifically, it collects data such as the user's viewing history, preferences, and feedback, and uses this data to recommend the most suitable content to the user using an AI model. This recommendation provides information and entertainment specially selected based on the user's preferences and interests.
[0978] As users continue to pay for the service over a long period, the system collects more data, and the AI learns more effectively. This improves the accuracy of the avatar's reproduction, allowing it to more accurately recreate the user's memories and personality.
[0979] As a concrete example, consider a case where a user prefers "action movies" and "comedy movies," liked "movie1," and disliked "movie2." Based on this data, an avatar is generated, and an example of a prompt message recommending content is as follows.
[0980] Example of a prompt:
[0981] Based on the user's data: {"user_id": "user123", "viewing_history": ["movie1", "movie2"], "preferences": ["action", "comedy"], "feedback": ["liked movie1", "disliked movie2"]}, generate an avatar that recreates the user's memories and personality.
[0982] By inputting this prompt into the OpenAI API, an avatar is generated that replicates the user's memories and personality. Based on this generated avatar, it becomes possible to recommend content that is most suitable for the user.
[0983] The flow of the specific processing in Application Example 3 will be explained using Figure 16.
[0984] Step 1:
[0985] The server collects life log data and characteristic data from messenger apps and internet search tools. Specifically, it retrieves user text messages, voice messages, search history, location information, etc. The input is various user data, and the output is the collected life log data and characteristic data.
[0986] Step 2:
[0987] The server inputs the collected lifelog data and feature data into the AI model as training data. Specifically, it preprocesses this data and converts it into a format that the AI model can learn from. The input is the preprocessed data, and the output is the training data used to train the AI model.
[0988] Step 3:
[0989] The server uses an AI model to generate an avatar that replicates the user's memories and personality. Specifically, it uses the OpenAI API to generate prompt statements based on the user's data and inputs them into the AI model. The input is the prompt statements, and the output is the generated avatar.
[0990] Step 4:
[0991] The server stores the generated avatars in the cloud. Specifically, it uses a cloud storage service to securely store the avatar data. The input is the generated avatar, and the output is the avatar data stored in the cloud.
[0992] Step 5:
[0993] The server collects data such as the user's viewing history, preferences, and feedback, and uses an AI model to recommend individually customized content. Specifically, it generates prompt messages based on the user's data and inputs them into the AI model. The input is the user's viewing history and preference data, and the output is the recommended content.
[0994] Step 6:
[0995] Users receive and view or use recommended content. Specifically, they use smartphones or other devices to view and enjoy the recommended content. The input is the recommended content, and the output is the user's viewing and usage history.
[0996] Step 7:
[0997] The server collects more data and continues training the AI model as users continue to pay for the service over a long period. Specifically, it periodically collects user usage data through a monthly subscription system and uses it to retrain the AI model. The input is the continuously collected user data, and the output is the retrained AI model.
[0998] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0999] "Example of form 1"
[1000] One embodiment of the present invention is a system that incorporates an emotion engine. This system recognizes the user's emotions and adjusts the avatar's response accordingly. Specifically, when the user is feeling happy, the avatar speaks in a cheerful tone. Conversely, when the user is sad, the avatar speaks in a calm tone. In this way, the emotion engine grasps the user's emotional state in real time and adjusts the avatar's response accordingly.
[1001] "Example of form 2"
[1002] Furthermore, the emotion engine controls the avatar's actions according to the user's emotional state. For example,
[1003] When the user is angry, the avatar chooses to offer an apology. Conversely, when the user is happy, the avatar chooses to offer a blessing. In this way, the emotion engine understands the user's emotional state and controls the avatar's actions accordingly.
[1004] "Example of form 3"
[1005] Furthermore, the emotion engine controls the avatar's facial expressions according to the user's emotional state. For example, when the user is happy, the avatar will smile. Conversely, when the user is sad, the avatar will make a sad face. In this way, the emotion engine understands the user's emotional state and controls the avatar's facial expressions accordingly.
[1006] The following describes the processing flow for each example of the form.
[1007] "Example of form 1"
[1008] Step 1: The emotion engine understands the user's emotional state in real time.
[1009] Step 2: The emotional engine adjusts the avatar's response according to the emotional state it has identified.
[1010] Step 3: The avatar delivers a pre-tuned response to the user.
[1011] "Example of form 2"
[1012] Step 1: The emotion engine understands the user's emotional state in real time.
[1013] Step 2: The emotional engine controls the avatar's actions according to the emotional state it has identified.
[1014] Step 3: The avatar performs controlled actions towards the user.
[1015] "Example of form 3"
[1016] Step 1: The emotion engine understands the user's emotional state in real time.
[1017] Step 2: The emotional engine controls the avatar's facial expressions according to the emotional state it has identified.
[1018] Step 3: The avatar displays controlled facial expressions to the user.
[1019] (Example 1)
[1020] Next, we will describe Example 1 of Form Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[1021] In modern society, there is a need for systems that can accurately understand users' behavioral patterns and emotional states and engage in appropriate dialogue accordingly. However, conventional systems have struggled to recognize users' emotions in real time and generate appropriate responses. Furthermore, data collection and analysis to recreate users' memories and personalities have been insufficient, resulting in a failure to improve user satisfaction.
[1022] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[1023] In this invention, the server includes means for linking life log data and characteristic data such as text and voiceprints from messenger apps, internet search tools, etc., means for the AI to reproduce the user's memories and personality using these as training data, means for generating an avatar of the user on the cloud, means for recognizing the user's emotions in real time using an emotion engine, and means for adjusting the tone and dialogue content of the avatar according to the user's emotional state. This makes it possible to conduct dialogue in real time according to the user's emotional state and reproduce the user's memories and personality with high accuracy.
[1024] A "messenger app" is a software application that allows users to send and receive text messages, voice messages, images, videos, and other content.
[1025] An "internet search tool" is a software application that allows users to search for information on the internet.
[1026] "Life log data" refers to data related to a user's daily life, including behavioral history, location information, and activity records.
[1027] "Text" refers to text data entered by the user, including messages, comments, and posts.
[1028] A "voiceprint" is characteristic data extracted from a user's voice data, and includes voice characteristics that can be used to identify an individual.
[1029] "Feature data" refers to data that shows specific patterns or characteristics extracted from user behavior and statements.
[1030] "Training data" refers to a dataset used to train a machine learning model, and includes input data and corresponding correct answer data.
[1031] "AI" is an abbreviation for artificial intelligence, which is a technology in which computers imitate human intelligence to learn, reason, recognize, and make judgments.
[1032] "Recreating the user's memories and personality" means using collected data to recreate the user's past actions and statements, and to mimic the user's personality and preferences.
[1033] "Cloud" refers to a collection of computer resources and services provided via the internet, serving as infrastructure for data storage and processing.
[1034] An "avatar" is a virtual entity that interacts with the user on their behalf, and is generated based on the user's memories and personality.
[1035] An "emotion engine" is a software component that analyzes and recognizes emotions from a user's text or voice.
[1036] "Recognizing in real time" means instantly analyzing emotions in response to user input and providing the results.
[1037] "Tone" refers to the tone or atmosphere of voice or text in a conversation, and is an element used to convey emotions and intentions.
[1038] "Adjusting the dialogue content" means changing the content and expression of the messages delivered by the avatar according to the user's emotional state.
[1039] This invention is a system that collects life log data and characteristic data such as text and voiceprints from messenger apps and internet search tools that users use daily, and uses this data to reconstruct the user's memories and personality. Furthermore, by combining it with an emotion engine, it enables dialogue that responds to the user's emotional state.
[1040] Data collection
[1041] When users use messenger apps (e.g., WhatsApp) or internet search tools (e.g., Google Search), life log data (e.g., chat history, search history) and characteristic data such as text and voiceprints are collected from these applications. Scripts that retrieve data via APIs are used for data collection.
[1042] Data preprocessing
[1043] The server receives the collected data and filters out unnecessary information. For example, it removes unwanted advertising messages from chat history and extracts only the user's messages. The Python pandas library is used for this process.
[1044] Data analysis and feature extraction
[1045] The server analyzes the pre-processed data to extract user behavior patterns, hobbies, interests, personality traits, and other characteristics. For example, natural language processing (NLP) techniques are used to extract emotions and topics from user statements. Libraries such as NLTK and spaCy are used for this analysis.
[1046] AI model training
[1047] The server uses the extracted feature data as training data to train a generative AI model (e.g., GPT-4). This allows the model to acquire the ability to reproduce the user's memories and personality. Frameworks such as TensorFlow and PyTorch are used for training.
[1048] emotion recognition
[1049] The server uses an emotion engine (e.g., IBM Watson's Sentiment Analysis API) to recognize emotions in real time from the user's text and voice. For example, if a user sends the message "I had a great time today!", the emotion engine recognizes the emotion as "happy".
[1050] Adjusting the reaction of the avatar
[1051] The device adjusts the avatar's tone and dialogue based on information from the emotion engine. For example, if the user is happy, the avatar will respond in a cheerful tone, saying, "That's wonderful! What happened?" Conversely, if the user is sad, the avatar will respond in a calm tone, saying, "That must have been tough. Is there anything I can do to help?"
[1052] Specific examples and prompt statements
[1053] Specific example
[1054] If a user sends the message "I had so much fun today!" via a messenger app, the server receives this message and recognizes the emotion of "fun" through its emotion engine. Based on this information, the device's avatar responds in a cheerful tone, "That's wonderful! What happened?"
[1055] Example of a prompt
[1056] "Design a system that analyzes the emotions of users from messages sent via messenger apps and has their avatar respond in an appropriate tone. Specifically, if the user is happy, the response should be in a cheerful tone; if they are sad, the response should be in a calm tone."
[1057] In this way, a system is built that enables dialogue tailored to the user's emotional state.
[1058] The flow of the specific processing in Example 1 will be explained using Figure 17.
[1059] Step 1:
[1060] When users use messenger apps or internet search tools, life log data and characteristic data such as text and voiceprints are collected from these applications.
[1061] Input: User's chat history, search history, voice data
[1062] Output: Collected life log data, text data, voiceprint data
[1063] Specific operation: Execute a script that retrieves data via the API and saves it to the database.
[1064] Step 2:
[1065] The server receives the collected data and filters out unnecessary information.
[1066] Input: Collected life log data, text data, voiceprint data
[1067] Output: Preprocessed data (data with unnecessary information removed)
[1068] Specific actions: Use the Python pandas library to create a dataframe and delete unnecessary rows and columns.
[1069] Step 3:
[1070] The server analyzes the pre-processed data to extract user behavior patterns and characteristics such as hobbies, interests, and personality.
[1071] Input: Preprocessed data
[1072] Output: User characteristic data (behavioral patterns, hobbies, interests, personality, etc.)
[1073] Specific operation: Uses natural language processing (NLP) techniques to extract sentiment and topics from text data. Libraries such as NLTK and spaCy are used.
[1074] Step 4:
[1075] The server uses the extracted feature data as training data to train a generative AI model.
[1076] Input: User characteristic data
[1077] Output: Trained generative AI model
[1078] Specific operation: Train a generative AI model (e.g., GPT-4) using frameworks such as TensorFlow or PyTorch.
[1079] Step 5:
[1080] The server uses an emotion engine to recognize emotions in real time from the user's text and voice.
[1081] Input: User's text data, audio data
[1082] Output: Recognized emotion data
[1083] Specific operation: Call an emotion engine (e.g., IBM Watson's Sentiment Analysis API) to analyze emotions.
[1084] Step 6:
[1085] The device adjusts the avatar's tone and dialogue based on information from the emotion engine.
[1086] Input: Recognized emotion data
[1087] Output: Adjusted dialogue
[1088] Specific actions: Based on emotional data, the avatar generates dialogue and responds to the user. For example, if the user is happy, the avatar will respond in a cheerful tone, "That's wonderful! What happened?"
[1089] (Application Example 1)
[1090] Next, we will describe Application Example 1 of Form Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[1091] Traditional content distribution services have faced challenges in providing personalized content recommendations that adequately reflect users' emotional states and individual hobbies and interests. Furthermore, it has been difficult to understand users' emotional states in real time and provide appropriate content accordingly.
[1092] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means. In this invention, the server includes means for linking life log data and characteristic data such as text and voiceprints from messenger apps, internet search tools, etc., means for the AI to reproduce the user's memories and personality using these as training data, means for generating an avatar of the person on the cloud, means including an emotion engine that grasps the user's emotional state in real time, and means for recommending optimal content according to the user's emotional state. This makes it possible to provide personalized content recommendations that are tailored to the user's emotional state and individual hobbies and interests.
[1093] A "messenger app" is an application that allows users to send and receive text messages, voice messages, images, videos, and other content.
[1094] An "internet search tool" is software or a web service that allows users to search for information on the internet.
[1095] "Life log data" refers to records of a user's actions and activities in their daily life.
[1096] "Feature data" refers to data that represents individual characteristics of a user, such as their written text or voiceprint.
[1097] "Training data" refers to a dataset used to train machine learning models.
[1098] "AI" is an abbreviation for artificial intelligence, which is a technology in which computers imitate human intelligence.
[1099] "Means for recreating a user's memories and personality" refers to technologies that recreate a user's past behavior and personality based on collected data.
[1100] "Cloud" refers to a collection of computer resources and services provided via the internet.
[1101] An "avatar" is a digital avatar that functions as a representation of the user.
[1102] An "emotion engine" is a technology that recognizes a user's emotional state in real time and generates a response accordingly.
[1103] "Content recommendation" is a technology that suggests the most suitable content based on the user's tastes, interests, and emotional state.
[1104] The system for implementing this invention is configured as follows: First, it collects life log data and characteristic data such as text and voiceprints from messenger apps and internet search tools installed on the user's smartphone or personal computer. This data reflects the user's behavioral patterns, hobbies, interests, personality, etc.
[1105] The server uses collected lifelog data and feature data as training data to recreate the user's memories and personality using a generative AI model. This generative AI model is implemented using programming languages such as Python and runs on the cloud. A digital avatar (incarnation) of the user is generated on the cloud, enabling conversations that reflect the user's memories and personality.
[1106] Furthermore, it incorporates an emotion engine that understands the user's emotional state in real time. This emotion engine analyzes the user's text and voice messages to recognize their emotional state. For example, if a user types "I'm a little tired today," the emotion engine detects the user's fatigue.
[1107] The server recommends the most suitable content to the user based on the emotional state obtained from the emotion engine. Content recommendations are based on the user's behavior patterns and emotional state, suggesting movies, music, and other content that will help the user relax.
[1108] For example, if a user types "I'm a little tired today," the emotion engine detects the user's fatigue and recommends relaxing movies or music. This process is achieved by inputting the following prompt into the generating AI model:
[1109] "Analyze the user's emotional state and recommend relaxing content. The user's input is 'I'm a little tired today.'"
[1110] In this way, personalized content recommendations tailored to the user's emotional state and individual hobbies and interests become possible.
[1111] The flow of a specific process in Application Example 1 will be explained using Figure 18.
[1112] Step 1:
[1113] Users use messenger apps and internet search tools installed on their smartphones or personal computers to send everyday messages and perform searches. This generates life log data and characteristic data such as text and voiceprints. The input is the user's messages and search queries, and the output is life log data and characteristic data.
[1114] Step 2:
[1115] The device sends the generated lifelog data and feature data to the server. The input is lifelog data and feature data, and the output is the data sent to the server. Specifically, the operation involves data collection and transmission.
[1116] Step 3:
[1117] The server inputs the received lifelog data and feature data into a generated AI model as training data to recreate the user's memories and personality. The input is lifelog data and feature data, and the output is a model that recreates the user's memories and personality. Specifically, the process involves data preprocessing, model training, and model generation.
[1118] Step 4:
[1119] The server generates a digital avatar (incarnation) of the user in the cloud. The input is a model that reproduces the user's memories and personality, and the output is the digital avatar in the cloud. Specifically, the server generates the avatar and places it in the cloud.
[1120] Step 5:
[1121] When a user enters a message, the terminal sends that message to the server. The input is the user's message, and the output is the message sent to the server. Specifically, the action involves sending a message.
[1122] Step 6:
[1123] The server inputs received messages into the emotion engine, allowing it to understand the user's emotional state in real time. The input is the user's message, and the output is the user's emotional state. Specifically, the system performs message analysis and emotional state recognition.
[1124] Step 7:
[1125] The server recommends optimal content based on the emotional state obtained from the emotion engine. The input is the user's emotional state, and the output is the recommended content. Specifically, the process involves evaluating the emotional state and selecting content.
[1126] Step 8:
[1127] The server sends recommended content to the terminal and presents it to the user. The input is the recommended content, and the output is the content presented to the user. Specifically, the operation involves sending and displaying content.
[1128] In this way, personalized content recommendations tailored to the user's emotional state and individual hobbies and interests are realized.
[1129] (Example 2)
[1130] Next, we will describe Example 2 of Form Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[1131] In modern society, users are often too busy to maintain communication with acquaintances, friends, and family. Furthermore, there is a need for those left behind to continue communicating with a deceased person after their passing. However, conventional technology has struggled to accurately reproduce a user's speech patterns and thought processes, and to engage in appropriate dialogue tailored to their emotional state.
[1132] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[1133] In this invention, the server includes means for linking life log data and characteristic data such as text and voiceprints from messenger apps, internet search tools, etc., means for artificial intelligence to recreate the user's memories and personality using these as training data, means for generating an avatar of the user on the cloud, means for analyzing the user's emotional state, means for generating dialogue with the avatar based on the emotional analysis results, and means for executing the generated dialogue. This makes it possible for the user to maintain communication with acquaintances, friends, and family even when they are absent, busy, or have passed away.
[1134] A "messenger app" is software that allows users to send and receive text messages, voice messages, images, videos, and other content.
[1135] An "internet search tool" is software that allows users to search for information on the internet and obtain the necessary data.
[1136] "Life log data" refers to data related to a user's daily life, including activity history, location information, and health data.
[1137] "Feature data" refers to data that represents the individual characteristics of a user, including writing style and voiceprints.
[1138] "Artificial intelligence" is a technology in which computers imitate human intelligence and perform learning, reasoning, and self-correction.
[1139] "Cloud" refers to a collection of computer resources and services provided via the internet.
[1140] An "avatar" is a virtual entity that mimics the user's speech patterns and thought processes, and engages in dialogue on their behalf.
[1141] "Emotional state" refers to the state of a user's current emotions, including joy, anger, sadness, etc.
[1142] "Emotion analysis" is the process of reading and analyzing emotions from a user's text and voice data.
[1143] "Dialogue generation" is the process of generating appropriate responses based on the user's emotional state and past dialogue history.
[1144] "Dialogue execution" is the process of actually sending the generated dialogue on behalf of the user.
[1145] This invention is a system that generates a user's avatar on the cloud, enabling communication with acquaintances, friends, and family even when the user is absent, busy, or has passed away. Specific embodiments of this system are described below.
[1146] First, the server collects user data. Specifically, it obtains life log data and characteristic data such as text and voiceprints from messenger apps and internet search tools. This data is used as training data to learn the user's language use and thought patterns.
[1147] Next, the server analyzes the collected data and learns the user's language use and thought patterns. Specifically, it analyzes the data using natural language processing techniques and learns user characteristics using generative AI models (for example, Google Cloud Natural Language API or OpenAI's GPT-4).
[1148] Subsequently, the server generates a user avatar based on the learned model. This avatar mimics the user's speech patterns and thought processes, and interacts on the user's behalf. The avatar is located in the cloud and sends and receives messages on the user's behalf.
[1149] Furthermore, the server uses an emotion engine to analyze the user's emotional state. Specifically, it analyzes the user's text and voice data to understand their emotional state in real time. For example, it uses Microsoft Azure's Text Analytics for Sentiment Analysis to perform emotion analysis.
[1150] Based on the emotion analysis results, the server generates dialogue for the avatar. For example, if the user is angry, it will choose words of apology; if they are happy, it will choose words of congratulations. The generated dialogue will be appropriate to the user's emotional state.
[1151] Finally, the device executes the generated conversation and sends the message on behalf of the user. This allows the user to maintain communication with acquaintances, friends, and family even when they are away or busy.
[1152] For example, if a user is too busy to reply, their avatar will send a message on their behalf. For instance, if a user receives a message from a friend asking "What are your plans for this weekend?", the avatar will reply "I'm planning to spend this weekend with my family" based on the user's past conversation history. Also, if a user is angry, the emotion engine will detect that emotion and the avatar will choose an apology. For example, if a user sends an angry message saying "How did this happen!", the avatar will reply "I'm sorry, I'll take care of it right away."
[1153] Examples of prompt messages include the following:
[1154] "Based on the user's past message history, generate an appropriate reply to the following message: 'What are your plans for this weekend?'"
[1155] "Generate an appropriate response for when a user is angry. User message: 'How did this happen!'"
[1156] In this way, the server and terminal work together to generate a user avatar and use an emotion engine to engage in appropriate conversations that correspond to the user's emotional state.
[1157] The flow of the specific processing in Example 2 will be explained using Figure 19.
[1158] Step 1:
[1159] The server collects user data. Specifically, it obtains life log data and characteristic data such as text and voiceprints from messenger apps and internet search tools. User account information and access permissions are required as input, and the collected data is obtained as output. For example, email data is collected using the Gmail API.
[1160] Step 2:
[1161] The server analyzes the collected data and learns the user's language use and thought patterns. Specifically, it uses natural language processing techniques to analyze the data and employs generative AI models (e.g., Google Cloud Natural Language API or OpenAI's GPT-4). The collected data is required as input, and the output is a model that captures the user's characteristics. For example, it analyzes past message history to extract patterns in the user's language use.
[1162] Step 3:
[1163] The server generates a user avatar based on a trained model. This avatar mimics the user's speech patterns and thought processes, and interacts on their behalf. The trained model is required as input, and the avatar, deployed on the cloud, is generated as output. For example, a user avatar is created and deployed on the cloud.
[1164] Step 4:
[1165] The server analyzes the user's emotional state using an emotion engine. Specifically, it analyzes the user's text and voice data to understand their emotional state in real time. New messages and voice data from the user are required as input, and the emotion analysis results are obtained as output. For example, emotion analysis is performed using Microsoft Azure's Text Analytics for Sentiment Analysis.
[1166] Step 5:
[1167] The server generates dialogue for the avatar based on the sentiment analysis results. For example, if the user is angry, it will choose words of apology, and if the user is happy, it will choose words of congratulations. The input requires the sentiment analysis results and a model that captures the user's characteristics, and the output is the generated dialogue. For example, if the user sends an angry message such as "How did this happen!", the avatar will reply, "I'm sorry, I'll take care of it right away."
[1168] Step 6:
[1169] The device executes the generated dialogue and sends the message on behalf of the user. The generated dialogue is required as input, and the sent message is obtained as output. For example, if the user receives a message from a friend asking "What are your plans for this weekend?", the avatar will reply "I'm planning to spend this weekend with my family."
[1170] In this way, the server and terminal work together to generate a user avatar and use an emotion engine to engage in appropriate conversations that correspond to the user's emotional state.
[1171] (Application Example 2)
[1172] Next, we will describe application example 2 of form example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 as the "terminal".
[1173] In modern society, there are problems such as users being too busy to fully utilize content distribution services, or being unable to communicate appropriately when users are absent. There is also the challenge of providing appropriate responses tailored to users' emotional states. Furthermore, there is a lack of means to maintain communication with acquaintances, friends, and family after a user's death.
[1174] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[1175] This invention includes a server that provides means for linking life log data and characteristic data such as text and voiceprints from messenger apps, internet search tools, etc., means for an AI to recreate the user's memories and personality using these as training data, means for generating an avatar of the user on the cloud, means for recommending content based on the user's viewing history and ratings, means for posting comments that mimic the user's speech patterns and thought patterns, and means for controlling the avatar's actions according to the user's emotional state. As a result, even if the user is busy, they can make full use of the content distribution service, appropriate communication can be maintained even when the user is absent, and appropriate responses can be made according to the user's emotional state. Furthermore, even after the user's death, communication with acquaintances, friends, and family can be maintained.
[1176] A "messenger app" is software that allows users to send and receive text messages, voice messages, images, videos, and other content.
[1177] An "internet search tool" is software that allows users to search for information on the internet.
[1178] "Life log data" refers to data related to a user's daily life, including activity history, location information, and health data.
[1179] "Feature data" refers to data used to identify an individual, such as a user's written text or voiceprint.
[1180] "Training data" refers to data used to train machine learning models, and consists of pairs of input data and corresponding correct answer data.
[1181] "AI" is an abbreviation for artificial intelligence, which is a technology in which computers imitate human intelligence to learn and reason.
[1182] "Means of recreating a user's memories and personality" refers to technologies that mimic a user's thought patterns and language use based on their past actions and statements.
[1183] "A means of creating a person's avatar on the cloud" refers to a technology that creates a digital alter ego of a user on a server accessible via the internet.
[1184] "Viewing history" refers to a record of content that a user has viewed in the past.
[1185] "Ratings" refer to records of evaluations and feedback given by users to content they have viewed.
[1186] "A means of recommending content" is a technology that suggests new content that users might be interested in, based on their viewing history and ratings.
[1187] "Methods for posting comments by mimicking language use and thought patterns" refers to technologies that generate and post comments that a user is likely to make, based on their past statements.
[1188] "Means of controlling the avatar's actions according to emotional state" refers to technology that analyzes the user's emotions and adjusts the avatar's actions and statements based on the results.
[1189] The embodiments for carrying out this invention will now be described. First, the entire system is built on the cloud and operates in conjunction with the user's terminal (smartphone or head-mounted display). The system consists of the following main components.
[1190] 1. Data Collection
[1191] The server collects life log data and characteristic data (such as text and voiceprints) from messenger apps and internet search tools. This data includes the user's activity history and statements.
[1192] 2. Incarnation generation
[1193] The server uses the collected data as training data to train a generative AI model (e.g., GPT-4). This model is used to recreate the user's memories and personality, specifically mimicking the user's speech patterns and thought processes.
[1194] 3. Avatar on the Cloud
[1195] The server creates a virtual avatar of the user in the cloud. This avatar communicates on the user's behalf when the user is absent or busy. It also maintains communication with acquaintances, friends, and family even after the user's death.
[1196] 4. Content Recommendation
[1197] The server uses a generative AI model to recommend new content based on the user's viewing history and ratings. For example, it might suggest the next movie a user should watch based on the genre of movies they've recently watched.
[1198] Example of a prompt
[1199] "The user recently watched an action movie. Please recommend three more movies to watch next."
[1200] 5. Post a comment
[1201] The server mimics the user's speech patterns and thought processes, posting comments on their behalf when they are absent or busy. This helps maintain the user's presence.
[1202] Example of a prompt
[1203] "Generate comments that users would post after watching 'Inception'."
[1204] 6. Emotional Response
[1205] The server uses an emotion engine (e.g., Emotion API) to analyze the user's emotional state. Based on the analysis, it adjusts the avatar's actions and statements. For example, it posts encouraging words when the user is sad and congratulatory words when the user is happy.
[1206] In this way, even busy users can fully utilize content delivery services, and appropriate communication can be maintained even when the user is absent. Furthermore, it becomes possible to respond appropriately to the user's emotional state, and communication with acquaintances, friends, and family can be maintained even after the user's death.
[1207] The flow of a specific process in Application Example 2 will be explained using Figure 20.
[1208] Step 1:
[1209] The server collects life log data and characteristic data (such as text and voiceprints) from messenger apps and internet search tools. Input is the user's activity history and statements, and output is the storage of this data in the cloud. Specifically, it retrieves data via an API and stores it in an AWS S3 bucket.
[1210] Step 2:
[1211] The server uses the collected data as training data to train a generative AI model (e.g., GPT-4). The input is the collected lifelog data and feature data, and the output is a model designed to reproduce the user's memories and personality. Specifically, the process involves preprocessing the data and feeding it to the AI model for learning.
[1212] Step 3:
[1213] The server generates a user avatar in the cloud. The input is a trained AI model, and the output is a digital replica of the user. Specifically, it deploys the AI model to a server in the cloud and provides an API to generate the user avatar.
[1214] Step 4:
[1215] The server recommends new content using a generative AI model based on the user's viewing history and ratings. The input is the user's viewing history and ratings, and the output is the recommended content. Specifically, it analyzes the viewing history and ratings, inputs prompts into the AI model, and generates recommendation results.
[1216] Step 5:
[1217] The server mimics the user's speech patterns and thought processes to post comments on their behalf when they are absent or busy. The input is the user's past statements, and the output is a generated comment. Specifically, it analyzes past statements, inputs prompts into an AI model, and generates comments.
[1218] Step 6:
[1219] The server analyzes the user's emotional state using an emotion engine (e.g., the Emotion API). The input is the user's current emotional data, and the output is the analysis result. Specifically, it acquires emotional data in real time and performs analysis using the Emotion API.
[1220] Step 7:
[1221] The server adjusts the avatar's actions and statements based on the analysis results. The input is the analysis results of the emotion engine, and the output is the adjusted actions and statements of the avatar. Specifically, it inputs prompt sentences into the AI model based on the analysis results, and generates appropriate actions and statements.
[1222] (Example 3)
[1223] Next, we will describe Embodiment 3 of Embodiment Example 3. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[1224] Conventional systems had low accuracy in recreating users' memories and personalities, and struggled to respond in real time to users' emotional states. Furthermore, insufficient long-term data collection and analysis prevented them from accurately reflecting users' emotional states. As a result, the user experience was limited, and the usefulness of the system was reduced.
[1225] The identification processing performed by the identification processing unit 290 of the data processing device 12 in Example 3 is realized by the following means. In this invention, the server includes means for linking life log data and characteristic data such as text and voiceprints from messenger apps, internet search tools, etc., means for the AI to reproduce the user's memories and personality using these as training data, means for generating an avatar of the person on the cloud, means for detecting the user's emotional state in real time, means for controlling the avatar's facial expression according to the user's emotional state, and means for increasing the accuracy of reproduction the longer the monthly subscription is continued. This makes it possible to reproduce the user's memories and personality with high accuracy and to respond in real time according to the emotional state.
[1226] A "messenger app" is software that allows users to send and receive text messages, voice messages, images, videos, and other content in real time.
[1227] An "internet search tool" is software that allows users to search for information on the internet and obtain the necessary data.
[1228] "Life log data" refers to data that includes a user's daily life history and activity records.
[1229] "Feature data" refers to unique data used to identify an individual, such as a user's text or voiceprint.
[1230] "Training data" refers to pairs of known inputs and outputs used to train a machine learning model.
[1231] "AI" is an abbreviation for artificial intelligence, a technology in which computers imitate human intelligence to learn, reason, and self-correct.
[1232] "Cloud" refers to a collection of computer resources and services provided via the internet.
[1233] An "avatar" is a virtual entity that recreates the user's memories and personality, and acts as a proxy for the user, engaging in dialogue and other actions.
[1234] "Emotional state" refers to the state of a user's current emotions and mood.
[1235] "Real-time" refers to the immediate collection, analysis, and response of data.
[1236] A "monthly subscription" is a pricing system where users pay a fixed fee each month to use a service.
[1237] "Reproducibility" refers to the degree to which a system can accurately reproduce a user's memories and personality.
[1238] Modes for carrying out the invention
[1239] This invention is a system that accurately reproduces a user's memories and personality and responds in real time according to their emotional state. Specific embodiments of this system are described below.
[1240] 1. Program generation
[1241] The server collects user data and generates a program that recreates the user's memories and personality using a generative AI model. This program also includes an emotion engine that controls the avatar's facial expressions according to the user's emotional state.
[1242] 2. Program Processing Description
[1243] The server collects more data as users continue to pay monthly subscriptions. This data includes the user's behavioral history, emotional state, and conversation content. The collected data is analyzed using generative AI models (e.g., OpenAI's GPT-4) and used to reconstruct the user's memory and personality.
[1244] The device uses hardware such as a camera and microphone to detect the user's emotional state in real time. For example, the camera captures the user's facial expressions, and the microphone analyzes the tone of the user's voice. This data is sent to an emotion engine to determine the user's emotional state.
[1245] The emotion engine controls the avatar's facial expressions based on the user's emotional state. For example, when the user is happy, the emotion engine makes the avatar smile. Conversely, when the user is sad, the emotion engine makes the avatar make a sad expression.
[1246] 3. Specific Examples and Examples of Prompt Statements
[1247] As a concrete example, consider a case where a user continues to pay for a subscription for one year. If a user continues to pay for a year, the system collects a vast amount of data about the user's behavior history and emotional state. Based on this data, a generative AI model can reproduce the user's memory and personality with very high accuracy.
[1248] Example of a prompt:
[1249] Based on data from a user who has continued paying for a year, please generate a program to recreate the user's memories and personality. Additionally, configure the emotion engine so that the avatar smiles when the user is happy and makes a sad expression when the user is sad.
[1250] In this way, the server collects data and generates a program that recreates the user's memories and personality using a generative AI model. The terminal detects the user's emotional state, and the emotion engine controls the avatar's facial expressions. As the user continues to pay for the service over a long period, the system collects more data, and the accuracy of the avatar's recreation improves. The flow of specific processing in Example 3 will be explained using Figure 21.
[1251] Program processing flow
[1252] Step 1: Collecting user data
[1253] Users log into the system and engage in everyday activities and interactions. The device uses a camera and microphone to capture the user's facial expressions and voice tone in real time. It also collects text data and activity history entered by the user.
[1254] Input: User facial expression data, voice tone, text data, behavioral history
[1255] Output: Collected user data
[1256] Specific actions:
[1257] The device's camera captures the user's face and collects facial expression data.
[1258] The device's microphone records the user's voice and analyzes their tone and emotions.
[1259] The system collects text data entered by users in chat.
[1260] Step 2: Sending data to the server
[1261] The device periodically sends the collected data to the server. The data is encrypted and transferred securely.
[1262] Input: Collected user data
[1263] Output: Data sent to the server
[1264] Specific actions:
[1265] The device encrypts the data it collects and sends it to the server.
[1266] The server saves the received data to the database.
[1267] Step 3: Data analysis on the server
[1268] The server analyzes the received data to identify user behavior patterns and emotional states. Data mining techniques and machine learning algorithms are used for this analysis.
[1269] Input: Data sent to the server
[1270] Output: Analysis results (user behavior patterns, emotional state)
[1271] Specific actions:
[1272] The server retrieves data from the database and applies an analysis algorithm.
[1273] The server identifies the user's behavioral patterns and emotional state, and saves the results.
[1274] Step 4: Recreating the user's memories and personality using a generative AI model.
[1275] The server uses a generated AI model (e.g., OpenAI's GPT-4) based on the analysis results to recreate the user's memories and personality. The generated model then simulates interactions and actions with the user.
[1276] Input: Analysis results
[1277] Output: Reconstructed user memories and personality
[1278] Specific actions:
[1279] The server inputs the analysis results into an AI model that recreates the user's memories and personality.
[1280] The generated model simulates user interactions and actions.
[1281] Step 5: Detecting the user's emotional state using the emotion engine
[1282] The device uses an emotion engine to detect the user's emotional state in real time. The emotion engine analyzes data from the camera and microphone to determine the user's emotional state.
[1283] Input: Real-time data from cameras and microphones
[1284] Output: User's emotional state
[1285] Specific actions:
[1286] The device's emotion engine analyzes facial expression data from the camera to determine the user's emotional state.
[1287] The device's emotion engine analyzes the tone of voice from the microphone to determine the user's emotional state.
[1288] Step 6: Controlling the avatar's facial expressions
[1289] After the emotion engine determines the user's emotional state, the device controls the avatar's facial expressions. For example, when the user is happy, the avatar will smile. Conversely, when the user is sad, the avatar will make a sad face.
[1290] Input: User's emotional state
[1291] Output: The avatar's expression
[1292] Specific actions:
[1293] The device controls the avatar's facial expressions based on the results of the emotion engine's judgment.
[1294] The avatar creates appropriate facial expressions according to the user's emotional state.
[1295] (Application Example 3)
[1296] Next, we will describe application example 3 of form example 3. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 as the "terminal".
[1297] Traditional avatar systems focus on recreating the user's memories and personality, but they have limitations in providing real-time responses based on the user's emotional state and delivering personalized content. Furthermore, the lack of means to analyze the user's emotional state could potentially degrade the quality of the user experience.
[1298] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 3 is realized by the following means. In this invention, the server includes means for linking life log data and characteristic data such as text and voiceprints from messenger apps, internet search tools, etc., means for the AI to reproduce the user's memories and personality using these as training data, means for generating an avatar of the person on the cloud, means for analyzing the user's emotional state in real time, and means for providing personalized content according to the analyzed emotional state. This makes it possible to provide real-time responses and personalized content according to the user's emotional state.
[1299] A "messenger app" is software that allows users to send and receive text messages and voice messages.
[1300] An "internet search tool" is software used to search for information on the internet.
[1301] "Life log data" refers to data related to a user's daily life, including activity history and location information.
[1302] "Feature data" refers to data that indicates the individual characteristics of a user, and includes things like text and voiceprints.
[1303] "Training data" refers to data used to train machine learning models.
[1304] "AI" is an abbreviation for artificial intelligence, which is a technology in which computers imitate human intelligence.
[1305] "Recreating the user's memories and personality" means imitating the user's character and thought patterns based on their past actions and statements.
[1306] "Creating a virtual avatar of oneself on the cloud" means creating a virtual alter ego of a user on a server accessible via the internet.
[1307] "Analyzing the user's emotional state in real time" means instantly determining the user's emotions at any given moment based on their facial expressions, voice, and other factors.
[1308] "Providing personalized content" means offering the most suitable information and entertainment based on the user's individual preferences and emotional state.
[1309] The system for carrying out this invention uses the following hardware and software to provide personalized content that responds to the user's emotional state.
[1310] hardware
[1311] 1. Camera: Used to capture the user's face. This utilizes the camera built into the smartphone or head-mounted display.
[1312] 2. Server: Use a cloud server for data processing and storage.
[1313] software
[1314] 1. OpenCV: This is a library for face detection. It analyzes video captured from a camera to detect faces.
[1315] 2. Keras: A deep learning library for implementing emotion recognition models. It predicts a user's emotions from detected facial images.
[1316] 3. Requests: This is a library for communicating with content delivery servers. It is used to retrieve content that matches the user's emotional state.
[1317] Data processing and data calculation
[1318] 1. Data Capture: Capture video from the camera in real time and perform face detection. Use OpenCV to detect faces, convert the detected faces to grayscale, and adjust their size.
[1319] 2. Emotion Prediction: The adjusted facial image is input into Keras's emotion recognition model to predict the user's emotion. There are seven types of emotions: anger, disgust, fear, happiness, sadness, surprise, and neutral expression.
[1320] 3. Content Delivery: Based on predicted sentiment, the Requests library is used to send a request to the content delivery server and retrieve appropriate content. The retrieved content is then displayed to the user.
[1321] Specific example
[1322] For example, if a user is using their smartphone and the camera captures their face and recognizes their emotion as "happy," the content delivery server might offer them a "comedy movie" or a "fun music playlist."
[1323] Example of a prompt
[1324] "Develop an application that analyzes user emotions in real time and provides personalized content based on those emotions. Use a Keras model for emotion recognition and the Requests library for communication with the content delivery server."
[1325] The flow of the specific processing in Application Example 3 will be explained using Figure 22.
[1326] Step 1:
[1327] The device captures the user's face in real time using its camera. The input is video data from the camera, and the output is the captured image frame. Specifically, the device activates the camera and continuously acquires video.
[1328] Step 2:
[1329] The device detects faces from captured image frames using OpenCV. The input is the captured image frame, and the output is the region of the detected face. Specifically, the device converts the image to grayscale and applies a face detection algorithm to determine the location of the face.
[1330] Step 3:
[1331] The device converts the detected face region to grayscale and resizes it to 48x48 pixels. The input is the detected face region, and the output is the resized face image. Specifically, the device normalizes the face image and converts it to a format suitable for machine learning models.
[1332] Step 4:
[1333] The device predicts emotions from resized face images using a Keras emotion recognition model. The input is a resized face image, and the output is the predicted emotion label. Specifically, the device inputs the face image into the model, calculates the probability distribution of emotions, and selects the emotion with the highest probability.
[1334] Step 5:
[1335] The device uses the Requests library to send a request to the content delivery server based on the predicted sentiment label. The input is the predicted sentiment label, and the output is the content retrieved from the content delivery server. Specifically, the device sends the sentiment label to the server in JSON format and requests appropriate content.
[1336] Step 6:
[1337] The server selects personalized content based on the received sentiment label and sends it back to the device. The input is the sentiment label, and the output is the selected content. Specifically, the server searches the database for content corresponding to the sentiment label and sends it to the device.
[1338] Step 7:
[1339] The terminal displays content received from the server to the user. The input is the content received from the server, and the output is the content displayed to the user. Specifically, the terminal displays the received content on the screen, making it viewable or usable by the user.
[1340] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[1341] The data generation model 58 is a form of so-called generative AI (Artificial Intelligence). One example of a data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1342] Other examples of generative AI include Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) are some examples.
[1343] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[1344] [Third Embodiment]
[1345] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[1346] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[1347] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1348] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[1349] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[1350] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[1351] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[1352] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[1353] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1354] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1355] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[1356] Next, the identification process performed by the identification processing unit 290 of the data processing device 12 will be described.
[1357] "Example of form 1"
[1358] One embodiment of the present invention involves collecting life log data and characteristic data such as text and voiceprints from messenger apps and internet search tools that users use daily. This data reflects the user's behavioral patterns, hobbies, interests, and personality. By using this data as training data to train an AI, the user's memories and personality can be reproduced.
[1359] "Example of form 2"
[1360] Next, a virtual avatar of the user is created in the cloud. This avatar is there to maintain communication with acquaintances, friends, and family even when the user is absent, busy, or after the user's death. Specifically, the avatar mimics the user's speech patterns and thought processes, recreating how the user normally interacts with others.
[1361] "Example of form 3"
[1362] Furthermore, the system of this invention employs a monthly subscription model. The longer a user continues to subscribe, the more data the system collects, and the more the AI learns. This increases the accuracy of the avatar's reproduction, making it possible to more accurately reproduce the user's memories and personality.
[1363] The following describes the processing flow for each example of the form.
[1364] "Example of form 1"
[1365] Step 1: Collect life log data and characteristic data such as text and voiceprints from messenger apps and internet search tools that the user uses on a daily basis.
[1366] Step 2: The collected data is used as training data to train the AI. Through this training, the AI understands the user's behavior patterns, hobbies, interests, personality, etc.
[1367] Step 3: As the AI learns, its ability to recreate the user's memories and personality improves.
[1368] "Example of form 2"
[1369] Step 1: Create a virtual avatar of the user in the cloud. This avatar will maintain communication with acquaintances, friends, and family even when the user is absent, busy, or has passed away.
[1370] Step 2: The avatar mimics the user's speech patterns and thought processes, recreating how the user typically interacts with others.
[1371] "Example of form 3"
[1372] Step 1: The system of the present invention employs a monthly subscription model.
[1373] Step 2: The longer a user continues to pay, the more data the system collects and the more the AI learns.
[1374] Step 3: This increases the accuracy of the avatar's reproduction, making it possible to more accurately recreate the user's memories and personality.
[1375] (Example 1)
[1376] Next, we will describe Embodiment 1 of Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[1377] In modern society, there is a need for technology that collects data reflecting users' behavioral patterns, interests, and personalities, and uses that data to reconstruct their memories and personalities. However, existing systems do not integrate the steps of data collection, preprocessing, storage, learning, evaluation, and generation, making it difficult to efficiently reconstruct users' memories and personalities. Furthermore, there is a lack of means to maintain communication with acquaintances, friends, and family even when the user is absent or after death. In addition, a sustainable billing model to improve the accuracy of the reconstruction has not been established.
[1378] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[1379] In this invention, the server includes means for collecting life log data and characteristic data such as text and voiceprints from messenger apps, internet search tools, etc.; means for preprocessing the collected data and converting it into a format suitable for training an AI model; means for storing the preprocessed data in a database; means for inputting the stored data as training data into an AI model and performing training to reproduce the user's memories and personality; means for evaluating the AI model after training is complete and confirming its performance; means for saving models with good evaluations; and means for generating an avatar of the user on the cloud. This makes it possible to efficiently collect and process data that reflects the user's behavior patterns, interests, and personality, and to reproduce the user's memories and personality with high accuracy. Furthermore, it is possible to maintain communication with acquaintances, friends, and family even when the user is absent or after death, and the accuracy of reproduction can be improved through a continuous billing model.
[1380] A "messenger app" is a software application that allows users to send and receive text messages, voice messages, images, videos, and other content.
[1381] An "internet search tool" is a software application that allows users to search for information on the internet.
[1382] "Life log data" refers to data that records a user's actions and activities in their daily life.
[1383] "Feature data" refers to data that reflects an individual's characteristics, such as a user's written text or voiceprint.
[1384] "Preprocessing" refers to the process of converting collected data into a format suitable for training an AI model.
[1385] A "database" is a system for efficiently storing, managing, and retrieving data.
[1386] "Training data" refers to data with correct labels that is used to train an AI model.
[1387] An "AI model" is a mathematical model that uses artificial intelligence technology to analyze data and perform predictions and classifications.
[1388] "Learning" is the process by which an AI model discovers patterns and rules based on training data.
[1389] "Evaluation" is the process of checking the performance of an AI model after it has completed its training.
[1390] "Cloud" refers to a collection of computer resources and services provided via the internet.
[1391] An "incarnation" is a digital entity that recreates the user's memories and personality.
[1392] A "billing model" is a business model that collects fees for the use of a service.
[1393] Modes for carrying out the invention
[1394] This invention is a system that collects life log data and characteristic data such as text and voiceprints from messenger apps and internet search tools that users use on a daily basis, and uses this data to reconstruct the user's memories and personality. A specific embodiment of this system is described below.
[1395] Data collection
[1396] The server collects life log data and characteristic data such as text and voiceprints from messenger apps and internet search tools that users use on a daily basis. This collection uses common software such as messenger apps and internet search tools. The collected data reflects the user's behavioral patterns, hobbies, interests, personality, etc.
[1397] Data preprocessing
[1398] The server preprocesses the collected data and converts it into a format suitable for training the AI model. Specifically, this involves tokenizing text data and converting audio data into spectrograms. This is done using Python libraries such as NLTK and Librosa. For example, NLTK is used for tokenizing text data, and Librosa is used for converting audio data into spectrograms.
[1399] Data storage
[1400] The server stores the pre-processed data in a database. This uses a database system such as MySQL or MongoDB. Specifically, it establishes a database connection and inserts the pre-processed data into the appropriate tables or collections.
[1401] AI model training
[1402] The server inputs pre-processed data as training data into an AI model and performs training to reproduce the user's memories and personality. For this training, generative AI models such as GPT-4 or BERT are used. Specifically, the server performs model initialization, data batch processing, and training.
[1403] Model evaluation and saving
[1404] The server evaluates the trained AI models and verifies their performance. Evaluation metrics such as accuracy and recall are used. Models that perform well are saved to the file system or cloud storage.
[1405] Cloud-based avatar generation
[1406] The server uses a well-rated model to generate an avatar of the user in the cloud. This avatar is used to recreate the user's memories and personality, and to maintain communication with acquaintances, friends, and family.
[1407] Specific example
[1408] Example 1: Data collection from messenger apps
[1409] The server collects messages that users send and receive using common messenger apps. The collected messages reflect the user's conversation patterns and interests. By preprocessing this message data and training an AI model, the system can replicate the user's conversational style.
[1410] Specific example 2: Data collection from internet search tools
[1411] The server collects search queries that users make using common internet search tools. These collected search queries reflect the user's interests and preferences. By preprocessing this search data and training an AI model, the system can reproduce the user's interests.
[1412] Example of a prompt
[1413] "Create a program that collects messages sent by users through common messenger apps and trains an AI model to replicate the user's conversational style."
[1414] "Create a program that collects search queries that users make using common internet search tools and trains an AI model to replicate those users' interests."
[1415] The flow of the specific processing in Example 1 will be explained using Figure 11.
[1416] Step 1: Data Collection
[1417] The server collects life log data and characteristic data such as text and voiceprints from messenger apps and internet search tools that users use on a daily basis. It receives log data from messenger apps and internet search tools as input and generates the collected raw data as output. Specifically, the server uses APIs and scraping techniques to acquire data and collects it using a secure communication protocol (e.g., HTTPS).
[1418] Step 2: Data Preprocessing
[1419] The server preprocesses the collected data and converts it into a format suitable for training the AI model. It receives raw data as input and generates preprocessed data as output. Specifically, the server performs the following operations:
[1420] Text data tokenization: Use Python's NLTK library to split text into words and phrases.
[1421] Spectrogram conversion of audio data: Convert audio data to a frequency spectrum using the Librosa library in Python.
[1422] Step 3: Save data
[1423] The server saves the pre-processed data to the database. It receives pre-processed data as input and generates data stored in the database as output. Specifically, the server performs the following operations:
[1424] Establishing a database connection: Connect to MySQL or MongoDB.
[1425] Data insertion: Insert pre-processed data into the appropriate tables or collections.
[1426] Step 4: Training the AI model
[1427] The server inputs pre-processed data as training data into the AI model and performs training to reproduce the user's memories and personality. It receives pre-processed data stored in a database as input and generates a trained AI model as output. Specifically, the server performs the following operations:
[1428] Model initialization: Initialize generative AI models such as GPT-4 and BERT.
[1429] Data batch processing: Pre-processed data is divided into batches and input into the model.
[1430] Running the training: Input data into the model and run the training.
[1431] Step 5: Evaluate and save the model.
[1432] The server evaluates the trained AI model and verifies its performance. It receives the trained AI model and evaluation data as input, and generates the evaluation results and a saved model as output. Specifically, the server performs the following operations:
[1433] Model evaluation: Calculate metrics such as accuracy and recall.
[1434] Model saving: Save models with good performance to the file system or cloud storage.
[1435] Step 6: Creating an avatar in the cloud
[1436] The server generates an avatar of the user in the cloud using a well-rated model. It receives a stored AI model as input and generates an avatar in the cloud as output. Specifically, the server performs the following operations:
[1437] Model Deployment: Deploy the AI model to the cloud environment.
[1438] Avatar Generation: Use the deployed model to generate an avatar that replicates the user's memories and personality.
[1439] (Application Example 1)
[1440] Next, we will describe Application Example 1 of Form Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[1441] In modern society, there is a growing need to provide personalized content to users by leveraging their life log data and characteristic data. However, conventional systems have struggled to recreate users' memories and personalities and recommend personalized content. Furthermore, there has been a lack of means to maintain communication with acquaintances and family when users are absent, busy, or even after death. A new system is needed to solve these problems.
[1442] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[1443] This invention includes a server that provides means for linking life log data and characteristic data such as text and voiceprints from messenger apps, internet search tools, etc., means for an AI to recreate the user's memories and personality using these as training data, means for generating an avatar of the user on the cloud, and means for automatically recommending content based on the user's preferences and interests based on the user's life log data and characteristic data. This makes it possible to recreate the user's memories and personality and provide personalized content. Furthermore, it enables the user to maintain communication with acquaintances and family even when they are absent, busy, or even after death.
[1444] A "messenger app" is an application that allows users to send and receive text messages, voice messages, images, videos, and other content.
[1445] An "internet search tool" is software or a web service that allows users to search for information on the internet.
[1446] "Life log data" refers to records of a user's actions and activities in their daily life.
[1447] "Feature data" refers to data used to identify individual users, such as user text or voiceprints.
[1448] "Training data" refers to a dataset used for AI to learn.
[1449] "AI" stands for artificial intelligence, which is a technology in which machines imitate human intelligence.
[1450] "Recreating the user's memories and personality" means that the AI imitates the user's thought and behavioral patterns based on the user's past actions and statements.
[1451] "Creating a digital avatar of oneself on the cloud" means creating a digital alter ego of a user using cloud computing technology.
[1452] "Automatically recommending content" means that AI suggests appropriate information and entertainment based on the user's preferences and interests.
[1453] The system for implementing this invention collects user life log data and characteristic data, reconstructs the user's memories and personality based on this data, and recommends personalized content. Specific embodiments of this system are described below.
[1454] System Configuration
[1455] hardware
[1456] Server: A server used for data collection, processing, storage, and training and inference of AI models.
[1457] Device: A device used by a user, such as a smartphone or computer.
[1458] software
[1459] Messenger app: An application that allows users to send and receive text messages, voice messages, images, videos, and other content.
[1460] Internet search tool: Software or web service used by users to search for information on the internet.
[1461] AI Model: A generative AI model that recreates the user's memories and personality based on the user's life log data and characteristic data, and recommends personalized content.
[1462] Data collection and processing
[1463] The server collects user lifelog and characteristic data from messenger apps and internet search tools. This includes the user's message history, search history, and voice data. The collected data is preprocessed and then used as training data for AI models.
[1464] AI model training and inference
[1465] The server trains an AI model based on the collected data. Specifically, text data is digitized using TfidfVectorizer, and user profiles are created using clustering algorithms (e.g., KMeans). Audio data is converted to text using speech recognition technology and processed similarly.
[1466] Content Recommendation
[1467] The server uses a pre-trained AI model to automatically recommend content based on the user's preferences and interests. This includes entertainment content such as movies, music, and articles. Recommended content is notified to the user's device.
[1468] Specific example
[1469] For example, if user A frequently talks about "movies" on a messenger app, the server collects that data and uses it to train an AI model. If user A's search history includes many searches for "latest movie reviews" and "movie trailers," the server creates a profile of user A based on this data. As a result, user A is recommended the latest movie review articles and movie trailer videos.
[1470] Example of a prompt
[1471] Develop an application that recommends content tailored to users' interests and preferences based on data collected from their messenger apps and search history. Specifically, implement a function that analyzes themes users frequently discuss and keywords they search for, and then automatically recommends content such as movies, music, and articles based on that analysis.
[1472] In this way, it is possible to build a system that utilizes users' life log data and characteristic data to deliver personalized content.
[1473] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[1474] Step 1:
[1475] The server collects user lifelog and characteristic data from messenger apps and internet search tools. Specifically, it retrieves user message history, search history, voice data, etc. This data is collected using APIs and database queries.
[1476] Input: Raw data from messenger apps and search tools
[1477] Output: Collected lifelog data and feature data
[1478] Step 2:
[1479] The server preprocesses the collected data. Specifically, text data is digitized using TfidfVectorizer, and speech data is converted to text using speech recognition technology. This converts the data into a format suitable for the AI model.
[1480] Input: Collected life log data and feature data
[1481] Output: Preprocessed data
[1482] Step 3:
[1483] The server trains the AI model based on pre-processed data. Specifically, it creates user profiles using clustering algorithms (e.g., KMeans). This allows user behavior patterns and interests to be reflected in the model.
[1484] Input: Preprocessed data
[1485] Output: Trained AI model
[1486] Step 4:
[1487] The server uses a pre-trained AI model to automatically recommend content based on the user's preferences and interests. Specifically, it selects entertainment content such as movies, music, and articles based on the user's profile.
[1488] Input: Pre-trained AI model, user profile
[1489] Output: Recommended content
[1490] Step 5:
[1491] The server notifies the user's device of the recommended content. Specifically, it presents the content to the user using push notifications or in-app messages.
[1492] Input: Recommended content
[1493] Output: Content displayed on the user's device
[1494] In this way, a system is built that utilizes users' life log data and characteristic data to deliver personalized content.
[1495] (Example 2)
[1496] Next, we will describe Example 2 of the Form Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[1497] In modern society, there is a need to maintain communication with acquaintances, friends, and family even when users are absent, busy, or have passed away. However, conventional technologies have struggled to accurately mimic users' speech patterns and thought processes to achieve natural dialogue. Furthermore, there has been a lack of effective means to utilize users' past messages and dialogue history to train generative AI models. This has resulted in the problem of user avatars generating unnatural responses.
[1498] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[1499] In this invention, the server includes means for linking life log data and characteristic data such as text and voiceprints from messenger apps, internet search tools, etc., means for a generative AI model to reproduce the user's memories and personality using these as training data, means for generating an avatar of the user on the cloud, means for collecting the user's past messages and dialogue history and storing it in a database, means for preprocessing the collected data to clean the text, tokenize it, and extract important phrases and patterns, means for training the generative AI model using the preprocessed data, means for generating an avatar of the user using the trained generative AI model, means for receiving messages from acquaintances and family, inputting them as prompts into the generative AI model, generating appropriate replies, and means for sending the generated replies to acquaintances and family. This makes it possible to accurately imitate the user's speech patterns and thought patterns and realize natural dialogue.
[1500] A "messenger app" is software that allows users to send and receive text messages, voice messages, images, videos, and other content.
[1501] An "internet search tool" is software that allows users to search for and retrieve information from the internet.
[1502] "Life log data" refers to data related to a user's daily life, including activity history, location information, and health data.
[1503] "Feature data" refers to data that represents the individual characteristics of a user, and includes things like writing style and voiceprints.
[1504] A "generative AI model" is a model that uses artificial intelligence technology to imitate a user's speech patterns and thought processes.
[1505] "Cloud" refers to a collection of computer resources and services provided via the internet.
[1506] An "avatar" is a virtual entity that mimics the user's speech patterns and thought processes, and engages in dialogue on their behalf.
[1507] A "database" is a system for efficiently storing, searching, and managing data.
[1508] "Preprocessing" refers to the process of converting data into a format suitable for analysis and model training.
[1509] "Tokenization" is the process of dividing text data into smaller units such as words and phrases.
[1510] A "prompt" is text input to a generative AI model that contains instructions for the model to generate an appropriate response.
[1511] This invention is a system for maintaining communication with acquaintances, friends, and family even when the user is absent, busy, or has passed away. This system links life log data and characteristic data such as text and voiceprints from messenger apps, internet search tools, etc., and uses this as training data for a generated AI model to recreate the user's memories and personality. A specific embodiment of this system is described below.
[1512] 1. Collection of user data
[1513] The server collects the user's past messages and conversation history. This includes emails, chat logs, and social media posts. The server stores this data in a database. For example, the server accesses email accounts and downloads past emails. It also uses chat application APIs to retrieve the user's chat logs.
[1514] 2. Data preprocessing
[1515] The server preprocesses the collected data. This includes text cleaning, tokenization, and extraction of important phrases and patterns. Specifically, the server removes unnecessary HTML tags and special characters from the text and divides the text into words and phrases (tokenization). Furthermore, it extracts frequently occurring phrases and patterns and stores them in a database.
[1516] 3. Training the Generative AI Model
[1517] The server trains a generative AI model using preprocessed data. For example, OpenAI's GPT-4 is used here. The server inputs the preprocessed data into the generative AI model and trains it. Once trained, the model becomes capable of mimicking the user's speech patterns and thought processes.
[1518] 4. The creation of an incarnation
[1519] The server uses a trained generative AI model to generate an avatar of the user. This avatar mimics the user's speech patterns and thought processes, recreating how the user typically interacts. The generated avatar is stored in a database.
[1520] 5. Receiving messages and generating replies
[1521] The device receives messages from acquaintances and family. The device inputs these messages as prompts into a generation AI model, which then generates an appropriate reply. For example, if the device receives the message "What are your plans for tomorrow?" from an acquaintance, it inputs this message into the generation AI model as a prompt. An example of a prompt would be: "User name: Taro Yamada, User's previous message: 'I'm busy today, so please check my plans for tomorrow,' Prompt: 'Generate a reply to confirm Taro Yamada's plans for tomorrow when he is busy.'"
[1522] 6. Send a reply
[1523] The device sends the generated reply to acquaintances and family. For example, it receives a reply from the generating AI model, "Tomorrow's schedule is a meeting at 10am and a client meeting at 2pm," and sends this to an acquaintance.
[1524] In this way, the server and terminal work together to create a user avatar, making it possible to maintain communication with acquaintances, friends, and family even when the user is absent, busy, or has passed away.
[1525] The flow of the specific processing in Example 2 will be explained using Figure 13.
[1526] Step 1: Collecting User Data
[1527] The server collects the user's past messages and conversation history. Inputs include emails, chat logs, and social media posts. The server stores this data in a database. Specifically, the server accesses email accounts and downloads past emails. It also uses chat application APIs to retrieve the user's chat logs. The output is the collected data stored in the database.
[1528] Step 2: Data preprocessing
[1529] The server preprocesses the collected data. The input includes the data collected in step 1. The server cleans, tokenizes, and extracts important phrases and patterns from the text. Specifically, the server removes unnecessary HTML tags and special characters from the text and divides the text into words and phrases (tokenization). Furthermore, it extracts frequently occurring phrases and patterns and stores them in the database. The output is the preprocessed data.
[1530] Step 3: Training the Generative AI Model
[1531] The server trains a generative AI model using preprocessed data. The input includes the data preprocessed in step 2. The server inputs the preprocessed data into the generative AI model and trains the model. Specifically, the server inputs data into the generative AI model (e.g., GPT-4) and adjusts the model's parameters. The output is a fully trained generative AI model.
[1532] Step 4: Creation of the Incarnation
[1533] The server generates an avatar of the user using a trained generative AI model. The input includes the generative AI model trained in step 3. The server inputs the user's characteristics into the generative AI model and generates an avatar. Specifically, the server inputs the user's speech patterns and thought patterns into the model and generates an avatar. As output, the generated avatar is saved in the database.
[1534] Step 5: Receiving messages and generating replies
[1535] The device receives messages from acquaintances and family. The input includes messages from acquaintances and family. The device inputs these messages as prompts into a generating AI model, which then generates an appropriate reply. Specifically, the device uses the messaging app's API to retrieve the received message and inputs it as a prompt into the generating AI model. The output is the reply from the generating AI model.
[1536] Step 6: Send a reply
[1537] The device sends the generated reply to acquaintances and family. The input includes the reply generated in step 5. Specifically, the device uses the messaging app's API to send the generated reply. The output is the reply sent to acquaintances and family.
[1538] In this way, the server and terminal work together to create a user avatar, making it possible to maintain communication with acquaintances, friends, and family even when the user is absent, busy, or has passed away.
[1539] (Application Example 2)
[1540] Next, we will describe application example 2 of form example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[1541] In modern society, it is often difficult to provide customer support when users are absent or busy. Furthermore, after a user's death, there are limited means of maintaining communication with acquaintances, friends, and family. Additionally, while virtual stores require the ability to mimic users' speech patterns and thought processes, the technology to achieve this is lacking.
[1542] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for linking life log data and characteristic data such as text and voiceprints from messenger apps, internet search tools, etc., means for the AI to reproduce the user's memories and personality using these as training data, means for generating an avatar of the user on the cloud, means for the avatar of the user to handle customer interactions in a virtual store, means for generating answers to customer questions using a generated AI model, means for mimicking the user's speech patterns and thought patterns using prompt sentences, and means for installation on smartphones and head-mounted displays. This makes it possible to handle customer interactions even when the user is absent or busy, and to maintain communication with acquaintances, friends, and family even after the user's death. Furthermore, it enables customer interactions in a virtual store that mimic the user's speech patterns and thought patterns.
[1543] A "messenger app" is software that allows users to send and receive text messages, voice messages, images, videos, and other content.
[1544] An "internet search tool" is software used to search for information on the internet and provide it to users.
[1545] "Life log data" refers to data related to a user's daily life, including activity history, location information, and health data.
[1546] "Feature data" refers to data used to identify an individual, such as a user's written text or voiceprint.
[1547] "Training data" refers to data used to train machine learning models.
[1548] "AI" is an abbreviation for artificial intelligence, which is a technology in which computers imitate human intelligence to learn and reason.
[1549] "Cloud" refers to a collection of computer resources and services provided via the internet.
[1550] An "avatar" is a virtual entity created by mimicking the user's speech patterns and thought processes.
[1551] A "virtual store" is a virtual store that exists on the internet, where users can purchase goods and services.
[1552] "Customer service" refers to responding to and addressing customer questions and requests.
[1553] A "generative AI model" is an artificial intelligence model that generates text and speech by mimicking a user's speech patterns and thought processes.
[1554] A "prompt" is text input to a generative AI model, instructing it on how to respond.
[1555] A "smartphone" is a mobile phone that is capable of connecting to the internet and running applications.
[1556] A "head-mounted display" is a display device worn on the head that provides virtual reality and augmented reality experiences.
[1557] The system for implementing this invention collects user lifelog data and characteristic data, and generates a user avatar on the cloud based on this data. Specific embodiments of this system are described below.
[1558] First, the server collects user lifelog data and characteristic data from messenger apps, internet search tools, and other sources. Lifelog data includes the user's daily activity history, location information, and health data. Characteristic data includes data used to identify the individual, such as the user's text and voiceprint.
[1559] Next, the server uses this data as training data to train a generative AI model. This generative AI model is used to mimic the user's speech patterns and thought patterns. Specifically, it uses prompt sentences to mimic the user's speech patterns and thought patterns.
[1560] The server generates a virtual avatar of the user in the cloud. This avatar exists to maintain communication with acquaintances, friends, and family even when the user is absent, busy, or has passed away. The avatar can handle customer interactions in a virtual store. For example, if a customer asks, "What are the features of this product?", the avatar uses a generative AI model to generate an appropriate answer.
[1561] This system will be implemented as an application installed on smartphones and head-mounted displays. It can improve the customer experience by allowing an avatar to handle customer interactions even when the user is absent or busy.
[1562] As a concrete example, by using the following prompt, the avatar can mimic the user's speech patterns and thought processes to provide customer service.
[1563] Example of a prompt:
[1564] You are a virtual store clerk. Please answer the following questions by mimicking the user's language and thought patterns.
[1565] Question: {question}
[1566] In this way, even when the user is absent or busy, the avatar can handle customer interactions, and communication with acquaintances, friends, and family can be maintained even after the user's death. Furthermore, it becomes possible to provide customer service in virtual stores that mimics the user's speech patterns and thought processes.
[1567] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[1568] Step 1:
[1569] The server collects user lifelog data and characteristic data from messenger apps and internet search tools. Inputs include user activity history, location information, health data, text, and voiceprints. This data is collected and stored in a database. The output is the collected lifelog data and characteristic data.
[1570] Step 2:
[1571] The server preprocesses the collected lifelog data and feature data. The data collected in step 1 is used as input. Data processing such as data cleaning, normalization, and feature extraction is performed to convert the data into a format suitable for training the generative AI model. The preprocessed data is obtained as output.
[1572] Step 3:
[1573] The server uses the pre-processed data as training data to train a generative AI model. The data pre-processed in step 2 is used as input. The generative AI model analyzes the data and extracts patterns to learn the user's language use and thought patterns. The trained generative AI model is obtained as output.
[1574] Step 4:
[1575] The server generates a user avatar in the cloud. The generative AI model trained in step 3 is used as input. Based on the generative AI model, a virtual entity is generated that mimics the user's speech patterns and thought patterns. The output is the avatar generated in the cloud.
[1576] Step 5:
[1577] The terminal runs an application that allows a user's avatar to interact with customers in a virtual store. The input includes the avatar generated in the cloud and questions from the customer. The terminal sends the customer's questions to the cloud and generates appropriate answers using a generative AI model. The output is the answer provided to the customer.
[1578] Step 6:
[1579] The server uses prompts to mimic the user's vocabulary and thought patterns. The input includes customer questions and prompts. The generative AI model uses the prompts to mimic the user's vocabulary and thought patterns and generates appropriate responses. The output is the mimicked response.
[1580] Step 7:
[1581] The device displays the customer's response through an application installed on a smartphone or head-mounted display. The input includes the response generated in step 6. The device displays the response to the customer, improving the customer experience. The output is the response displayed to the customer.
[1582] (Example 3)
[1583] Next, we will describe Embodiment 3 of Embodiment Example 3. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[1584] In modern society, there is a need to digitize and permanently store and reproduce users' memories and personalities. However, conventional systems have difficulty accurately reproducing users' memories and personalities, and they lack mechanisms to improve the accuracy of reproduction through long-term use by the user. Therefore, a system is needed that can reproduce users' memories and personalities with high accuracy and further improve accuracy through long-term use.
[1585] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 3 is realized by the following means.
[1586] In this invention, the server includes means for linking life log data and characteristic data such as text and voiceprints from messenger apps, internet search tools, etc., means for the AI to reproduce the user's memories and personality using these as training data, means for generating an avatar of the user on the cloud, means for collecting data provided by the user and storing it in a database, means for preprocessing the collected data and converting it into a format suitable for the generative AI model, means for training the generative AI model using the preprocessed data, means for receiving prompt sentences from the user and generating a response using the generative AI model, and means for providing the generated response to the user. This makes it possible to reproduce the user's memories and personality with high accuracy and to further improve accuracy with long-term use.
[1587] A "messenger app" is software that allows users to send and receive text messages and voice messages.
[1588] An "internet search tool" is software that allows users to search for information on the internet.
[1589] "Life log data" refers to data related to a user's daily life, including behavioral history and activity records.
[1590] "Feature data" refers to data that indicates the individual characteristics of a user, and includes things like text and voiceprints.
[1591] "Training data" refers to data used to train a generative AI model, and is a dataset that has been assigned correct labels.
[1592] A "generative AI model" is an artificial intelligence model used to recreate a user's memories and personality.
[1593] "Cloud" refers to a collection of computer resources and services provided via the internet.
[1594] An "incarnation" is a digital entity that recreates the user's memories and personality.
[1595] A "database" is a system for systematically storing and managing collected data.
[1596] "Preprocessing" refers to the process of converting collected data into a format suitable for generating AI models.
[1597] A "prompt message" is a sentence of a question or request that a user enters into the system.
[1598] A "response" is the answer that a generative AI model generates in response to a prompt.
[1599] This invention is a system that digitizes a user's memories and personality and generates an avatar on the cloud. The system collects life log data and feature data from messenger apps, internet search tools, etc., and uses this as training data for the generated AI model. A specific embodiment of this system is described below.
[1600] First, users provide life log data such as text messages, voice data, and image data through messenger apps or internet search tools. The device sends this data to the server, and the server stores the received data in a database.
[1601] Next, the server preprocesses the collected data. Specifically, text data is tokenized using natural language processing (NLP) techniques, and audio data is converted to text using speech recognition software. Image data is used for feature extraction using image recognition techniques. Machine learning frameworks such as TensorFlow and PyTorch are used for these preprocessing steps.
[1602] The pre-processed data is used to train the generative AI model. The server uses this data to train the generative AI model, teaching it patterns to reproduce the user's memories and personality. As training progresses, the model's accuracy improves.
[1603] When a user inputs a question or request as a prompt to the system, the terminal sends this prompt to the server. The server uses a generative AI model to generate an appropriate response to the prompt. The generated response is sent to the terminal and provided to the user.
[1604] For example, if a user asks, "What's my favorite movie?", the server uses a generative AI model to search the user's past data for information about their favorite movie and generates a response such as, "Your favorite movie is Inception." Also, if a user asks, "Where was the last place I traveled to?", the server generates a response such as, "You went to Hawaii last summer."
[1605] This system employs a monthly subscription model, and the longer a user continues to subscribe, the more data is collected and the more the AI learns. This increases the accuracy of the avatar's reproduction, making it possible to more accurately reproduce the user's memories and personality. The flow of the specific processing in Example 3 will be explained using Figure 15.
[1606] Step 1:
[1607] Users provide lifelog data such as text messages, voice data, and image data through messenger apps and internet search tools. The device sends this data to the server. The input is the lifelog data from the user, and the output is the data sent to the server.
[1608] Step 2:
[1609] The server stores the received data in a database. Specifically, it stores text messages, audio data, and image data in corresponding database tables. The input is the data sent from the terminal, and the output is the data stored in the database.
[1610] Step 3:
[1611] The server preprocesses the collected data. Text data is tokenized using natural language processing (NLP) techniques, and important keywords are extracted. Audio data is converted to text using speech recognition software. Image data is used for feature extraction using image recognition techniques. The input is raw data stored in a database, and the output is preprocessed data.
[1612] Step 4:
[1613] The server trains a generative AI model using preprocessed data. Specifically, it uses machine learning frameworks such as TensorFlow and PyTorch to train the model to reproduce patterns that replicate the user's memories and personality. The input is preprocessed data, and the output is the trained generative AI model.
[1614] Step 5:
[1615] The user inputs questions or requests to the system as prompts. The terminal sends these prompts to the server. The input is the prompt from the user, and the output is the prompt sent to the server.
[1616] Step 6:
[1617] The server inputs the received prompt into a generative AI model and generates an appropriate response. The generative AI model generates the most appropriate answer based on the user's past data. The input is the prompt and the trained generative AI model, and the output is the generated response.
[1618] Step 7:
[1619] The server sends the generated response to the terminal. The terminal displays this response to the user. The input is the generated response, and the output is the response provided to the user.
[1620] As a concrete example, if a user asks, "What's my favorite movie?", the server uses a generative AI model to search the user's past data for information about their "favorite movie" and generates a response such as, "Your favorite movie is Inception." Similarly, if a user asks, "Where was the last place I traveled to?", the server generates a response such as, "You went to Hawaii last summer."
[1621] (Application Example 3)
[1622] Next, we will describe application example 3 of form example 3. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[1623] Conventional systems had low accuracy in recreating user memories and personalities, making it difficult to recommend individually customized content. Furthermore, the lack of mechanisms to improve system accuracy through long-term user engagement meant that improving the user experience was a challenge.
[1624] The identification processing performed by the identification processing unit 290 of the data processing device 12 in Application Example 3 is realized by the following means. In this invention, the server includes means for linking life log data and characteristic data such as text and voiceprints from messenger apps, internet search tools, etc., means for the AI to reproduce the user's memories and personality using these as training data, means for generating an avatar of the person on the cloud, and means for collecting user data, generating an avatar using an AI model, and recommending individually customized content. This makes it possible to reproduce the user's memories and personality with high accuracy and recommend individually customized content.
[1625] A "messenger app" is software that allows users to send and receive text messages, voice messages, images, videos, and other content.
[1626] An "internet search tool" is software that allows users to search for information on the internet.
[1627] "Life log data" refers to data related to a user's daily life, including activity history, location information, and health data.
[1628] "Feature data" refers to data used to identify an individual, such as a user's written text or voiceprint.
[1629] "Training data" refers to data used to train machine learning models.
[1630] "AI" is an abbreviation for artificial intelligence, which is a technology in which computers imitate human intelligence.
[1631] "Recreating the user's memories and personality" means imitating the user's past experiences and character based on collected data.
[1632] "Cloud" refers to a collection of computer resources and services provided via the internet.
[1633] An "incarnation" is a virtual entity that recreates the user's memories and personality.
[1634] "Personally customized content" refers to information and entertainment specially selected based on the user's preferences and interests.
[1635] An "AI model" is a mathematical model that uses artificial intelligence algorithms to analyze data and perform predictions and classifications.
[1636] A "monthly subscription system" is a pricing model in which users can access a service by paying a fixed fee each month.
[1637] The system for implementing this invention is configured as follows: The server has means for linking life log data and characteristic data such as text and voiceprints from messenger apps, internet search tools, etc. This makes it possible to collect data about the user's daily life and data for identifying the individual.
[1638] Next, the server has the means for the AI to recreate the user's memories and personality using the collected data as training data. Specifically, it uses an artificial intelligence (AI) model to mimic the user's past experiences and personality based on the collected life log data and feature data. This AI model is implemented using the OpenAI API.
[1639] Furthermore, the server has the means to generate an avatar of the user in the cloud. The generated avatar is a virtual entity that reproduces the user's memories and personality, and is managed in the cloud.
[1640] Furthermore, the server has the means to collect user data, generate avatars using AI models, and recommend individually customized content. Specifically, it collects data such as the user's viewing history, preferences, and feedback, and uses this data to recommend the most suitable content to the user using an AI model. This recommendation provides information and entertainment specially selected based on the user's preferences and interests.
[1641] As users continue to pay for the service over a long period, the system collects more data, and the AI learns more effectively. This improves the accuracy of the avatar's reproduction, allowing it to more accurately recreate the user's memories and personality.
[1642] As a concrete example, consider a case where a user prefers "action movies" and "comedy movies," liked "movie1," and disliked "movie2." Based on this data, an avatar is generated, and an example of a prompt message recommending content is as follows.
[1643] Example of a prompt:
[1644] Based on the user's data: {"user_id": "user123", "viewing_history": ["movie1", "movie2"], "preferences": ["action", "comedy"], "feedback": ["liked movie1", "disliked movie2"]}, generate an avatar that recreates the user's memories and personality.
[1645] By inputting this prompt into the OpenAI API, an avatar is generated that replicates the user's memories and personality. Based on this generated avatar, it becomes possible to recommend content that is most suitable for the user.
[1646] The flow of the specific processing in Application Example 3 will be explained using Figure 16.
[1647] Step 1:
[1648] The server collects life log data and characteristic data from messenger apps and internet search tools. Specifically, it retrieves user text messages, voice messages, search history, location information, etc. The input is various user data, and the output is the collected life log data and characteristic data.
[1649] Step 2:
[1650] The server inputs the collected lifelog data and feature data into the AI model as training data. Specifically, it preprocesses this data and converts it into a format that the AI model can learn from. The input is the preprocessed data, and the output is the training data used to train the AI model.
[1651] Step 3:
[1652] The server uses an AI model to generate an avatar that replicates the user's memories and personality. Specifically, it uses the OpenAI API to generate prompt statements based on the user's data and inputs them into the AI model. The input is the prompt statements, and the output is the generated avatar.
[1653] Step 4:
[1654] The server stores the generated avatars in the cloud. Specifically, it uses a cloud storage service to securely store the avatar data. The input is the generated avatar, and the output is the avatar data stored in the cloud.
[1655] Step 5:
[1656] The server collects data such as the user's viewing history, preferences, and feedback, and uses an AI model to recommend individually customized content. Specifically, it generates prompt messages based on the user's data and inputs them into the AI model. The input is the user's viewing history and preference data, and the output is the recommended content.
[1657] Step 6:
[1658] Users receive and view or use recommended content. Specifically, they use smartphones or other devices to view and enjoy the recommended content. The input is the recommended content, and the output is the user's viewing and usage history.
[1659] Step 7:
[1660] The server collects more data and continues training the AI model as users continue to pay for the service over a long period. Specifically, it periodically collects user usage data through a monthly subscription system and uses it to retrain the AI model. The input is the continuously collected user data, and the output is the retrained AI model.
[1661] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[1662] "Example of form 1"
[1663] One embodiment of the present invention is a system that incorporates an emotion engine. This system recognizes the user's emotions and adjusts the avatar's response accordingly. Specifically, when the user is feeling happy, the avatar speaks in a cheerful tone. Conversely, when the user is sad, the avatar speaks in a calm tone. In this way, the emotion...
Claims
1. A data processing system including a processor and storage, The processor, in accordance with the program stored in the storage, A data acquisition means that collects user life log data from messenger apps and internet search tools, collects text related to the user, and further collects data representing the user's voiceprint, thereby obtaining characteristic data including the life log data and the text and voiceprint. An avatar providing means that, based on the life log data and characteristic data, uses a generating AI to learn the user's behavioral patterns, hobbies, interests, personality, language expression, and thought patterns, and provides an avatar that generates responses to messages from people other than the user when the user is absent or after death. A reproduction accuracy adjustment means adjusts the reproduction accuracy of the user by the avatar by adjusting the collection period of learning data in the learning of the avatar provision means according to the billing period of the user under the monthly subscription system, A response output means that outputs response data to a message input to a generative AI that has been trained based on the aforementioned feature data. A system configured to perform the following actions.
2. The system according to claim 1, wherein the data representing the voiceprint is collected via a terminal used by the user.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A