System
The system addresses the challenge of recreating deceased individuals for real-time conversations by generating AI models from uploaded data, allowing users to interact with them realistically and emotionally, incorporating current trends.
Patent Information
- Application Number
- JP2024137209
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-08-16
- Publication Date
- 2026-02-27
AI Technical Summary
Conventional technologies fail to realistically recreate the appearance, voice, and personality of deceased individuals for real-time conversations, lacking the ability to provide a meaningful interaction.
A system that allows users to upload information about the deceased, including photos, audio, and personality data, which generates appearance, voice, and personality models, enabling real-time dialogue through AI models that are continuously updated with trend data to maintain realism.
Enables users to have realistic and emotionally comforting conversations with deceased loved ones, ensuring the dialogue includes the latest information and maintains a sense of presence.
Smart Images

Figure 2026034088000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] When a loved one passes away, family and close friends often wish to spend their days feeling the presence of the deceased. However, with conventional technology, it has been difficult to realistically experience a reunion or conversation with the deceased. Due to current technological limitations, it is not possible to faithfully reproduce the appearance, voice, and personality of the deceased and have a real-time conversation. The objective of this invention is to realize a re-creation of a conversation with the deceased based on information about the deceased while they were alive, so that it feels as if they are actually there. [Means for solving the problem]
[0005] To solve the above problems, the present invention provides the following means. First, a means is provided for a user to upload information about the deceased to the system. This means includes the deceased's photograph, voice, hobbies, preferences, personality information, etc. The present invention also includes a means for generating an appearance, voice, and personality model of the deceased based on the uploaded information. This allows realistic video and audio of the deceased to be reproduced. The present invention also includes a means for generating a dialogue with the deceased in real time based on the generated model. This dialogue is transmitted to the user's terminal for display and playback. In addition, the present invention also includes a means for continuously collecting trend data based on the appearance, voice, and personality model of the deceased and updating the model based on that data. This ensures that the dialogue with the deceased always includes the latest information. Furthermore, a means is provided for analyzing voice data to extract voice characteristics and generating realistic voice using a voice synthesizer. These means enable dialogue with the deceased to be realized in a manner that is close to reality.
[0006] "Information about the deceased" refers to data such as photographs, audio recordings, hobbies, preferences, and personality information that the deceased owned while they were alive.
[0007] "Means for generating" refers to algorithms or programs that analyze the uploaded information about the deceased and construct an appearance model, voice model, and personality model of the deceased.
[0008] "Means for generating in real time" refers to an algorithm or program that generates a dialogue with a user in real time based on the generated model, allowing the user to experience the dialogue immediately.
[0009] A "user terminal" is a device used by a user to experience an interaction, and includes a smartphone, tablet, personal computer, etc.
[0010] "Trend data" refers to information about trends of the times, such as the latest news, fashions, technology, and cultural changes.
[0011] "Speech synthesizer" refers to a device or program that generates realistic speech based on the characteristics of analyzed speech data.
[0012] "Continuous collection means" refers to a mechanism for regularly obtaining the latest trend data from the Internet and continuously updating the AI model based on this data.
[0013] A "dialogue generation algorithm" refers to a set of calculation procedures or rules for generating appropriate responses based on user input. [Brief explanation of the drawings]
[0014] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11]FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0015] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0016] First, the terms used in the following description will be explained.
[0017] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0018] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0019] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0020] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0021] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0022] [First embodiment]
[0023] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0024] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0025] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0026] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0027] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0028] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0029] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0030] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0031] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0032] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0033] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0034] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0035] The embodiment of this invention is a system that allows users to upload information about the deceased, generates an AI model based on that information, and provides real-time dialogue with the deceased. This will be described in detail below with specific examples.
[0036] System Overview
[0037] The system provides an interface for users to upload information about the deceased, including photos, audio files, hobbies, preferences, and personality information. This information is sent to a server, which then generates appearance, voice, and personality models of the deceased.
[0038] Initial Setup
[0039] Users first log in to the system and provide information about the deceased. On the registration screen, users upload photos and audio files of the deceased, and enter text information about their hobbies, preferences, and personality.
[0040] The device sends the uploaded information to a server, which then stores the deceased person's information in a database.
[0041] Data analysis and model generation
[0042] The server extracts facial features from uploaded photos to generate facial recognition data, analyzes voice characteristics from audio files to generate a voice model, and uses natural language processing technology to analyze text information and create a model of the deceased's personality and speech patterns.
[0043] The server then builds a model of the deceased person's appearance, voice, and personality, and combines these to generate an AI model that serves as the basis for real-time interaction with the deceased person, with their appearance and voice appropriate to their current age.
[0044] Dialogue generation and execution
[0045] The user accesses the system using a terminal and selects the mode to start a conversation. When the request to start a conversation is sent to the server, the server generates video and audio of the deceased in real time based on the generated AI model, allowing the user to instantly experience a conversation with the deceased.
[0046] The device analyzes the user's voice input and converts it into text data, which is then sent to the server, which uses a dialogue generation algorithm to generate a response, which is then sent back to the device and played back as audio and video.
[0047] Continuous model updates
[0048] The server regularly collects current trend data and keeps the AI model up to date, ensuring that conversations with the deceased always contain the latest information and maintain realism.
[0049] Specific examples
[0050] For example, consider a user who wants to reunite with their deceased grandmother. The user uploads photos of the grandmother, recordings of her voice, and information about her favorite activities and hobbies. The server analyzes this information and creates a model of the grandmother's current appearance, voice, and personality.
[0051] Next, the user launches the app and selects "Interact with Grandmother." This causes the device to display the generated video and audio of the grandmother. When the user asks, "How was your day?", the server generates a response based on the grandmother's hobbies and past data, such as, "I was tending to the flowers in the garden today. It was a lot of fun." The user can enjoy the interaction as if the deceased were still alive.
[0052] In this way, the present invention is a system that enables users to have daily conversations with deceased loved ones who are important to them, and provides them with mental comfort.
[0053] The processing flow will be explained below.
[0054] Step 1:
[0055] Users log into the system and upload information about the deceased.
[0056] Users upload photos and audio files of the deceased on the registration screen.
[0057] The user inputs text information about the deceased's hobbies, preferences, and personality.
[0058] Step 2:
[0059] The terminal transmits the uploaded data to the server.
[0060] The device sends photos, audio files, and text information to the server.
[0061] Step 3:
[0062] The server stores the received data in a database.
[0063] The server stores the photo data in an image database.
[0064] The server stores the voice data in a voice database.
[0065] The server stores the text information in a text database.
[0066] Step 4:
[0067] The server generates face authentication data based on the photograph data.
[0068] The server uses image analysis algorithms to extract facial feature points.
[0069] The server generates a face recognition model based on the feature points.
[0070] Step 5:
[0071] The server analyzes the speech data to generate a speech model.
[0072] The server uses a voice analysis algorithm to extract voice characteristics.
[0073] The server configures the voice synthesizer based on the extracted features.
[0074] Step 6:
[0075] The server generates a personality model based on the text information.
[0076] The server analyzes the text information using natural language processing techniques.
[0077] The server models the deceased's speaking style and personality based on the analysis results.
[0078] Step 7:
[0079] The server integrates the generated appearance model, voice model, and personality model to generate an AI model.
[0080] The server combines the data from each model to create a single integrated model.
[0081] Step 8:
[0082] The user initiates the interaction using a dedicated application.
[0083] The user presses a button within the application to start an interaction.
[0084] The terminal sends a request to start a conversation to the server.
[0085] Step 9:
[0086] The server generates video and audio of the deceased in real time based on an AI model.
[0087] The server uses an image generation algorithm to generate an appearance of the deceased person according to their current age.
[0088] The server generates the audio using a voice synthesizer.
[0089] The server transmits the generated video and audio to the terminal.
[0090] Step 10:
[0091] The terminal plays back the received video and audio and presents them to the user.
[0092] The device uses playback software to play back the video and audio in sync.
[0093] Step 11:
[0094] The user speaks to the deceased person.
[0095] The user speaks about the specific content of the interaction (e.g., "How was your day?").
[0096] Step 12:
[0097] The terminal converts the user's voice input into text data.
[0098] The device uses voice recognition software to convert the speech into text.
[0099] The terminal transmits the text data to the server.
[0100] Step 13:
[0101] The server executes a dialogue generation algorithm based on the received text data.
[0102] The server parses the text data and generates an appropriate response.
[0103] The server converts the generated response into text data and audio data.
[0104] Step 14:
[0105] The server transmits the generated text data and voice data to the terminal.
[0106] Step 15:
[0107] The terminal reproduces audio and video based on the transmitted data and presents a response to the user.
[0108] The terminal uses playback software to play back the audio and video in sync.
[0109] Through these steps, users can have a conversation with the deceased and have a realistic experience.
[0110] Example 1
[0111] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0112] Current technology does not yet provide a method for communicating with the deceased in real time. Therefore, there is a need to realize a dialogue with the deceased and provide users with peace of mind. There is also a need to enable the deceased to have information based on current trends. There is a need for technology that can integrate a wide range of information, such as audio, video, and personality, to enable natural dialogue.
[0113] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0114] In this invention, the server comprises: a means for a user to upload information about the deceased;
[0115] A means for the server to receive the information uploaded by the user and store it in a database;
[0116] A means for the server to analyze the uploaded photo data using image processing technology and generate facial recognition data;
[0117] A means for the server to analyze the uploaded voice data using a voice analysis technology and generate a voice model;
[0118] A server analyzes the uploaded text data using natural language processing technology to generate a personality model;
[0119] A means of integrating the generated models to build an AI model;
[0120] means for transmitting a request to a server instructing a user to initiate a dialogue using a terminal;
[0121] A means for the server to generate video and audio of the deceased in real time based on the generated AI model;
[0122] A means for converting the user's voice into text data using voice recognition technology on the user's terminal and transmitting the text data to the server;
[0123] A means for the server to analyze the text data and generate a response using a dialogue generation algorithm;
[0124] means for transmitting the generated response to a user's terminal and playing it as audio and video;
[0125] A means for the server to continuously collect trend data based on the appearance model, voice model, and personality model of the deceased and update the AI model based on the collected trend data;
[0126] A means for the server to analyze the uploaded voice data and extract voice characteristics;
[0127] The system also includes a means for the server to generate realistic voice using speech synthesis technology based on the extracted voice characteristics. This allows the server to conduct natural dialogue in real time using information about the deceased, providing the user with peace of mind. Furthermore, continuous model updates ensure that the dialogue contains the latest information, maintaining greater realism.
[0128] "User" refers to the individual or entity who accesses the system, provides information about the deceased, and initiates the interaction.
[0129] "Server" refers to the computer system that receives, stores, analyzes information about the deceased sent by users and generates an AI model.
[0130] "Terminal" refers to a device through which a user accesses the system, initiates a dialogue, provides voice input, and plays back responses from the server.
[0131] A "database" refers to a system for systematically storing and managing information about the deceased (photographs, audio data, text information, etc.).
[0132] "Image processing technology" refers to the technology that analyzes uploaded photo data and generates facial recognition data.
[0133] "Voice analysis technology" refers to technology that analyzes uploaded voice data and generates a voice model.
[0134] "Natural language processing technology" refers to technology that analyzes uploaded text data and generates a personality model.
[0135] "AI model" refers to an artificial intelligence model that integrates the appearance, voice, and personality models of the deceased, allowing the deceased to interact in real time.
[0136] "Speech recognition technology" refers to technology that converts a user's voice into text data.
[0137] A "dialogue generation algorithm" refers to an algorithm for generating an appropriate response based on input text from a user.
[0138] "Speech synthesis technology" refers to the technology that generates realistic speech based on text data.
[0139] "Trend data" refers to the latest data that reflects current information, trends, and user preferences.
[0140] This invention is a system that allows users to upload information about the deceased, generates an AI model based on that information, and provides real-time dialogue with the deceased. Specific embodiments of this system are described in detail below.
[0141] Initial Setup
[0142] Users first log in to the system and provide information about the deceased. Through the registration screen, users upload or enter information such as:
[0143] Photos (JPEG, PNG, etc.)
[0144] Audio files (MP3, WAV, etc.)
[0145] Text information about hobbies, preferences, and personality
[0146] This information is sent from the user's terminal to the server.
[0147] Receiving and storing data
[0148] The server receives the deceased person's information sent from the device and stores it securely in a database, including:
[0149] Photo data
[0150] Audio data
[0151] Text data
[0152] Data analysis and model generation
[0153] The server uses the following specific software techniques to analyze the received data:
[0154] Image processing: Using image processing libraries such as OpenCV, facial features are extracted from photos and facial recognition data is generated.
[0155] Speech analysis: Using speech analysis tools such as Google® Cloud Speech-to-Text API and IBM Watson®, we extract speech characteristics from audio files and generate a speech model.
[0156] Text analysis: Using natural language processing libraries such as spaCy and GPT-3®, text information is analyzed to model the personality and speech patterns of the deceased.
[0157] By combining this data, the server generates an AI model that serves as the basis for real-time interaction with the deceased person, with their current age-appropriate appearance and voice.
[0158] Dialogue generation and execution
[0159] A user accesses the system using his / her terminal and selects a mode to start a conversation, which sends a request to start a conversation to the server.
[0160] In response to this request, the server begins generating video and audio of the deceased person in real time based on the generated AI model. Specifically, the process involves the following steps:
[0161] Video Generation: Using 3D modeling tools and animation software, real-time video of the deceased is generated.
[0162] Voice generation: Generate the voice of the deceased using text-to-speech technology (TTS, e.g., Google Cloud Text-to-Speech).
[0163] The device analyzes the user's voice input and converts it into text data using speech recognition technology (e.g., Google Cloud Speech-to-Text), which is then sent to the server.
[0164] The server analyzes the received user text data and generates a response using a dialogue generation algorithm (e.g., GPT-3). The generated response text is converted back into speech and sent to the device.
[0165] Continuous model updates
[0166] The server periodically collects user interaction history and current trend data to update the AI model. This ensures that conversations with the deceased always contain the latest information, maintaining realism. For example, the deceased can talk naturally about their current hobbies or the latest news.
[0167] Specific examples
[0168] For example, consider a user who wants to reunite with their deceased grandmother. The user uploads photos of the grandmother, recordings of her voice, and information about her favorite activities and hobbies. The server analyzes this information and creates a model of the grandmother's current appearance, voice, and personality.
[0169] Next, the user launches the app and selects "Interact with Grandmother." This causes the device to display the generated video and audio of the grandmother. When the user asks, "How was your day?", the server generates a response based on the grandmother's hobbies and past data: "I was tending to the flowers in the garden today. It was a lot of fun." Through this interaction, the user can enjoy the experience as if the deceased were still alive.
[0170] This system allows users to have daily conversations with deceased loved ones who are important to them, providing them with spiritual comfort.
[0171] Prompt Sentence Examples
[0172] For example, when generating a dialogue using GPT-3, you can enter a prompt like this:
[0173] "The user begins a conversation with their grandmother, who likes to spend time in her garden tending to her flowers. The user asks, 'How was your day?' The grandmother talks about her day's events."
[0174] By using this prompt, the AI can generate natural dialogue that meets the user's expectations.
[0175] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0176] Step 1:
[0177] Users log in to the system and provide information about the deceased. Specifically, they access a registration screen and upload photos (JPEG, PNG, etc.), audio files (MP3, WAV, etc.), and text information about the deceased's hobbies, preferences, and personality.
[0178] Input: Photos, audio files, and text information from the user
[0179] Output: Uploaded data
[0180] Step 2:
[0181] The terminal receives the information uploaded by the user and transmits it to the server.
[0182] Input: Photos, audio files, and text information from the user
[0183] Output: Data transferred to the server
[0184] Step 3:
[0185] The server receives the data sent from the terminal and stores it in a database.
[0186] Input: Uploaded information (photos, audio files, text information)
[0187] Output: Information stored in the database
[0188] Step 4:
[0189] The server analyzes the stored photo data using image processing technology (e.g., OpenCV) and generates facial recognition data.
[0190] Input: Photo data
[0191] Output: Facial recognition data
[0192] Specific operation: Use OpenCV to extract facial feature points and convert them into data.
[0193] Step 5:
[0194] The server analyzes the stored voice data using voice analysis technology (such as Google Cloud Speech-to-Text) and generates a voice model.
[0195] Input: Audio data
[0196] Output: Audio model
[0197] Specific operation: Analyzes audio files, extracts audio characteristics, and models them.
[0198] Step 6:
[0199] The server analyzes the stored text data using natural language processing technology (such as spaCy or GPT-3) to model the personality and speaking style of the deceased.
[0200] Input: Text data
[0201] Output: personality model
[0202] Specific behavior: Analyze text data and model writing style and word usage.
[0203] Step 7:
[0204] The server combines facial recognition data, voice models, and personality models to generate an AI model.
[0205] Input: Face recognition data, voice model, personality model
[0206] Output: AI model
[0207] Specific operation: Each model is integrated and generated as a single AI model.
[0208] Step 8:
[0209] The user sends a request to start a dialogue from the terminal to the server.
[0210] Input: Dialogue-initiating request
[0211] Output: Request sent to the server
[0212] Step 9:
[0213] When the server receives a request to start a dialogue, it generates video and audio of the deceased in real time based on the generated AI model.
[0214] Input: Dialogue start request, AI model
[0215] Output: Generated video and audio
[0216] How it works: Images are generated using 3D modeling tools and animation software, and audio is generated using TTS technology.
[0217] Step 10:
[0218] The device receives the user's voice input, converts it into text data using voice recognition technology (e.g., Google Cloud Speech-to-Text), and sends it to the server.
[0219] Input: User voice input
[0220] Output: Text data
[0221] Specific operation: Converts speech into text and sends it to the server.
[0222] Step 11:
[0223] The server analyzes the text data, generates a response using a dialogue generation algorithm (e.g., GPT-3), converts the response text back into speech, and sends it to the terminal.
[0224] Input: Text data
[0225] Output: Generated response (audio and text)
[0226] Specific operation: GPT-3 is used to analyze text data, generate responses, and convert them into speech.
[0227] Step 12:
[0228] The terminal receives the response sent from the server and plays it back as audio and video.
[0229] Input: Generated response (audio and video)
[0230] Output: Replayed response
[0231] Specific operation: Plays back received audio and video in real time.
[0232] Step 13:
[0233] The server periodically collects dialogue history and current trend data to update the AI model.
[0234] Input: Dialogue history, trend data
[0235] Output: Updated AI model
[0236] What it does: Rebuild and update the model based on new data.
[0237] The above are the specific processing steps and operations of the program for this system.
[0238] (Application example 1)
[0239] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0240] Conventional memorials and memorial services have only been able to preserve memories of the deceased in static forms such as photographs and videos, and have been unable to provide an interactive dialogue experience. Furthermore, there has been a lack of systems that allow people to find spiritual comfort through dialogue with the deceased. The present invention aims to enable real-time dialogue with the deceased, providing a deeper emotional experience, especially in memorial spaces in brick-and-mortar stores.
[0241] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0242] In this invention, the server includes a means for uploading information about the deceased, a means for analyzing the uploaded information to generate appearance, voice, and personality models, and a means for generating dialogue with the deceased in real time based on the generated models, thereby enabling an interactive dialogue experience to be provided on a device installed in the memorial space of a physical store.
[0243] "Means for uploading information about the deceased" refers to an interface that allows users to submit online photos, audio files, and text information about the deceased's personality and hobbies.
[0244] "Means of analyzing uploaded information to generate appearance, voice, and personality models" refers to technology in which the server extracts facial features from photographs of the deceased, analyzes voice characteristics from audio files, and models personality patterns from text information.
[0245] "Means for generating dialogue with the deceased in real time based on the generated model" refers to a technology that combines appearance, voice, and personality models generated by the server to generate natural dialogue in response to user input and provides it in real time.
[0246] "Means for displaying and playing on devices installed within the memorial space, including display and playback means for provision in physical stores" refers to technology that displays conversations with the deceased on dedicated devices within physical stores and plays them back as audio.
[0247] "Means for continuously collecting trend data and updating the model based on that data" refers to technology in which the server periodically retrieves the latest information and dynamically updates the deceased person model to improve its accuracy.
[0248] "Means for analyzing uploaded voice data and extracting voice characteristics" refers to technology that identifies unique voice patterns from audio files uploaded by users.
[0249] "Means for generating realistic voice using a voice synthesizer" refers to technology that generates artificial voice based on the characteristics of the deceased's voice.
[0250] "Means for providing an interactive dialogue experience in a memorial space within a physical store" refers to devices or systems installed to allow visitors to have real-time conversations with models of the deceased.
[0251] The embodiments for carrying out the present invention are as follows.
[0252] System Program
[0253] The system primarily consists of a user terminal, a server, and devices installed within the memorial space. The user first uploads information about the deceased, including photos, audio files, hobbies, preferences, and personality information. The server analyzes this information and generates appearance, voice, and personality models of the deceased. These models are integrated to create an AI model that enables real-time interaction with the deceased.
[0254] Processing Description
[0255] Uploading and initial settings from the user's device
[0256] Users first log in to the system through a dedicated interface and upload photos, audio files, and written information about the deceased. The user's device then sends this information to the server.
[0257] Server-based analysis and model generation
[0258] The server extracts facial and vocal features from uploaded photos and audio files. It also uses natural language processing technology to analyze the deceased's personality and dialogue patterns from text information. Specifically, it uses facial recognition software (e.g., OpenCV), voice analysis software (e.g., Google Speech-to-Text API), and natural language processing libraries (e.g., Hugging Face Transformers). Using these technologies, the server builds and integrates appearance, voice, and personality models of the deceased to generate an AI model.
[0259] Real-time dialogue generation and display
[0260] When a user begins a dialogue with the deceased using smart glasses or a head-mounted display installed in the memorial space, the terminal analyzes the user's voice input and converts it into text data. The server uses a dialogue generation algorithm based on this text data to generate a response, which is then sent back to the terminal. The terminal then plays back the response as audio and displays a video of the deceased. A video generation library (e.g., OpenCV) is used to render the video.
[0261] Continuous model updates
[0262] The server periodically collects the latest trend data and updates the deceased person model. This ensures that interactions with the deceased always reflect the latest information, maintaining realism. Trend data collection and analysis utilizes the latest databases and cloud computing technology.
[0263] Specific examples
[0264] For example, consider a case where a user uses the system to reminisce about a deceased family member. The user uploads photos of the deceased, audio recordings, and information about their hobbies and preferences to the system. The server analyzes this information and generates an AI model of the deceased.
[0265] Next, the user puts on the smart glasses installed in the memorial space and selects "Talk to Grandma." The system then generates video and audio of the deceased in real time. When the user asks, "How was your day?", the server generates a response based on the deceased's hobbies and past data, such as, "Today I was tending to the flowers in the garden. It was a lot of fun."
[0266] Prompt Sentence Examples
[0267] "Generate a dialogue between me and my grandmother, sharing memories from the past."
[0268] "Build a realistic dialogue system that can answer questions about deceased family members' hobbies and preferences."
[0269] In this way, the present invention provides users with a very realistic interactive experience with the deceased, creating new value in the memorial space of physical stores.
[0270] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0271] Step 1:
[0272] User upload of information about deceased persons
[0273] Input: Photos of the deceased, audio files, and text information about their hobbies, preferences, and personality.
[0274] How it works: A user uploads information related to the deceased through an interface.
[0275] Output: Uploaded data is sent to the system.
[0276] Step 2:
[0277] Sending data by the device
[0278] Input: The deceased person's information uploaded in Step 1
[0279] Operation: The user terminal sends the collected data to the server.
[0280] Output: The data is received on the server and is ready for analysis.
[0281] Step 3:
[0282] Data analysis and model generation by the server
[0283] Input: Received photos, audio files, and text information of the deceased
[0284] Operation:
[0285] Extract facial features from photos using facial recognition software (e.g., OpenCV).
[0286] Analyze speech features using speech analysis software (e.g., Google Speech-to-Text API).
[0287] Use natural language processing libraries (e.g., Hugging Face Transformers) to model personality and speech patterns from text information.
[0288] Output: Appearance, voice, and personality models of the deceased are generated and integrated.
[0289] Step 4:
[0290] AI model generation by the server
[0291] Input: Appearance model, voice model, and personality model generated in Step 3
[0292] How it works: Each model of the deceased person is combined to generate an AI model capable of real-time interaction.
[0293] Output: The completed AI model is saved and used for future dialogue generation.
[0294] Step 5:
[0295] User initiated interaction
[0296] Input: Access to smart glasses or head-mounted displays installed within the memorial space
[0297] Action: The user puts on the device, launches the application and selects an interaction mode.
[0298] Output: A conversation initiation request is sent to the server.
[0299] Step 6:
[0300] Server-generated dialogue
[0301] Input: User voice input (real-time conversation)
[0302] Operation:
[0303] It uses speech recognition technology to convert the user's speech into text.
[0304] It uses natural language processing techniques to analyze the text and generate appropriate responses.
[0305] A speech synthesizer is used to convert the generated text response into speech.
[0306] Use an image generation library (e.g., OpenCV) to render a video that matches the video of the deceased.
[0307] Output: Audio and visual responses of the deceased model are generated.
[0308] Step 7:
[0309] Display and playback of terminal interactions
[0310] Input: Audio and video data generated in step 6
[0311] Operation:
[0312] The terminal plays the received audio data and displays the video data.
[0313] The deceased's response is output to the user in the memorial space.
[0314] Output: The user experiences an interactive dialogue with the deceased person.
[0315] Step 8:
[0316] Continuously updating the model with the server
[0317] Input: Latest trend data and interaction history
[0318] Operation:
[0319] The server periodically updates the deceased person model, incorporating new trend information and past interaction data.
[0320] Improve the accuracy of your model based on updated data.
[0321] Output: The AI model always reflects the latest information, maintaining the quality of real-time interactions.
[0322] Through the above steps, the present invention provides the user with a realistic interaction experience with the deceased.
[0323] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0324] The embodiment of this invention is a system that allows a user to upload information about a deceased person, generates an AI model based on that information, provides real-time dialogue with the deceased, and dynamically adjusts the content of the dialogue based on the user's emotions. This will be described in detail below with specific examples.
[0325] System Overview
[0326] The system provides an interface for users to upload information about the deceased, including photos, audio files, hobbies, preferences, and personality. This information is sent to a server, which then generates an appearance model, voice model, and personality model of the deceased. The system also incorporates an emotion engine that recognizes the user's emotions and dynamically changes the content of the dialogue.
[0327] Initial Setup
[0328] Users first log in to the system and provide information about the deceased. On the registration screen, users upload photos and audio files of the deceased, and enter text information about their hobbies, preferences, and personality.
[0329] The device sends the uploaded information to a server, which then stores the deceased person's information in a database.
[0330] Data analysis and model generation
[0331] The server extracts facial features from uploaded photos to generate facial recognition data, analyzes voice characteristics from audio files to generate a voice model, and uses natural language processing technology to analyze text information and create a model of the deceased's personality and speech patterns.
[0332] The server then builds a model of the deceased person's appearance, voice, and personality, and combines these to generate an AI model that serves as the basis for real-time interaction with the deceased person, with their appearance and voice appropriate to their current age.
[0333] Dialogue generation and execution
[0334] The user accesses the system using a terminal and selects the mode to start a conversation. When the request to start a conversation is sent to the server, the server generates video and audio of the deceased in real time based on the generated AI model, allowing the user to instantly experience a conversation with the deceased.
[0335] The device analyzes the user's voice input and converts it into text data, which is then sent to the server, which uses a dialogue generation algorithm to generate a response, which is then sent back to the device and played back as audio and video.
[0336] Emotion recognition and dialogue adjustment
[0337] The device analyzes the user's voice and video and recognizes the user's emotions using an emotion engine, which extracts emotions from the user's tone of voice and facial expressions.
[0338] The server dynamically adjusts the dialogue content based on the user's emotional data recognized by the emotion engine, thereby generating more appropriate responses to the user and improving the realism of the dialogue.
[0339] Continuous model updates
[0340] The server regularly collects current trend data and keeps the AI model up to date. This ensures that conversations with the deceased always contain the latest information, maintaining realism. The emotion engine also continually learns, improving the accuracy of user emotion recognition.
[0341] Specific examples
[0342] For example, consider a user who wants to reunite with their deceased grandmother. The user uploads photos of the grandmother, recordings of her voice, and information about her favorite activities and hobbies. The server analyzes this information and creates a model of the grandmother's current appearance, voice, and personality.
[0343] Next, the user launches the app and selects "Talk to Grandma." This causes the device to display the generated video and audio of the grandmother. When the user asks, "How was your day?", the server generates a response based on the grandmother's hobbies and past data, such as, "I was taking care of the flowers in the garden today. It was a lot of fun." Furthermore, the emotion engine recognizes the user's emotions, and if the user is excited, it adjusts the content of the conversation, such as, "Can I see those flowers?"
[0344] In this way, the present invention is a system that allows users to have daily conversations with deceased loved ones who are important to them, and provides emotionally sensitive responses, thereby achieving greater mental comfort.
[0345] The processing flow will be explained below.
[0346] Step 1:
[0347] Users log into the system and upload information about the deceased.
[0348] Users select and upload photos and audio files of the deceased on the registration screen.
[0349] The user inputs the deceased's hobbies, preferences, and personality information in text format.
[0350] Step 2:
[0351] The terminal transmits the uploaded data to the server.
[0352] The terminal transmits the photos, audio files, and text information to the server.
[0353] Step 3:
[0354] The server stores the received data in a database.
[0355] The server stores the photo data in an image database.
[0356] The server stores the voice data in a voice database.
[0357] The server stores the text information in a text database.
[0358] Step 4:
[0359] The server generates face authentication data based on the photograph data.
[0360] The server uses image analysis algorithms to extract facial feature points.
[0361] The server generates a face recognition model based on the feature points.
[0362] Step 5:
[0363] The server analyzes the speech data to generate a speech model.
[0364] The server uses a voice analysis algorithm to extract voice characteristics.
[0365] The server configures the voice synthesizer based on the extracted features.
[0366] Step 6:
[0367] The server generates a personality model based on the text information.
[0368] The server analyzes the text information using natural language processing techniques.
[0369] The server models the deceased's speaking style and personality based on the analysis results.
[0370] Step 7:
[0371] The server integrates the generated appearance model, voice model, and personality model to generate an AI model.
[0372] The server combines the data from each model to create an integrated model.
[0373] Step 8:
[0374] The user initiates the interaction using a dedicated application.
[0375] The user presses a button within the application to start an interaction.
[0376] The terminal sends a request to start a conversation to the server.
[0377] Step 9:
[0378] The server generates video and audio of the deceased in real time based on an AI model.
[0379] The server uses an image generation algorithm to generate an appearance of the deceased person according to their current age.
[0380] The server generates the audio using a voice synthesizer.
[0381] The server transmits the generated video and audio to the terminal.
[0382] Step 10:
[0383] The terminal plays back the received video and audio and presents them to the user.
[0384] The device uses playback software to play back the video and audio in sync.
[0385] Step 11:
[0386] The user speaks to the deceased person.
[0387] The user speaks about the specific content of the interaction (e.g., "How was your day?").
[0388] Step 12:
[0389] The terminal converts the user's voice input into text data.
[0390] The device uses voice recognition software to convert the speech into text.
[0391] The terminal transmits the text data to the server.
[0392] Step 13:
[0393] The server executes a dialogue generation algorithm based on the received text data.
[0394] The server parses the text data and generates an appropriate response.
[0395] The server converts the generated response into text data and audio data.
[0396] Step 14:
[0397] The server transmits the generated text data and voice data to the terminal.
[0398] Step 15:
[0399] The terminal reproduces audio and video based on the transmitted data and presents a response to the user.
[0400] The terminal uses playback software to play back the audio and video in sync.
[0401] Step 16:
[0402] The device analyzes the user's voice and video and recognizes the user's emotions using an emotion engine.
[0403] The device uses voice analysis and facial recognition to extract the user's emotions.
[0404] Step 17:
[0405] The server dynamically adjusts the content of the dialogue based on the user's emotion data recognized by the emotion engine.
[0406] The server adjusts the response to be brighter if the user's sentiment is positive.
[0407] The server adjusts its response to tone down the user's emotions if they are negative.
[0408] Through these steps, users can have a conversation with the deceased and have a realistic, emotional experience.
[0409] Example 2
[0410] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0411] In modern society, there is a demand for technology that can recreate memories of deceased loved ones. However, conventional methods have difficulty in providing a real-time conversational experience with the deceased, and they are also insufficient in adjusting the conversation to take the user's emotions into account. Therefore, there is a need for a system that provides real-time, emotionally sensitive conversations based on information about the deceased.
[0412] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0413] In this invention, the server includes means for uploading information about the deceased, means for analyzing the uploaded information to generate appearance, voice, and personality models, means for generating a real-time dialogue with the deceased based on the generated model, means for converting the user's voice into text data and generating a response using a dialogue generation algorithm, means for transmitting the generated dialogue to the user's terminal for display and playback, and means for recognizing the user's emotions and dynamically adjusting the content of the dialogue, thereby enabling a real-time dialogue experience with the deceased and providing appropriate responses that are in tune with the user's emotions.
[0414] "Information about the deceased" refers to digital data such as photographs, audio files, hobbies, preferences, and personality information of the deceased.
[0415] "Means for uploading" refers to the interface and technology that allows a user to enter information about a deceased person into the system and transmit it to the server.
[0416] "Means of analyzing and generating appearance, voice, and personality models" refers to technology that allows a server to use uploaded data to recreate the appearance, voice, and personality of the deceased as a digital model.
[0417] "Means for generating real-time interactions" refers to the functionality and algorithms that use the generated model to create interactions with the deceased person in real time.
[0418] "Means for converting into text data" refers to speech recognition technology for converting a user's voice input into text form.
[0419] "Dialogue generation algorithm" refers to an algorithm for generating appropriate responses based on input from a user.
[0420] "Means for displaying and playing" refers to the technology for displaying the response sent from the server as audio or video on the user's device.
[0421] "Means for recognizing emotions and dynamically adjusting dialogue content" refers to technology that analyzes a user's emotions and changes the dialogue content and responses to match those emotions.
[0422] "Trend data" refers to data that refers to current trends and the latest information.
[0423] "Speech synthesizer" refers to technology for converting text data into voice data and playing it back.
[0424] The embodiment of this invention is a system that allows a user to upload information about a deceased person, generates an AI model based on that information, provides real-time dialogue with the deceased, and dynamically adjusts the content of the dialogue based on the user's emotions. This will be described in detail below with specific examples.
[0425] System Overview
[0426] The system provides an interface for users to upload information about the deceased, including photos, audio files, hobbies, preferences, and personality. This information is sent to a server, which then generates an appearance model, voice model, and personality model of the deceased. The system also incorporates an emotion engine that recognizes the user's emotions and dynamically changes the content of the dialogue.
[0427] Initial Setup
[0428] The user first logs in to the system and provides information about the deceased. On the registration screen, the user uploads photos and audio files of the deceased, and enters text information about their hobbies, preferences, and personality. For example, the user might upload a photo of their grandmother working in the garden or an audio recording of their grandmother's life. After receiving this information, the device sends it to the server, which then stores the information in a database.
[0429] Data analysis and model generation
[0430] The server uses an image analysis engine to extract the facial features of the deceased from uploaded photos. For example, it uses OpenCV and facial recognition APIs to recognize features such as the position of the eyes and mouth and skin color. The server then uses a voice analysis engine to analyze the voice features of the uploaded audio file. For example, it uses voice recognition technology to extract specific patterns and voice qualities from the voice. It also uses a natural language processing engine (e.g., GPT-3) to model the personality and speaking patterns of the deceased from text information. The server combines the results of these analyses to create three models: an appearance model, a voice model, and a personality model of the deceased, and generates an AI model based on this information.
[0431] Dialogue generation and execution
[0432] To start a conversation, the user accesses the system and selects the corresponding mode. For example, they select "Dialogue with Grandma" from the app menu. The device sends this request to the server as an API request. The server uses the generated AI model to generate video and audio of the deceased person in real time to display to the user and sends them to the device. The device receives the user's voice input, converts it into text using speech recognition technology, and sends this text data to the server. The server uses a dialogue generation engine based on the text data to generate an appropriate response. For example, in response to the question, "How was your day?", the server generates a response such as, "I was taking care of the flowers in the garden today. It was a lot of fun." The device receives the response from the server, generates audio data using speech synthesis technology, and plays it back to the user as audio and video.
[0433] Emotion recognition and dialogue adjustment
[0434] The device analyzes the user's voice and video using an emotion engine (e.g., Amazon Rekognition or Microsoft® Azure® Emotion API) to extract the user's emotional data. The server then adjusts the content of the dialogue based on this emotional data. For example, if the user is excited, the server generates a response that reflects their emotion, such as, "Can I see that flower?"
[0435] Continuous model updates
[0436] The server periodically collects current trend data and updates the AI model, ensuring that conversations with the deceased reflect the latest information and maintain realism. The emotion engine also continuously learns, improving the accuracy of user emotion recognition.
[0437] Example operation
[0438] If a user wants to reunite with their deceased grandmother, they upload photos and audio files of the grandmother to the system, along with information about her hobbies and preferences, such as "she liked gardening" and "she was good at cooking." The server analyzes this data and generates an AI model of the grandmother. When the user selects "Interact with Grandmother," the device displays real-time video and audio of the grandmother, and the conversation begins. When the user asks, "How was your day?", the server generates a response based on the grandmother's hobbies, such as, "I was tending to the flowers in the garden today. It was so much fun." Furthermore, if the user is excited, the server dynamically adjusts the content of the conversation, such as, "Can I see those flowers?" This system strives to create a realistic conversation with a deceased person who is important to the user, providing responses that are in tune with their emotions.
[0439] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0440] Step 1: Upload your information
[0441] Users log into the system and upload photos, audio files, hobbies, preferences, and personality information about the deceased.
[0442] Input: photos, audio files, hobbies, preferences, personality information
[0443] Output: Information of the deceased person sent to the server
[0444] Specific operation: The user selects various files from the interface and clicks the "Upload" button, and the device issues an API request to send this data to the server.
[0445] Step 2: Receiving and storing data
[0446] The terminal transmits the uploaded information to the server.
[0447] Input: User-uploaded information data about the deceased
[0448] Output: Information data of the deceased person stored on the server
[0449] Specific operation: The terminal splits the data received from the user and sends it to the server's API endpoint in the appropriate format. The server records the received data in the database and returns a confirmation response.
[0450] Step 3: Appearance model generation
[0451] The server analyzes the received photo and extracts facial features.
[0452] Input: Photo data
[0453] Output: Appearance model including facial recognition data
[0454] Specific operation: The server uses an image analysis engine (e.g., OpenCV) to extract features such as eyes, mouth, and skin color from the photo. Based on these features, it generates an appearance model.
[0455] Step 4: Generate a voice model
[0456] The server analyzes the uploaded audio file and extracts audio characteristics.
[0457] Input: Audio file
[0458] Output: Audio model
[0459] Specific operation: The server uses a voice analysis engine to analyze characteristics such as pitch, tone, and speed from the audio file, and generates a voice model based on this.
[0460] Step 5: Generate a personality model
[0461] The server analyzes the text information using natural language processing technology and models the personality and speech patterns of the deceased.
[0462] Input: Text data of hobbies, preferences, and personality information
[0463] Output: personality model
[0464] Specific operation: The server uses a natural language processing engine (e.g., GPT-3) to analyze text data and extract characteristic phrases and contexts, thereby generating a personality model.
[0465] Step 6: Integrating the AI model
[0466] The server integrates the appearance model, voice model, and personality model to generate an AI model.
[0467] Input: Appearance model, voice model, personality model
[0468] Output: Integrated AI model
[0469] Specific operation: The server combines the data from each model to generate a unified AI model with a consistent personality.
[0470] Step 7: Creating and Executing Real-Time Interactions
[0471] A user accesses the system using a terminal and selects an interaction mode.
[0472] Input: User's interaction initiation request
[0473] Output: Real-time dialogue display by an AI model of the deceased person
[0474] Specific operation: The device sends a request to the server, which then generates video and audio of the deceased in real time based on the generated AI model and sends them to the device. The server then receives the user's dialogue input, converts it into text using voice recognition technology, and generates a dialogue response.
[0475] Step 8: Emotion recognition and dialogue adjustment
[0476] The device analyzes the user's voice and video to recognize the user's emotions.
[0477] Input: User's audio and video data
[0478] Output: User emotion data
[0479] How it works: The device uses an emotion engine (e.g., Amazon Rekognition) to analyze the user's tone of voice and facial expressions to extract emotional data. The server then dynamically adjusts its response based on this data.
[0480] Step 9: Continuously updating the model
[0481] The server periodically collects trend data and updates the AI model.
[0482] Input: Latest trend data
[0483] Output: Updated AI model
[0484] How it works: The server collects the latest news and trend information from the internet and reflects it in the AI model. The emotion engine also continuously learns and improves the accuracy of user emotion recognition.
[0485] These are the processing steps of the program for this system. Each step works together to provide the user with a real-time conversational experience with the deceased, generating emotionally sensitive responses.
[0486] (Application example 2)
[0487] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0488] While existing systems exist that simulate conversations with deceased loved ones, they lack realism and are unable to provide conversations that are sensitive to the user's emotions. Furthermore, the deceased's model is fixed and not continuously updated, resulting in inconsistent conversation content. Furthermore, the system lacks the ability to properly analyze the user's voice input and generate dialogue responses, leaving a need for improved user experience.
[0489] The identification processing by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for uploading information about the deceased, means for analyzing the uploaded information to generate appearance, voice, and personality models, means for generating dialogue with the deceased in real time based on the generated models, means for recognizing the user's emotions and dynamically adjusting the dialogue content based on the emotions, and means for analyzing the user's voice input and generating appropriate dialogue responses based on the analysis results. This enables highly realistic dialogue that is in tune with the user's emotions, providing an experience that more vividly revives memories of the deceased.
[0490] The "means for uploading information about the deceased" refers to a means for providing an interface for a user to send information about the deceased, such as photos, audio, and text, to the system.
[0491] "Means for analyzing uploaded information to generate appearance, voice, and personality models" refers to means for generating digital models to reproduce the face, voice, and personality of the deceased based on uploaded information about the deceased.
[0492] "Means for generating dialogue with the deceased in real time based on the generated model" refers to means for generating dialogue in real time using a model of the appearance, voice, and personality of the deceased.
[0493] "Means for transmitting the generated dialogue to the user's device and displaying / playing it back" refers to means for transmitting the generated dialogue with the deceased to the user's device, such as a smartphone or PC, and playing it back as audio and video.
[0494] "Means for recognizing the user's emotions and dynamically adjusting the dialogue content based on those emotions" refers to means for detecting the user's emotions from the tone of voice and facial expression, and appropriately changing the dialogue content in accordance with those emotions.
[0495] The "means for analyzing uploaded voice data and extracting voice characteristics" refers to a means for extracting voice tones and speaking style characteristics from uploaded voice data.
[0496] "Means for generating realistic voice using a voice synthesizer" refers to means for generating realistic voice using voice synthesis technology based on extracted voice characteristics.
[0497] "Means for analyzing a user's voice input and generating an appropriate dialogue response based on the analysis results" refers to means for analyzing the content of what the user has said and generating an optimal response based on that content.
[0498] The system for implementing this invention is configured by combining the following means. The system allows the user to upload information about the deceased, generates an AI model based on that information, and engages in real-time dialogue with the deceased. It also has the ability to recognize the user's emotions and dynamically adjust the content of the dialogue. This allows the system to provide the user with a highly realistic dialogue experience.
[0499] System configuration
[0500] server:
[0501] Store information about the deceased in a database.
[0502] Photographs, audio, and text information of the deceased are analyzed to generate appearance, voice, and personality models.
[0503] The generated models are integrated to create a generative AI model.
[0504] To analyze a user's emotions and dynamically adjust dialogue content based on the emotions.
[0505] Real-time video and audio of the deceased are generated and transmitted to the user's device.
[0506] Device:
[0507] A user interface is provided and a means is provided for uploading information about the deceased.
[0508] It captures the user's audio and video inputs and sends them to a server for sentiment analysis.
[0509] Display and play back the dialogue sent from the server.
[0510] Implementation details
[0511] Users first log in to the system and upload photos, audio, and text information about the deceased, such as their hobbies, preferences, and personality. This information is then sent to the server and stored in a database.
[0512] The server analyzes the uploaded information to generate facial recognition data, voice feature data, and personality data. Based on this data, it builds a model of the deceased person's appearance, voice, and personality, thereby completing the generative AI model.
[0513] When a user selects the dialogue mode and sends a request to start dialogue to the server, the server generates video and audio of the deceased in real time based on the generative AI model and sends them to the user's device.
[0514] The device analyzes the user's voice input and sends the text data to the server, which uses a dialogue generation algorithm to generate a response and sends it back to the device, where it is played back as audio and video.
[0515] The device analyzes the user's voice and facial expressions and uses an emotion engine to recognize the user's emotions. This recognized emotion data is sent to the server, which then dynamically adjusts the content of the conversation based on that data.
[0516] Specific examples
[0517] For example, consider a case where a user wants to reunite with a deceased parent. The user uploads photos of the parent, a recording of their voice, and information about their favorite activities and hobbies to the system. The server analyzes this information to create a model of the parent's current appearance, voice, and personality. Based on this model, the user can experience a real-time conversation with the parent. If the user says, "I was busy at work today," the system might generate a response such as, "Don't work too hard." Furthermore, if the user becomes emotional, the system will adjust the conversation based on that emotion to provide a more appropriate response.
[0518] Prompt Sentence Examples
[0519] User input: "I had a busy day at work today."
[0520] Emotion: "Fatigue" (user looks tired)
[0521] Generates response: "Don't try too hard."
[0522] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0523] Step 1:
[0524] The user logs in to the system and uploads information about the deceased. In this step, the user provides the system with photos, audio, and text information about the deceased, such as their hobbies, preferences, and personality. The input is the digital data uploaded by the user, and the output is that this data is sent to the server and stored in the database.
[0525] Step 2:
[0526] The server analyzes uploaded photos to generate facial recognition data. It also analyzes uploaded audio files to extract voice feature data and create a personality model of the deceased person based on text information. The input is the deceased person's photo, audio, and text information, and the output is facial recognition data, voice feature data, and the generation of a personality model. Specific operations include using a facial recognition algorithm to extract facial features and using voice analysis software to extract voice features.
[0527] Step 3:
[0528] The server integrates the generated facial recognition data, voice feature data, and personality model to build a generative AI model of the deceased. This generative AI model is used to comprehensively recreate the appearance, voice, and personality of the deceased. The input is various analytical data, and the output is the generation of an integrated generative AI model. Specifically, an integration algorithm is used to combine each element of the model into one.
[0529] Step 4:
[0530] The user selects a dialogue mode and sends a dialogue start request to the server. The input is the dialogue start request by the user's operation, and the output is that this request is sent to the server.
[0531] Step 5:
[0532] The server generates video and audio of the deceased in real time based on the generative AI model and transmits them to the user's device. The input is a request to start a dialogue between the generative AI model and the user, and the output is the generation and transmission of video and audio data of the deceased. Specific operations include synthesizing the voice using a voice synthesizer and generating video in real time using CG technology.
[0533] Step 6:
[0534] The device analyzes the user's voice input and sends the resulting text data to the server. The input is the user's voice, and the output is the conversion of that voice into text and transmission to the server. Specifically, the process of converting voice into text is carried out using voice recognition software.
[0535] Step 7:
[0536] The server generates a response based on the user's voice input using a dialogue generation algorithm and sends it back to the device. The input is the user's voice text and a generative AI model, and the output is the generated response text and voice data. Specifically, this involves the process of generating appropriate dialogue content using a generative AI model and natural language processing technology.
[0537] Step 8:
[0538] The terminal plays the generated response as audio and video. The input is the response data sent from the server, and the output is the display and playback of the dialogue for the user. Specifically, the operation involves playing back audio data and displaying video.
[0539] Step 9:
[0540] The device analyzes the user's voice and facial expressions and uses an emotion engine to recognize the user's emotions. This emotion data is sent to a server. The input is real-time audio and video data, and the output is the emotion recognition results. Specifically, emotion analysis software is used to detect emotions from voice tone and facial expressions.
[0541] Step 10:
[0542] The server dynamically adjusts the dialogue content based on the user's emotional data. The input is the emotional data and the current dialogue content, and the output is the adjusted dialogue content. Specifically, the server selects dialogue content that matches the emotion and updates the generative AI model.
[0543] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0544] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (registered trademark) (Internet search engine).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0545] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0546] [Second embodiment]
[0547] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0548] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0549] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0550] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0551] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0552] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0553] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0554] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0555] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0556] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0557] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0558] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0559] The embodiment of this invention is a system that allows users to upload information about the deceased, generates an AI model based on that information, and provides real-time dialogue with the deceased. This will be described in detail below with specific examples.
[0560] System Overview
[0561] The system provides an interface for users to upload information about the deceased, including photos, audio files, hobbies, preferences, and personality information. This information is sent to a server, which then generates appearance, voice, and personality models of the deceased.
[0562] Initial Setup
[0563] Users first log in to the system and provide information about the deceased. On the registration screen, users upload photos and audio files of the deceased, and enter text information about their hobbies, preferences, and personality.
[0564] The device sends the uploaded information to a server, which then stores the deceased person's information in a database.
[0565] Data analysis and model generation
[0566] The server extracts facial features from uploaded photos to generate facial recognition data, analyzes voice characteristics from audio files to generate a voice model, and uses natural language processing technology to analyze text information and create a model of the deceased's personality and speech patterns.
[0567] The server then builds a model of the deceased person's appearance, voice, and personality, and combines these to generate an AI model that serves as the basis for real-time interaction with the deceased person, with their appearance and voice appropriate to their current age.
[0568] Dialogue generation and execution
[0569] The user accesses the system using a terminal and selects the mode to start a conversation. When the request to start a conversation is sent to the server, the server generates video and audio of the deceased in real time based on the generated AI model, allowing the user to instantly experience a conversation with the deceased.
[0570] The device analyzes the user's voice input and converts it into text data, which is then sent to the server, which uses a dialogue generation algorithm to generate a response, which is then sent back to the device and played back as audio and video.
[0571] Continuous model updates
[0572] The server regularly collects current trend data and keeps the AI model up to date, ensuring that conversations with the deceased always contain the latest information and maintain realism.
[0573] Specific examples
[0574] For example, consider a user who wants to reunite with their deceased grandmother. The user uploads photos of the grandmother, recordings of her voice, and information about her favorite activities and hobbies. The server analyzes this information and creates a model of the grandmother's current appearance, voice, and personality.
[0575] Next, the user launches the app and selects "Interact with Grandmother." This causes the device to display the generated video and audio of the grandmother. When the user asks, "How was your day?", the server generates a response based on the grandmother's hobbies and past data, such as, "I was tending to the flowers in the garden today. It was a lot of fun." The user can enjoy the interaction as if the deceased were still alive.
[0576] In this way, the present invention is a system that enables users to have daily conversations with deceased loved ones who are important to them, and provides them with mental comfort.
[0577] The processing flow will be explained below.
[0578] Step 1:
[0579] Users log into the system and upload information about the deceased.
[0580] Users upload photos and audio files of the deceased on the registration screen.
[0581] The user inputs text information about the deceased's hobbies, preferences, and personality.
[0582] Step 2:
[0583] The terminal transmits the uploaded data to the server.
[0584] The device sends photos, audio files, and text information to the server.
[0585] Step 3:
[0586] The server stores the received data in a database.
[0587] The server stores the photo data in an image database.
[0588] The server stores the voice data in a voice database.
[0589] The server stores the text information in a text database.
[0590] Step 4:
[0591] The server generates face authentication data based on the photograph data.
[0592] The server uses image analysis algorithms to extract facial feature points.
[0593] The server generates a face recognition model based on the feature points.
[0594] Step 5:
[0595] The server analyzes the speech data to generate a speech model.
[0596] The server uses a voice analysis algorithm to extract voice characteristics.
[0597] The server configures the voice synthesizer based on the extracted features.
[0598] Step 6:
[0599] The server generates a personality model based on the text information.
[0600] The server analyzes the text information using natural language processing techniques.
[0601] The server models the deceased's speaking style and personality based on the analysis results.
[0602] Step 7:
[0603] The server integrates the generated appearance model, voice model, and personality model to generate an AI model.
[0604] The server combines the data from each model to create a single integrated model.
[0605] Step 8:
[0606] The user initiates the interaction using a dedicated application.
[0607] The user presses a button within the application to start an interaction.
[0608] The terminal sends a request to start a conversation to the server.
[0609] Step 9:
[0610] The server generates video and audio of the deceased in real time based on an AI model.
[0611] The server uses an image generation algorithm to generate an appearance of the deceased person according to their current age.
[0612] The server generates the audio using a voice synthesizer.
[0613] The server transmits the generated video and audio to the terminal.
[0614] Step 10:
[0615] The terminal plays back the received video and audio and presents them to the user.
[0616] The device uses playback software to play back the video and audio in sync.
[0617] Step 11:
[0618] The user speaks to the deceased person.
[0619] The user speaks about the specific content of the interaction (e.g., "How was your day?").
[0620] Step 12:
[0621] The terminal converts the user's voice input into text data.
[0622] The device uses voice recognition software to convert the speech into text.
[0623] The terminal transmits the text data to the server.
[0624] Step 13:
[0625] The server executes a dialogue generation algorithm based on the received text data.
[0626] The server parses the text data and generates an appropriate response.
[0627] The server converts the generated response into text data and audio data.
[0628] Step 14:
[0629] The server transmits the generated text data and voice data to the terminal.
[0630] Step 15:
[0631] The terminal reproduces audio and video based on the transmitted data and presents a response to the user.
[0632] The terminal uses playback software to play back the audio and video in sync.
[0633] Through these steps, users can have a conversation with the deceased and have a realistic experience.
[0634] Example 1
[0635] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0636] Current technology does not yet provide a method for communicating with the deceased in real time. Therefore, there is a need to realize a dialogue with the deceased and provide users with peace of mind. There is also a need to enable the deceased to have information based on current trends. There is a need for technology that can integrate a wide range of information, such as audio, video, and personality, to enable natural dialogue.
[0637] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0638] In this invention, the server comprises: a means for a user to upload information about the deceased;
[0639] A means for the server to receive the information uploaded by the user and store it in a database;
[0640] A means for the server to analyze the uploaded photo data using image processing technology and generate facial recognition data;
[0641] A means for the server to analyze the uploaded voice data using a voice analysis technology and generate a voice model;
[0642] A server analyzes the uploaded text data using natural language processing technology to generate a personality model;
[0643] A means of integrating the generated models to build an AI model;
[0644] means for transmitting a request to a server instructing a user to initiate a dialogue using a terminal;
[0645] A means for the server to generate video and audio of the deceased in real time based on the generated AI model;
[0646] A means for converting the user's voice into text data using voice recognition technology on the user's terminal and transmitting the text data to the server;
[0647] A means for the server to analyze the text data and generate a response using a dialogue generation algorithm;
[0648] means for transmitting the generated response to a user's terminal and playing it as audio and video;
[0649] A means for the server to continuously collect trend data based on the appearance model, voice model, and personality model of the deceased and update the AI model based on the collected trend data;
[0650] A means for the server to analyze the uploaded voice data and extract voice characteristics;
[0651] The system also includes a means for the server to generate realistic voice using speech synthesis technology based on the extracted voice characteristics. This allows the server to conduct natural dialogue in real time using information about the deceased, providing the user with peace of mind. Furthermore, continuous model updates ensure that the dialogue contains the latest information, maintaining greater realism.
[0652] "User" refers to the individual or entity who accesses the system, provides information about the deceased, and initiates the interaction.
[0653] "Server" refers to the computer system that receives, stores, analyzes information about the deceased sent by users and generates an AI model.
[0654] "Terminal" refers to a device through which a user accesses the system, initiates a dialogue, provides voice input, and plays back responses from the server.
[0655] A "database" refers to a system for systematically storing and managing information about the deceased (photographs, audio data, text information, etc.).
[0656] "Image processing technology" refers to the technology that analyzes uploaded photo data and generates facial recognition data.
[0657] "Voice analysis technology" refers to technology that analyzes uploaded voice data and generates a voice model.
[0658] "Natural language processing technology" refers to technology that analyzes uploaded text data and generates a personality model.
[0659] "AI model" refers to an artificial intelligence model that integrates the appearance, voice, and personality models of the deceased, allowing the deceased to interact in real time.
[0660] "Speech recognition technology" refers to technology that converts a user's voice into text data.
[0661] A "dialogue generation algorithm" refers to an algorithm for generating an appropriate response based on input text from a user.
[0662] "Speech synthesis technology" refers to the technology that generates realistic speech based on text data.
[0663] "Trend data" refers to the latest data that reflects current information, trends, and user preferences.
[0664] This invention is a system that allows users to upload information about the deceased, generates an AI model based on that information, and provides real-time dialogue with the deceased. Specific embodiments of this system are described in detail below.
[0665] Initial Setup
[0666] Users first log in to the system and provide information about the deceased. Through the registration screen, users upload or enter information such as:
[0667] Photos (JPEG, PNG, etc.)
[0668] Audio files (MP3, WAV, etc.)
[0669] Text information about hobbies, preferences, and personality
[0670] This information is sent from the user's terminal to the server.
[0671] Receiving and storing data
[0672] The server receives the deceased person's information sent from the device and stores it securely in a database, including:
[0673] Photo data
[0674] Audio data
[0675] Text data
[0676] Data analysis and model generation
[0677] The server uses the following specific software techniques to analyze the received data:
[0678] Image processing: Using image processing libraries such as OpenCV, facial features are extracted from photos and facial recognition data is generated.
[0679] Speech analysis: Use speech analysis tools such as Google Cloud Speech-to-Text API and IBM Watson to extract speech characteristics from audio files and generate a speech model.
[0680] Text analysis: Using natural language processing libraries such as spaCy and GPT-3, text information is analyzed to model the personality and speech patterns of the deceased.
[0681] By combining this data, the server generates an AI model that serves as the basis for real-time interaction with the deceased person, with their current age-appropriate appearance and voice.
[0682] Dialogue generation and execution
[0683] A user accesses the system using his / her terminal and selects a mode to start a conversation, which sends a request to start a conversation to the server.
[0684] In response to this request, the server begins generating video and audio of the deceased person in real time based on the generated AI model. Specifically, the process involves the following steps:
[0685] Video Generation: Using 3D modeling tools and animation software, real-time video of the deceased is generated.
[0686] Voice generation: Generate the voice of the deceased using text-to-speech technology (TTS, e.g., Google Cloud Text-to-Speech).
[0687] The device analyzes the user's voice input and converts it into text data using speech recognition technology (e.g., Google Cloud Speech-to-Text), which is then sent to the server.
[0688] The server analyzes the received user text data and generates a response using a dialogue generation algorithm (e.g., GPT-3). The generated response text is converted back into speech and sent to the device.
[0689] Continuous model updates
[0690] The server periodically collects user interaction history and current trend data to update the AI model. This ensures that conversations with the deceased always contain the latest information, maintaining realism. For example, the deceased can talk naturally about their current hobbies or the latest news.
[0691] Specific examples
[0692] For example, consider a user who wants to reunite with their deceased grandmother. The user uploads photos of the grandmother, recordings of her voice, and information about her favorite activities and hobbies. The server analyzes this information and creates a model of the grandmother's current appearance, voice, and personality.
[0693] Next, the user launches the app and selects "Interact with Grandmother." This causes the device to display the generated video and audio of the grandmother. When the user asks, "How was your day?", the server generates a response based on the grandmother's hobbies and past data: "I was tending to the flowers in the garden today. It was a lot of fun." Through this interaction, the user can enjoy the experience as if the deceased were still alive.
[0694] This system allows users to have daily conversations with deceased loved ones who are important to them, providing them with spiritual comfort.
[0695] Prompt Sentence Examples
[0696] For example, when generating a dialogue using GPT-3, you can enter a prompt like this:
[0697] "The user begins a conversation with their grandmother, who likes to spend time in her garden tending to her flowers. The user asks, 'How was your day?' The grandmother talks about her day's events."
[0698] By using this prompt, the AI can generate natural dialogue that meets the user's expectations.
[0699] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0700] Step 1:
[0701] Users log in to the system and provide information about the deceased. Specifically, they access a registration screen and upload photos (JPEG, PNG, etc.), audio files (MP3, WAV, etc.), and text information about the deceased's hobbies, preferences, and personality.
[0702] Input: Photos, audio files, and text information from the user
[0703] Output: Uploaded data
[0704] Step 2:
[0705] The terminal receives the information uploaded by the user and transmits it to the server.
[0706] Input: Photos, audio files, and text information from the user
[0707] Output: Data transferred to the server
[0708] Step 3:
[0709] The server receives the data sent from the terminal and stores it in a database.
[0710] Input: Uploaded information (photos, audio files, text information)
[0711] Output: Information stored in the database
[0712] Step 4:
[0713] The server analyzes the stored photo data using image processing technology (e.g., OpenCV) and generates facial recognition data.
[0714] Input: Photo data
[0715] Output: Facial recognition data
[0716] Specific operation: Use OpenCV to extract facial feature points and convert them into data.
[0717] Step 5:
[0718] The server analyzes the stored voice data using voice analysis technology (such as Google Cloud Speech-to-Text) and generates a voice model.
[0719] Input: Audio data
[0720] Output: Audio model
[0721] Specific operation: Analyzes audio files, extracts audio characteristics, and models them.
[0722] Step 6:
[0723] The server analyzes the stored text data using natural language processing technology (such as spaCy or GPT-3) to model the personality and speaking style of the deceased.
[0724] Input: Text data
[0725] Output: personality model
[0726] Specific behavior: Analyze text data and model writing style and word usage.
[0727] Step 7:
[0728] The server combines facial recognition data, voice models, and personality models to generate an AI model.
[0729] Input: Face recognition data, voice model, personality model
[0730] Output: AI model
[0731] Specific operation: Each model is integrated and generated as a single AI model.
[0732] Step 8:
[0733] The user sends a request to start a dialogue from the terminal to the server.
[0734] Input: Dialogue-initiating request
[0735] Output: Request sent to the server
[0736] Step 9:
[0737] When the server receives a request to start a dialogue, it generates video and audio of the deceased in real time based on the generated AI model.
[0738] Input: Dialogue start request, AI model
[0739] Output: Generated video and audio
[0740] How it works: Images are generated using 3D modeling tools and animation software, and audio is generated using TTS technology.
[0741] Step 10:
[0742] The device receives the user's voice input, converts it into text data using voice recognition technology (e.g., Google Cloud Speech-to-Text), and sends it to the server.
[0743] Input: User voice input
[0744] Output: Text data
[0745] Specific operation: Converts speech into text and sends it to the server.
[0746] Step 11:
[0747] The server analyzes the text data, generates a response using a dialogue generation algorithm (e.g., GPT-3), converts the response text back into speech, and sends it to the terminal.
[0748] Input: Text data
[0749] Output: Generated response (audio and text)
[0750] Specific operation: GPT-3 is used to analyze text data, generate responses, and convert them into speech.
[0751] Step 12:
[0752] The terminal receives the response sent from the server and plays it back as audio and video.
[0753] Input: Generated response (audio and video)
[0754] Output: Replayed response
[0755] Specific operation: Plays back received audio and video in real time.
[0756] Step 13:
[0757] The server periodically collects dialogue history and current trend data to update the AI model.
[0758] Input: Dialogue history, trend data
[0759] Output: Updated AI model
[0760] What it does: Rebuild and update the model based on new data.
[0761] The above are the specific processing steps and operations of the program for this system.
[0762] (Application example 1)
[0763] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0764] Conventional memorials and memorial services have only been able to preserve memories of the deceased in static forms such as photographs and videos, and have been unable to provide an interactive dialogue experience. Furthermore, there has been a lack of systems that allow people to find spiritual comfort through dialogue with the deceased. The present invention aims to enable real-time dialogue with the deceased, providing a deeper emotional experience, especially in memorial spaces in brick-and-mortar stores.
[0765] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0766] In this invention, the server includes a means for uploading information about the deceased, a means for analyzing the uploaded information to generate appearance, voice, and personality models, and a means for generating dialogue with the deceased in real time based on the generated models, thereby enabling an interactive dialogue experience to be provided on a device installed in the memorial space of a physical store.
[0767] "Means for uploading information about the deceased" refers to an interface that allows users to submit online photos, audio files, and text information about the deceased's personality and hobbies.
[0768] "Means of analyzing uploaded information to generate appearance, voice, and personality models" refers to technology in which the server extracts facial features from photographs of the deceased, analyzes voice characteristics from audio files, and models personality patterns from text information.
[0769] "Means for generating dialogue with the deceased in real time based on the generated model" refers to a technology that combines appearance, voice, and personality models generated by the server to generate natural dialogue in response to user input and provides it in real time.
[0770] "Means for displaying and playing on devices installed within the memorial space, including display and playback means for provision in physical stores" refers to technology that displays conversations with the deceased on dedicated devices within physical stores and plays them back as audio.
[0771] "Means for continuously collecting trend data and updating the model based on that data" refers to technology in which the server periodically retrieves the latest information and dynamically updates the deceased person model to improve its accuracy.
[0772] "Means for analyzing uploaded voice data and extracting voice characteristics" refers to technology that identifies unique voice patterns from audio files uploaded by users.
[0773] "Means for generating realistic voice using a voice synthesizer" refers to technology that generates artificial voice based on the characteristics of the deceased's voice.
[0774] "Means for providing an interactive dialogue experience in a memorial space within a physical store" refers to devices or systems installed to allow visitors to have real-time conversations with models of the deceased.
[0775] The embodiments for carrying out the present invention are as follows.
[0776] System Program
[0777] The system primarily consists of a user terminal, a server, and devices installed within the memorial space. The user first uploads information about the deceased, including photos, audio files, hobbies, preferences, and personality information. The server analyzes this information and generates appearance, voice, and personality models of the deceased. These models are integrated to create an AI model that enables real-time interaction with the deceased.
[0778] Processing Description
[0779] Uploading and initial settings from the user's device
[0780] Users first log in to the system through a dedicated interface and upload photos, audio files, and written information about the deceased. The user's device then sends this information to the server.
[0781] Server-based analysis and model generation
[0782] The server extracts facial and vocal features from uploaded photos and audio files. It also uses natural language processing technology to analyze the deceased's personality and dialogue patterns from text information. Specifically, it uses facial recognition software (e.g., OpenCV), voice analysis software (e.g., Google Speech-to-Text API), and natural language processing libraries (e.g., Hugging Face Transformers). Using these technologies, the server builds and integrates appearance, voice, and personality models of the deceased to generate an AI model.
[0783] Real-time dialogue generation and display
[0784] When a user begins a dialogue with the deceased using smart glasses or a head-mounted display installed in the memorial space, the terminal analyzes the user's voice input and converts it into text data. The server uses a dialogue generation algorithm based on this text data to generate a response, which is then sent back to the terminal. The terminal then plays back the response as audio and displays a video of the deceased. A video generation library (e.g., OpenCV) is used to render the video.
[0785] Continuous model updates
[0786] The server periodically collects the latest trend data and updates the deceased person model. This ensures that interactions with the deceased always reflect the latest information, maintaining realism. Trend data collection and analysis utilizes the latest databases and cloud computing technology.
[0787] Specific examples
[0788] For example, consider a case where a user uses the system to reminisce about a deceased family member. The user uploads photos of the deceased, audio recordings, and information about their hobbies and preferences to the system. The server analyzes this information and generates an AI model of the deceased.
[0789] Next, the user puts on the smart glasses installed in the memorial space and selects "Talk to Grandma." The system then generates video and audio of the deceased in real time. When the user asks, "How was your day?", the server generates a response based on the deceased's hobbies and past data, such as, "Today I was tending to the flowers in the garden. It was a lot of fun."
[0790] Prompt Sentence Examples
[0791] "Generate a dialogue between me and my grandmother, sharing memories from the past."
[0792] "Build a realistic dialogue system that can answer questions about deceased family members' hobbies and preferences."
[0793] In this way, the present invention provides users with a very realistic interactive experience with the deceased, creating new value in the memorial space of physical stores.
[0794] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0795] Step 1:
[0796] User upload of information about deceased persons
[0797] Input: Photos of the deceased, audio files, and text information about their hobbies, preferences, and personality.
[0798] How it works: A user uploads information related to the deceased through an interface.
[0799] Output: Uploaded data is sent to the system.
[0800] Step 2:
[0801] Sending data by the device
[0802] Input: The deceased person's information uploaded in Step 1
[0803] Operation: The user terminal sends the collected data to the server.
[0804] Output: The data is received on the server and is ready for analysis.
[0805] Step 3:
[0806] Data analysis and model generation by the server
[0807] Input: Received photos, audio files, and text information of the deceased
[0808] Operation:
[0809] Extract facial features from photos using facial recognition software (e.g., OpenCV).
[0810] Analyze speech features using speech analysis software (e.g., Google Speech-to-Text API).
[0811] Use natural language processing libraries (e.g., Hugging Face Transformers) to model personality and speech patterns from text information.
[0812] Output: Appearance, voice, and personality models of the deceased are generated and integrated.
[0813] Step 4:
[0814] AI model generation by the server
[0815] Input: Appearance model, voice model, and personality model generated in Step 3
[0816] How it works: Each model of the deceased person is combined to generate an AI model capable of real-time interaction.
[0817] Output: The completed AI model is saved and used for future dialogue generation.
[0818] Step 5:
[0819] User initiated interaction
[0820] Input: Access to smart glasses or head-mounted displays installed within the memorial space
[0821] Action: The user puts on the device, launches the application and selects an interaction mode.
[0822] Output: A conversation initiation request is sent to the server.
[0823] Step 6:
[0824] Server-generated dialogue
[0825] Input: User voice input (real-time conversation)
[0826] Operation:
[0827] It uses speech recognition technology to convert the user's speech into text.
[0828] It uses natural language processing techniques to analyze the text and generate appropriate responses.
[0829] A speech synthesizer is used to convert the generated text response into speech.
[0830] Use an image generation library (e.g., OpenCV) to render a video that matches the video of the deceased.
[0831] Output: Audio and visual responses of the deceased model are generated.
[0832] Step 7:
[0833] Display and playback of terminal interactions
[0834] Input: Audio and video data generated in step 6
[0835] Operation:
[0836] The terminal plays the received audio data and displays the video data.
[0837] The deceased's response is output to the user in the memorial space.
[0838] Output: The user experiences an interactive dialogue with the deceased person.
[0839] Step 8:
[0840] Continuously updating the model with the server
[0841] Input: Latest trend data and interaction history
[0842] Operation:
[0843] The server periodically updates the deceased person model, incorporating new trend information and past interaction data.
[0844] Improve the accuracy of your model based on updated data.
[0845] Output: The AI model always reflects the latest information, maintaining the quality of real-time interactions.
[0846] Through the above steps, the present invention provides the user with a realistic interaction experience with the deceased.
[0847] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0848] The embodiment of this invention is a system that allows a user to upload information about a deceased person, generates an AI model based on that information, provides real-time dialogue with the deceased, and dynamically adjusts the content of the dialogue based on the user's emotions. This will be described in detail below with specific examples.
[0849] System Overview
[0850] The system provides an interface for users to upload information about the deceased, including photos, audio files, hobbies, preferences, and personality. This information is sent to a server, which then generates an appearance model, voice model, and personality model of the deceased. The system also incorporates an emotion engine that recognizes the user's emotions and dynamically changes the content of the dialogue.
[0851] Initial Setup
[0852] Users first log in to the system and provide information about the deceased. On the registration screen, users upload photos and audio files of the deceased, and enter text information about their hobbies, preferences, and personality.
[0853] The device sends the uploaded information to a server, which then stores the deceased person's information in a database.
[0854] Data analysis and model generation
[0855] The server extracts facial features from uploaded photos to generate facial recognition data, analyzes voice characteristics from audio files to generate a voice model, and uses natural language processing technology to analyze text information and create a model of the deceased's personality and speech patterns.
[0856] The server then builds a model of the deceased person's appearance, voice, and personality, and combines these to generate an AI model that serves as the basis for real-time interaction with the deceased person, with their appearance and voice appropriate to their current age.
[0857] Dialogue generation and execution
[0858] The user accesses the system using a terminal and selects the mode to start a conversation. When the request to start a conversation is sent to the server, the server generates video and audio of the deceased in real time based on the generated AI model, allowing the user to instantly experience a conversation with the deceased.
[0859] The device analyzes the user's voice input and converts it into text data, which is then sent to the server, which uses a dialogue generation algorithm to generate a response, which is then sent back to the device and played back as audio and video.
[0860] Emotion recognition and dialogue adjustment
[0861] The device analyzes the user's voice and video and recognizes the user's emotions using an emotion engine, which extracts emotions from the user's tone of voice and facial expressions.
[0862] The server dynamically adjusts the dialogue content based on the user's emotional data recognized by the emotion engine, thereby generating more appropriate responses to the user and improving the realism of the dialogue.
[0863] Continuous model updates
[0864] The server regularly collects current trend data and keeps the AI model up to date. This ensures that conversations with the deceased always contain the latest information, maintaining realism. The emotion engine also continually learns, improving the accuracy of user emotion recognition.
[0865] Specific examples
[0866] For example, consider a user who wants to reunite with their deceased grandmother. The user uploads photos of the grandmother, recordings of her voice, and information about her favorite activities and hobbies. The server analyzes this information and creates a model of the grandmother's current appearance, voice, and personality.
[0867] Next, the user launches the app and selects "Talk to Grandma." This causes the device to display the generated video and audio of the grandmother. When the user asks, "How was your day?", the server generates a response based on the grandmother's hobbies and past data, such as, "I was taking care of the flowers in the garden today. It was a lot of fun." Furthermore, the emotion engine recognizes the user's emotions, and if the user is excited, it adjusts the content of the conversation, such as, "Can I see those flowers?"
[0868] In this way, the present invention is a system that allows users to have daily conversations with deceased loved ones who are important to them, and provides emotionally sensitive responses, thereby achieving greater mental comfort.
[0869] The processing flow will be explained below.
[0870] Step 1:
[0871] Users log into the system and upload information about the deceased.
[0872] Users select and upload photos and audio files of the deceased on the registration screen.
[0873] The user inputs the deceased's hobbies, preferences, and personality information in text format.
[0874] Step 2:
[0875] The terminal transmits the uploaded data to the server.
[0876] The terminal transmits the photos, audio files, and text information to the server.
[0877] Step 3:
[0878] The server stores the received data in a database.
[0879] The server stores the photo data in an image database.
[0880] The server stores the voice data in a voice database.
[0881] The server stores the text information in a text database.
[0882] Step 4:
[0883] The server generates face authentication data based on the photograph data.
[0884] The server uses image analysis algorithms to extract facial feature points.
[0885] The server generates a face recognition model based on the feature points.
[0886] Step 5:
[0887] The server analyzes the speech data to generate a speech model.
[0888] The server uses a voice analysis algorithm to extract voice characteristics.
[0889] The server configures the voice synthesizer based on the extracted features.
[0890] Step 6:
[0891] The server generates a personality model based on the text information.
[0892] The server analyzes the text information using natural language processing techniques.
[0893] The server models the deceased's speaking style and personality based on the analysis results.
[0894] Step 7:
[0895] The server integrates the generated appearance model, voice model, and personality model to generate an AI model.
[0896] The server combines the data from each model to create an integrated model.
[0897] Step 8:
[0898] The user initiates the interaction using a dedicated application.
[0899] The user presses a button within the application to start an interaction.
[0900] The terminal sends a request to start a conversation to the server.
[0901] Step 9:
[0902] The server generates video and audio of the deceased in real time based on an AI model.
[0903] The server uses an image generation algorithm to generate an appearance of the deceased person according to their current age.
[0904] The server generates the audio using a voice synthesizer.
[0905] The server transmits the generated video and audio to the terminal.
[0906] Step 10:
[0907] The terminal plays back the received video and audio and presents them to the user.
[0908] The device uses playback software to play back the video and audio in sync.
[0909] Step 11:
[0910] The user speaks to the deceased person.
[0911] The user speaks about the specific content of the interaction (e.g., "How was your day?").
[0912] Step 12:
[0913] The terminal converts the user's voice input into text data.
[0914] The device uses voice recognition software to convert the speech into text.
[0915] The terminal transmits the text data to the server.
[0916] Step 13:
[0917] The server executes a dialogue generation algorithm based on the received text data.
[0918] The server parses the text data and generates an appropriate response.
[0919] The server converts the generated response into text data and audio data.
[0920] Step 14:
[0921] The server transmits the generated text data and voice data to the terminal.
[0922] Step 15:
[0923] The terminal reproduces audio and video based on the transmitted data and presents a response to the user.
[0924] The terminal uses playback software to play back the audio and video in sync.
[0925] Step 16:
[0926] The device analyzes the user's voice and video and recognizes the user's emotions using an emotion engine.
[0927] The device uses voice analysis and facial recognition to extract the user's emotions.
[0928] Step 17:
[0929] The server dynamically adjusts the content of the dialogue based on the user's emotion data recognized by the emotion engine.
[0930] The server adjusts the response to be brighter if the user's sentiment is positive.
[0931] The server adjusts its response to tone down the user's emotions if they are negative.
[0932] Through these steps, users can have a conversation with the deceased and have a realistic, emotional experience.
[0933] Example 2
[0934] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0935] In modern society, there is a demand for technology that can recreate memories of deceased loved ones. However, conventional methods have difficulty in providing a real-time conversational experience with the deceased, and they are also insufficient in adjusting the conversation to take the user's emotions into account. Therefore, there is a need for a system that provides real-time, emotionally sensitive conversations based on information about the deceased.
[0936] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0937] In this invention, the server includes means for uploading information about the deceased, means for analyzing the uploaded information to generate appearance, voice, and personality models, means for generating a real-time dialogue with the deceased based on the generated model, means for converting the user's voice into text data and generating a response using a dialogue generation algorithm, means for transmitting the generated dialogue to the user's terminal for display and playback, and means for recognizing the user's emotions and dynamically adjusting the content of the dialogue, thereby enabling a real-time dialogue experience with the deceased and providing appropriate responses that are in tune with the user's emotions.
[0938] "Information about the deceased" refers to digital data such as photographs, audio files, hobbies, preferences, and personality information of the deceased.
[0939] "Means for uploading" refers to the interface and technology that allows a user to enter information about a deceased person into the system and transmit it to the server.
[0940] "Means of analyzing and generating appearance, voice, and personality models" refers to technology that allows a server to use uploaded data to recreate the appearance, voice, and personality of the deceased as a digital model.
[0941] "Means for generating real-time interactions" refers to the functionality and algorithms that use the generated model to create interactions with the deceased person in real time.
[0942] "Means for converting into text data" refers to speech recognition technology for converting a user's voice input into text form.
[0943] "Dialogue generation algorithm" refers to an algorithm for generating appropriate responses based on input from a user.
[0944] "Means for displaying and playing" refers to the technology for displaying the response sent from the server as audio or video on the user's device.
[0945] "Means for recognizing emotions and dynamically adjusting dialogue content" refers to technology that analyzes a user's emotions and changes the dialogue content and responses to match those emotions.
[0946] "Trend data" refers to data that refers to current trends and the latest information.
[0947] "Speech synthesizer" refers to technology for converting text data into voice data and playing it back.
[0948] The embodiment of this invention is a system that allows a user to upload information about a deceased person, generates an AI model based on that information, provides real-time dialogue with the deceased, and dynamically adjusts the content of the dialogue based on the user's emotions. This will be described in detail below with specific examples.
[0949] System Overview
[0950] The system provides an interface for users to upload information about the deceased, including photos, audio files, hobbies, preferences, and personality. This information is sent to a server, which then generates an appearance model, voice model, and personality model of the deceased. The system also incorporates an emotion engine that recognizes the user's emotions and dynamically changes the content of the dialogue.
[0951] Initial Setup
[0952] The user first logs in to the system and provides information about the deceased. On the registration screen, the user uploads photos and audio files of the deceased, and enters text information about their hobbies, preferences, and personality. For example, the user might upload a photo of their grandmother working in the garden or an audio recording of their grandmother's life. After receiving this information, the device sends it to the server, which then stores the information in a database.
[0953] Data analysis and model generation
[0954] The server uses an image analysis engine to extract the facial features of the deceased from uploaded photos. For example, it uses OpenCV and facial recognition APIs to recognize features such as the position of the eyes and mouth and skin color. The server then uses a voice analysis engine to analyze the voice features of the uploaded audio file. For example, it uses voice recognition technology to extract specific patterns and voice qualities from the voice. It also uses a natural language processing engine (e.g., GPT-3) to model the personality and speaking patterns of the deceased from text information. The server combines the results of these analyses to create three models: an appearance model, a voice model, and a personality model of the deceased, and generates an AI model based on this information.
[0955] Dialogue generation and execution
[0956] To start a conversation, the user accesses the system and selects the corresponding mode. For example, they select "Dialogue with Grandma" from the app menu. The device sends this request to the server as an API request. The server uses the generated AI model to generate video and audio of the deceased person in real time to display to the user and sends them to the device. The device receives the user's voice input, converts it into text using speech recognition technology, and sends this text data to the server. The server uses a dialogue generation engine based on the text data to generate an appropriate response. For example, in response to the question, "How was your day?", the server generates a response such as, "I was taking care of the flowers in the garden today. It was a lot of fun." The device receives the response from the server, generates audio data using speech synthesis technology, and plays it back to the user as audio and video.
[0957] Emotion recognition and dialogue adjustment
[0958] The device analyzes the user's voice and video using an emotion engine (e.g., Amazon Rekognition or Microsoft Azure Emotion API) to extract the user's emotional data. The server then adjusts the content of the dialogue based on this emotional data. For example, if the user is excited, the server generates a response that reflects their emotion, such as, "Can I see that flower?"
[0959] Continuous model updates
[0960] The server periodically collects current trend data and updates the AI model, ensuring that conversations with the deceased reflect the latest information and maintain realism. The emotion engine also continuously learns, improving the accuracy of user emotion recognition.
[0961] Example operation
[0962] If a user wants to reunite with their deceased grandmother, they upload photos and audio files of the grandmother to the system, along with information about her hobbies and preferences, such as "she liked gardening" and "she was good at cooking." The server analyzes this data and generates an AI model of the grandmother. When the user selects "Interact with Grandmother," the device displays real-time video and audio of the grandmother, and the conversation begins. When the user asks, "How was your day?", the server generates a response based on the grandmother's hobbies, such as, "I was tending to the flowers in the garden today. It was so much fun." Furthermore, if the user is excited, the server dynamically adjusts the content of the conversation, such as, "Can I see those flowers?" This system strives to create a realistic conversation with a deceased person who is important to the user, providing responses that are in tune with their emotions.
[0963] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0964] Step 1: Upload your information
[0965] Users log into the system and upload photos, audio files, hobbies, preferences, and personality information about the deceased.
[0966] Input: photos, audio files, hobbies, preferences, personality information
[0967] Output: Information of the deceased person sent to the server
[0968] Specific operation: The user selects various files from the interface and clicks the "Upload" button, and the device issues an API request to send this data to the server.
[0969] Step 2: Receiving and storing data
[0970] The terminal transmits the uploaded information to the server.
[0971] Input: User-uploaded information data about the deceased
[0972] Output: Information data of the deceased person stored on the server
[0973] Specific operation: The terminal splits the data received from the user and sends it to the server's API endpoint in the appropriate format. The server records the received data in the database and returns a confirmation response.
[0974] Step 3: Appearance model generation
[0975] The server analyzes the received photo and extracts facial features.
[0976] Input: Photo data
[0977] Output: Appearance model including facial recognition data
[0978] Specific operation: The server uses an image analysis engine (e.g., OpenCV) to extract features such as eyes, mouth, and skin color from the photo. Based on these features, it generates an appearance model.
[0979] Step 4: Generate a voice model
[0980] The server analyzes the uploaded audio file and extracts audio characteristics.
[0981] Input: Audio file
[0982] Output: Audio model
[0983] Specific operation: The server uses a voice analysis engine to analyze characteristics such as pitch, tone, and speed from the audio file, and generates a voice model based on this.
[0984] Step 5: Generate a personality model
[0985] The server analyzes the text information using natural language processing technology and models the personality and speech patterns of the deceased.
[0986] Input: Text data of hobbies, preferences, and personality information
[0987] Output: personality model
[0988] Specific operation: The server uses a natural language processing engine (e.g., GPT-3) to analyze text data and extract characteristic phrases and contexts, thereby generating a personality model.
[0989] Step 6: Integrating the AI model
[0990] The server integrates the appearance model, voice model, and personality model to generate an AI model.
[0991] Input: Appearance model, voice model, personality model
[0992] Output: Integrated AI model
[0993] Specific operation: The server combines the data from each model to generate a unified AI model with a consistent personality.
[0994] Step 7: Creating and Executing Real-Time Interactions
[0995] A user accesses the system using a terminal and selects an interaction mode.
[0996] Input: User's interaction initiation request
[0997] Output: Real-time dialogue display by an AI model of the deceased person
[0998] Specific operation: The device sends a request to the server, which then generates video and audio of the deceased in real time based on the generated AI model and sends them to the device. The server then receives the user's dialogue input, converts it into text using voice recognition technology, and generates a dialogue response.
[0999] Step 8: Emotion recognition and dialogue adjustment
[1000] The device analyzes the user's voice and video to recognize the user's emotions.
[1001] Input: User's audio and video data
[1002] Output: User emotion data
[1003] How it works: The device uses an emotion engine (e.g., Amazon Rekognition) to analyze the user's tone of voice and facial expressions to extract emotional data. The server then dynamically adjusts its response based on this data.
[1004] Step 9: Continuously updating the model
[1005] The server periodically collects trend data and updates the AI model.
[1006] Input: Latest trend data
[1007] Output: Updated AI model
[1008] How it works: The server collects the latest news and trend information from the internet and reflects it in the AI model. The emotion engine also continuously learns and improves the accuracy of user emotion recognition.
[1009] These are the processing steps of the program for this system. Each step works together to provide the user with a real-time conversational experience with the deceased, generating emotionally sensitive responses.
[1010] (Application example 2)
[1011] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[1012] While existing systems exist that simulate conversations with deceased loved ones, they lack realism and are unable to provide conversations that are sensitive to the user's emotions. Furthermore, the deceased's model is fixed and not continuously updated, resulting in inconsistent conversation content. Furthermore, the system lacks the ability to properly analyze the user's voice input and generate dialogue responses, leaving a need for improved user experience.
[1013] The identification processing by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for uploading information about the deceased, means for analyzing the uploaded information to generate appearance, voice, and personality models, means for generating dialogue with the deceased in real time based on the generated models, means for recognizing the user's emotions and dynamically adjusting the dialogue content based on the emotions, and means for analyzing the user's voice input and generating appropriate dialogue responses based on the analysis results. This enables highly realistic dialogue that is in tune with the user's emotions, providing an experience that more vividly revives memories of the deceased.
[1014] The "means for uploading information about the deceased" refers to a means for providing an interface for a user to send information about the deceased, such as photos, audio, and text, to the system.
[1015] "Means for analyzing uploaded information to generate appearance, voice, and personality models" refers to means for generating digital models to reproduce the face, voice, and personality of the deceased based on uploaded information about the deceased.
[1016] "Means for generating dialogue with the deceased in real time based on the generated model" refers to means for generating dialogue in real time using a model of the appearance, voice, and personality of the deceased.
[1017] "Means for transmitting the generated dialogue to the user's device and displaying / playing it back" refers to means for transmitting the generated dialogue with the deceased to the user's device, such as a smartphone or PC, and playing it back as audio and video.
[1018] "Means for recognizing the user's emotions and dynamically adjusting the dialogue content based on those emotions" refers to means for detecting the user's emotions from the tone of voice and facial expression, and appropriately changing the dialogue content in accordance with those emotions.
[1019] The "means for analyzing uploaded voice data and extracting voice characteristics" refers to a means for extracting voice tones and speaking style characteristics from uploaded voice data.
[1020] "Means for generating realistic voice using a voice synthesizer" refers to means for generating realistic voice using voice synthesis technology based on extracted voice characteristics.
[1021] "Means for analyzing a user's voice input and generating an appropriate dialogue response based on the analysis results" refers to means for analyzing the content of what the user has said and generating an optimal response based on that content.
[1022] The system for implementing this invention is configured by combining the following means. The system allows the user to upload information about the deceased, generates an AI model based on that information, and engages in real-time dialogue with the deceased. It also has the ability to recognize the user's emotions and dynamically adjust the content of the dialogue. This allows the system to provide the user with a highly realistic dialogue experience.
[1023] System configuration
[1024] server:
[1025] Store information about the deceased in a database.
[1026] Photographs, audio, and text information of the deceased are analyzed to generate appearance, voice, and personality models.
[1027] The generated models are integrated to create a generative AI model.
[1028] To analyze a user's emotions and dynamically adjust dialogue content based on the emotions.
[1029] Real-time video and audio of the deceased are generated and transmitted to the user's device.
[1030] Device:
[1031] A user interface is provided and a means is provided for uploading information about the deceased.
[1032] It captures the user's audio and video inputs and sends them to a server for sentiment analysis.
[1033] Display and play back the dialogue sent from the server.
[1034] Implementation details
[1035] Users first log in to the system and upload photos, audio, and text information about the deceased, such as their hobbies, preferences, and personality. This information is then sent to the server and stored in a database.
[1036] The server analyzes the uploaded information to generate facial recognition data, voice feature data, and personality data. Based on this data, it builds a model of the deceased person's appearance, voice, and personality, thereby completing the generative AI model.
[1037] When a user selects the dialogue mode and sends a request to start dialogue to the server, the server generates video and audio of the deceased in real time based on the generative AI model and sends them to the user's device.
[1038] The device analyzes the user's voice input and sends the text data to the server, which uses a dialogue generation algorithm to generate a response and sends it back to the device, where it is played back as audio and video.
[1039] The device analyzes the user's voice and facial expressions and uses an emotion engine to recognize the user's emotions. This recognized emotion data is sent to the server, which then dynamically adjusts the content of the conversation based on that data.
[1040] Specific examples
[1041] For example, consider a case where a user wants to reunite with a deceased parent. The user uploads photos of the parent, a recording of their voice, and information about their favorite activities and hobbies to the system. The server analyzes this information to create a model of the parent's current appearance, voice, and personality. Based on this model, the user can experience a real-time conversation with the parent. If the user says, "I was busy at work today," the system might generate a response such as, "Don't work too hard." Furthermore, if the user becomes emotional, the system will adjust the conversation based on that emotion to provide a more appropriate response.
[1042] Prompt Sentence Examples
[1043] User input: "I had a busy day at work today."
[1044] Emotion: "Fatigue" (user looks tired)
[1045] Generates response: "Don't try too hard."
[1046] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1047] Step 1:
[1048] The user logs in to the system and uploads information about the deceased. In this step, the user provides the system with photos, audio, and text information about the deceased, such as their hobbies, preferences, and personality. The input is the digital data uploaded by the user, and the output is that this data is sent to the server and stored in the database.
[1049] Step 2:
[1050] The server analyzes uploaded photos to generate facial recognition data. It also analyzes uploaded audio files to extract voice feature data and create a personality model of the deceased person based on text information. The input is the deceased person's photo, audio, and text information, and the output is facial recognition data, voice feature data, and the generation of a personality model. Specific operations include using a facial recognition algorithm to extract facial features and using voice analysis software to extract voice features.
[1051] Step 3:
[1052] The server integrates the generated facial recognition data, voice feature data, and personality model to build a generative AI model of the deceased. This generative AI model is used to comprehensively recreate the appearance, voice, and personality of the deceased. The input is various analytical data, and the output is the generation of an integrated generative AI model. Specifically, an integration algorithm is used to combine each element of the model into one.
[1053] Step 4:
[1054] The user selects a dialogue mode and sends a dialogue start request to the server. The input is the dialogue start request by the user's operation, and the output is that this request is sent to the server.
[1055] Step 5:
[1056] The server generates video and audio of the deceased in real time based on the generative AI model and transmits them to the user's device. The input is a request to start a dialogue between the generative AI model and the user, and the output is the generation and transmission of video and audio data of the deceased. Specific operations include synthesizing the voice using a voice synthesizer and generating video in real time using CG technology.
[1057] Step 6:
[1058] The device analyzes the user's voice input and sends the resulting text data to the server. The input is the user's voice, and the output is the conversion of that voice into text and transmission to the server. Specifically, the process of converting voice into text is carried out using voice recognition software.
[1059] Step 7:
[1060] The server generates a response based on the user's voice input using a dialogue generation algorithm and sends it back to the device. The input is the user's voice text and a generative AI model, and the output is the generated response text and voice data. Specifically, this involves the process of generating appropriate dialogue content using a generative AI model and natural language processing technology.
[1061] Step 8:
[1062] The terminal plays the generated response as audio and video. The input is the response data sent from the server, and the output is the display and playback of the dialogue for the user. Specifically, the operation involves playing back audio data and displaying video.
[1063] Step 9:
[1064] The device analyzes the user's voice and facial expressions and uses an emotion engine to recognize the user's emotions. This emotion data is sent to a server. The input is real-time audio and video data, and the output is the emotion recognition results. Specifically, emotion analysis software is used to detect emotions from voice tone and facial expressions.
[1065] Step 10:
[1066] The server dynamically adjusts the dialogue content based on the user's emotional data. The input is the emotional data and the current dialogue content, and the output is the adjusted dialogue content. Specifically, the server selects dialogue content that matches the emotion and updates the generative AI model.
[1067] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1068] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1069] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[1070] [Third embodiment]
[1071] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[1072] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[1073] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1074] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[1075] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1076] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1077] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1078] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1079] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1080] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1081] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1082] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[1083] The embodiment of this invention is a system that allows users to upload information about the deceased, generates an AI model based on that information, and provides real-time dialogue with the deceased. This will be described in detail below with specific examples.
[1084] System Overview
[1085] The system provides an interface for users to upload information about the deceased, including photos, audio files, hobbies, preferences, and personality information. This information is sent to a server, which then generates appearance, voice, and personality models of the deceased.
[1086] Initial Setup
[1087] Users first log in to the system and provide information about the deceased. On the registration screen, users upload photos and audio files of the deceased, and enter text information about their hobbies, preferences, and personality.
[1088] The device sends the uploaded information to a server, which then stores the deceased person's information in a database.
[1089] Data analysis and model generation
[1090] The server extracts facial features from uploaded photos to generate facial recognition data, analyzes voice characteristics from audio files to generate a voice model, and uses natural language processing technology to analyze text information and create a model of the deceased's personality and speech patterns.
[1091] The server then builds a model of the deceased person's appearance, voice, and personality, and combines these to generate an AI model that serves as the basis for real-time interaction with the deceased person, with their appearance and voice appropriate to their current age.
[1092] Dialogue generation and execution
[1093] The user accesses the system using a terminal and selects the mode to start a conversation. When the request to start a conversation is sent to the server, the server generates video and audio of the deceased in real time based on the generated AI model, allowing the user to instantly experience a conversation with the deceased.
[1094] The device analyzes the user's voice input and converts it into text data, which is then sent to the server, which uses a dialogue generation algorithm to generate a response, which is then sent back to the device and played back as audio and video.
[1095] Continuous model updates
[1096] The server regularly collects current trend data and keeps the AI model up to date, ensuring that conversations with the deceased always contain the latest information and maintain realism.
[1097] Specific examples
[1098] For example, consider a user who wants to reunite with their deceased grandmother. The user uploads photos of the grandmother, recordings of her voice, and information about her favorite activities and hobbies. The server analyzes this information and creates a model of the grandmother's current appearance, voice, and personality.
[1099] Next, the user launches the app and selects "Interact with Grandmother." This causes the device to display the generated video and audio of the grandmother. When the user asks, "How was your day?", the server generates a response based on the grandmother's hobbies and past data, such as, "I was tending to the flowers in the garden today. It was a lot of fun." The user can enjoy the interaction as if the deceased were still alive.
[1100] In this way, the present invention is a system that enables users to have daily conversations with deceased loved ones who are important to them, and provides them with mental comfort.
[1101] The processing flow will be explained below.
[1102] Step 1:
[1103] Users log into the system and upload information about the deceased.
[1104] Users upload photos and audio files of the deceased on the registration screen.
[1105] The user inputs text information about the deceased's hobbies, preferences, and personality.
[1106] Step 2:
[1107] The terminal transmits the uploaded data to the server.
[1108] The device sends photos, audio files, and text information to the server.
[1109] Step 3:
[1110] The server stores the received data in a database.
[1111] The server stores the photo data in an image database.
[1112] The server stores the voice data in a voice database.
[1113] The server stores the text information in a text database.
[1114] Step 4:
[1115] The server generates face authentication data based on the photograph data.
[1116] The server uses image analysis algorithms to extract facial feature points.
[1117] The server generates a face recognition model based on the feature points.
[1118] Step 5:
[1119] The server analyzes the speech data to generate a speech model.
[1120] The server uses a voice analysis algorithm to extract voice characteristics.
[1121] The server configures the voice synthesizer based on the extracted features.
[1122] Step 6:
[1123] The server generates a personality model based on the text information.
[1124] The server analyzes the text information using natural language processing techniques.
[1125] The server models the deceased's speaking style and personality based on the analysis results.
[1126] Step 7:
[1127] The server integrates the generated appearance model, voice model, and personality model to generate an AI model.
[1128] The server combines the data from each model to create a single integrated model.
[1129] Step 8:
[1130] The user initiates the interaction using a dedicated application.
[1131] The user presses a button within the application to start an interaction.
[1132] The terminal sends a request to start a conversation to the server.
[1133] Step 9:
[1134] The server generates video and audio of the deceased in real time based on an AI model.
[1135] The server uses an image generation algorithm to generate an appearance of the deceased person according to their current age.
[1136] The server generates the audio using a voice synthesizer.
[1137] The server transmits the generated video and audio to the terminal.
[1138] Step 10:
[1139] The terminal plays back the received video and audio and presents them to the user.
[1140] The device uses playback software to play back the video and audio in sync.
[1141] Step 11:
[1142] The user speaks to the deceased person.
[1143] The user speaks about the specific content of the interaction (e.g., "How was your day?").
[1144] Step 12:
[1145] The terminal converts the user's voice input into text data.
[1146] The device uses voice recognition software to convert the speech into text.
[1147] The terminal transmits the text data to the server.
[1148] Step 13:
[1149] The server executes a dialogue generation algorithm based on the received text data.
[1150] The server parses the text data and generates an appropriate response.
[1151] The server converts the generated response into text data and audio data.
[1152] Step 14:
[1153] The server transmits the generated text data and voice data to the terminal.
[1154] Step 15:
[1155] The terminal reproduces audio and video based on the transmitted data and presents a response to the user.
[1156] The terminal uses playback software to play back the audio and video in sync.
[1157] Through these steps, users can have a conversation with the deceased and have a realistic experience.
[1158] Example 1
[1159] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1160] Current technology does not yet provide a method for communicating with the deceased in real time. Therefore, there is a need to realize a dialogue with the deceased and provide users with peace of mind. There is also a need to enable the deceased to have information based on current trends. There is a need for technology that can integrate a wide range of information, such as audio, video, and personality, to enable natural dialogue.
[1161] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1162] In this invention, the server comprises: a means for a user to upload information about the deceased;
[1163] A means for the server to receive the information uploaded by the user and store it in a database;
[1164] A means for the server to analyze the uploaded photo data using image processing technology and generate facial recognition data;
[1165] A means for the server to analyze the uploaded voice data using a voice analysis technology and generate a voice model;
[1166] A server analyzes the uploaded text data using natural language processing technology to generate a personality model;
[1167] A means of integrating the generated models to build an AI model;
[1168] means for transmitting a request to a server instructing a user to initiate a dialogue using a terminal;
[1169] A means for the server to generate video and audio of the deceased in real time based on the generated AI model;
[1170] A means for converting the user's voice into text data using voice recognition technology on the user's terminal and transmitting the text data to the server;
[1171] A means for the server to analyze the text data and generate a response using a dialogue generation algorithm;
[1172] means for transmitting the generated response to a user's terminal and playing it as audio and video;
[1173] A means for the server to continuously collect trend data based on the appearance model, voice model, and personality model of the deceased and update the AI model based on the collected trend data;
[1174] A means for the server to analyze the uploaded voice data and extract voice characteristics;
[1175] The system also includes a means for the server to generate realistic voice using speech synthesis technology based on the extracted voice characteristics. This allows the server to conduct natural dialogue in real time using information about the deceased, providing the user with peace of mind. Furthermore, continuous model updates ensure that the dialogue contains the latest information, maintaining greater realism.
[1176] "User" refers to the individual or entity who accesses the system, provides information about the deceased, and initiates the interaction.
[1177] "Server" refers to the computer system that receives, stores, analyzes information about the deceased sent by users and generates an AI model.
[1178] "Terminal" refers to a device through which a user accesses the system, initiates a dialogue, provides voice input, and plays back responses from the server.
[1179] A "database" refers to a system for systematically storing and managing information about the deceased (photographs, audio data, text information, etc.).
[1180] "Image processing technology" refers to the technology that analyzes uploaded photo data and generates facial recognition data.
[1181] "Voice analysis technology" refers to technology that analyzes uploaded voice data and generates a voice model.
[1182] "Natural language processing technology" refers to technology that analyzes uploaded text data and generates a personality model.
[1183] "AI model" refers to an artificial intelligence model that integrates the appearance, voice, and personality models of the deceased, allowing the deceased to interact in real time.
[1184] "Speech recognition technology" refers to technology that converts a user's voice into text data.
[1185] A "dialogue generation algorithm" refers to an algorithm for generating an appropriate response based on input text from a user.
[1186] "Speech synthesis technology" refers to the technology that generates realistic speech based on text data.
[1187] "Trend data" refers to the latest data that reflects current information, trends, and user preferences.
[1188] This invention is a system that allows users to upload information about the deceased, generates an AI model based on that information, and provides real-time dialogue with the deceased. Specific embodiments of this system are described in detail below.
[1189] Initial Setup
[1190] Users first log in to the system and provide information about the deceased. Through the registration screen, users upload or enter information such as:
[1191] Photos (JPEG, PNG, etc.)
[1192] Audio files (MP3, WAV, etc.)
[1193] Text information about hobbies, preferences, and personality
[1194] This information is sent from the user's terminal to the server.
[1195] Receiving and storing data
[1196] The server receives the deceased person's information sent from the device and stores it securely in a database, including:
[1197] Photo data
[1198] Audio data
[1199] Text data
[1200] Data analysis and model generation
[1201] The server uses the following specific software techniques to analyze the received data:
[1202] Image processing: Using image processing libraries such as OpenCV, facial features are extracted from photos and facial recognition data is generated.
[1203] Speech analysis: Use speech analysis tools such as Google Cloud Speech-to-Text API and IBM Watson to extract speech characteristics from audio files and generate a speech model.
[1204] Text analysis: Using natural language processing libraries such as spaCy and GPT-3, text information is analyzed to model the personality and speech patterns of the deceased.
[1205] By combining this data, the server generates an AI model that serves as the basis for real-time interaction with the deceased person, with their current age-appropriate appearance and voice.
[1206] Dialogue generation and execution
[1207] A user accesses the system using his / her terminal and selects a mode to start a conversation, which sends a request to start a conversation to the server.
[1208] In response to this request, the server begins generating video and audio of the deceased person in real time based on the generated AI model. Specifically, the process involves the following steps:
[1209] Video Generation: Using 3D modeling tools and animation software, real-time video of the deceased is generated.
[1210] Voice generation: Generate the voice of the deceased using text-to-speech technology (TTS, e.g., Google Cloud Text-to-Speech).
[1211] The device analyzes the user's voice input and converts it into text data using speech recognition technology (e.g., Google Cloud Speech-to-Text), which is then sent to the server.
[1212] The server analyzes the received user text data and generates a response using a dialogue generation algorithm (e.g., GPT-3). The generated response text is converted back into speech and sent to the device.
[1213] Continuous model updates
[1214] The server periodically collects user interaction history and current trend data to update the AI model. This ensures that conversations with the deceased always contain the latest information, maintaining realism. For example, the deceased can talk naturally about their current hobbies or the latest news.
[1215] Specific examples
[1216] For example, consider a user who wants to reunite with their deceased grandmother. The user uploads photos of the grandmother, recordings of her voice, and information about her favorite activities and hobbies. The server analyzes this information and creates a model of the grandmother's current appearance, voice, and personality.
[1217] Next, the user launches the app and selects "Interact with Grandmother." This causes the device to display the generated video and audio of the grandmother. When the user asks, "How was your day?", the server generates a response based on the grandmother's hobbies and past data: "I was tending to the flowers in the garden today. It was a lot of fun." Through this interaction, the user can enjoy the experience as if the deceased were still alive.
[1218] This system allows users to have daily conversations with deceased loved ones who are important to them, providing them with spiritual comfort.
[1219] Prompt Sentence Examples
[1220] For example, when generating a dialogue using GPT-3, you can enter a prompt like this:
[1221] "The user begins a conversation with their grandmother, who likes to spend time in her garden tending to her flowers. The user asks, 'How was your day?' The grandmother talks about her day's events."
[1222] By using this prompt, the AI can generate natural dialogue that meets the user's expectations.
[1223] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1224] Step 1:
[1225] Users log in to the system and provide information about the deceased. Specifically, they access a registration screen and upload photos (JPEG, PNG, etc.), audio files (MP3, WAV, etc.), and text information about the deceased's hobbies, preferences, and personality.
[1226] Input: Photos, audio files, and text information from the user
[1227] Output: Uploaded data
[1228] Step 2:
[1229] The terminal receives the information uploaded by the user and transmits it to the server.
[1230] Input: Photos, audio files, and text information from the user
[1231] Output: Data transferred to the server
[1232] Step 3:
[1233] The server receives the data sent from the terminal and stores it in a database.
[1234] Input: Uploaded information (photos, audio files, text information)
[1235] Output: Information stored in the database
[1236] Step 4:
[1237] The server analyzes the stored photo data using image processing technology (e.g., OpenCV) and generates facial recognition data.
[1238] Input: Photo data
[1239] Output: Facial recognition data
[1240] Specific operation: Use OpenCV to extract facial feature points and convert them into data.
[1241] Step 5:
[1242] The server analyzes the stored voice data using voice analysis technology (such as Google Cloud Speech-to-Text) and generates a voice model.
[1243] Input: Audio data
[1244] Output: Audio model
[1245] Specific operation: Analyzes audio files, extracts audio characteristics, and models them.
[1246] Step 6:
[1247] The server analyzes the stored text data using natural language processing technology (such as spaCy or GPT-3) to model the personality and speaking style of the deceased.
[1248] Input: Text data
[1249] Output: personality model
[1250] Specific behavior: Analyze text data and model writing style and word usage.
[1251] Step 7:
[1252] The server combines facial recognition data, voice models, and personality models to generate an AI model.
[1253] Input: Face recognition data, voice model, personality model
[1254] Output: AI model
[1255] Specific operation: Each model is integrated and generated as a single AI model.
[1256] Step 8:
[1257] The user sends a request to start a dialogue from the terminal to the server.
[1258] Input: Dialogue-initiating request
[1259] Output: Request sent to the server
[1260] Step 9:
[1261] When the server receives a request to start a dialogue, it generates video and audio of the deceased in real time based on the generated AI model.
[1262] Input: Dialogue start request, AI model
[1263] Output: Generated video and audio
[1264] How it works: Images are generated using 3D modeling tools and animation software, and audio is generated using TTS technology.
[1265] Step 10:
[1266] The device receives the user's voice input, converts it into text data using voice recognition technology (e.g., Google Cloud Speech-to-Text), and sends it to the server.
[1267] Input: User voice input
[1268] Output: Text data
[1269] Specific operation: Converts speech into text and sends it to the server.
[1270] Step 11:
[1271] The server analyzes the text data, generates a response using a dialogue generation algorithm (e.g., GPT-3), converts the response text back into speech, and sends it to the terminal.
[1272] Input: Text data
[1273] Output: Generated response (audio and text)
[1274] Specific operation: GPT-3 is used to analyze text data, generate responses, and convert them into speech.
[1275] Step 12:
[1276] The terminal receives the response sent from the server and plays it back as audio and video.
[1277] Input: Generated response (audio and video)
[1278] Output: Replayed response
[1279] Specific operation: Plays back received audio and video in real time.
[1280] Step 13:
[1281] The server periodically collects dialogue history and current trend data to update the AI model.
[1282] Input: Dialogue history, trend data
[1283] Output: Updated AI model
[1284] What it does: Rebuild and update the model based on new data.
[1285] The above are the specific processing steps and operations of the program for this system.
[1286] (Application example 1)
[1287] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1288] Conventional memorials and memorial services have only been able to preserve memories of the deceased in static forms such as photographs and videos, and have been unable to provide an interactive dialogue experience. Furthermore, there has been a lack of systems that allow people to find spiritual comfort through dialogue with the deceased. The present invention aims to enable real-time dialogue with the deceased, providing a deeper emotional experience, especially in memorial spaces in brick-and-mortar stores.
[1289] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1290] In this invention, the server includes a means for uploading information about the deceased, a means for analyzing the uploaded information to generate appearance, voice, and personality models, and a means for generating dialogue with the deceased in real time based on the generated models, thereby enabling an interactive dialogue experience to be provided on a device installed in the memorial space of a physical store.
[1291] "Means for uploading information about the deceased" refers to an interface that allows users to submit online photos, audio files, and text information about the deceased's personality and hobbies.
[1292] "Means of analyzing uploaded information to generate appearance, voice, and personality models" refers to technology in which the server extracts facial features from photographs of the deceased, analyzes voice characteristics from audio files, and models personality patterns from text information.
[1293] "Means for generating dialogue with the deceased in real time based on the generated model" refers to a technology that combines appearance, voice, and personality models generated by the server to generate natural dialogue in response to user input and provides it in real time.
[1294] "Means for displaying and playing on devices installed within the memorial space, including display and playback means for provision in physical stores" refers to technology that displays conversations with the deceased on dedicated devices within physical stores and plays them back as audio.
[1295] "Means for continuously collecting trend data and updating the model based on that data" refers to technology in which the server periodically retrieves the latest information and dynamically updates the deceased person model to improve its accuracy.
[1296] "Means for analyzing uploaded voice data and extracting voice characteristics" refers to technology that identifies unique voice patterns from audio files uploaded by users.
[1297] "Means for generating realistic voice using a voice synthesizer" refers to technology that generates artificial voice based on the characteristics of the deceased's voice.
[1298] "Means for providing an interactive dialogue experience in a memorial space within a physical store" refers to devices or systems installed to allow visitors to have real-time conversations with models of the deceased.
[1299] The embodiments for carrying out the present invention are as follows.
[1300] System Program
[1301] The system primarily consists of a user terminal, a server, and devices installed within the memorial space. The user first uploads information about the deceased, including photos, audio files, hobbies, preferences, and personality information. The server analyzes this information and generates appearance, voice, and personality models of the deceased. These models are integrated to create an AI model that enables real-time interaction with the deceased.
[1302] Processing Description
[1303] Uploading and initial settings from the user's device
[1304] Users first log in to the system through a dedicated interface and upload photos, audio files, and written information about the deceased. The user's device then sends this information to the server.
[1305] Server-based analysis and model generation
[1306] The server extracts facial and vocal features from uploaded photos and audio files. It also uses natural language processing technology to analyze the deceased's personality and dialogue patterns from text information. Specifically, it uses facial recognition software (e.g., OpenCV), voice analysis software (e.g., Google Speech-to-Text API), and natural language processing libraries (e.g., Hugging Face Transformers). Using these technologies, the server builds and integrates appearance, voice, and personality models of the deceased to generate an AI model.
[1307] Real-time dialogue generation and display
[1308] When a user begins a dialogue with the deceased using smart glasses or a head-mounted display installed in the memorial space, the terminal analyzes the user's voice input and converts it into text data. The server uses a dialogue generation algorithm based on this text data to generate a response, which is then sent back to the terminal. The terminal then plays back the response as audio and displays a video of the deceased. A video generation library (e.g., OpenCV) is used to render the video.
[1309] Continuous model updates
[1310] The server periodically collects the latest trend data and updates the deceased person model. This ensures that interactions with the deceased always reflect the latest information, maintaining realism. Trend data collection and analysis utilizes the latest databases and cloud computing technology.
[1311] Specific examples
[1312] For example, consider a case where a user uses the system to reminisce about a deceased family member. The user uploads photos of the deceased, audio recordings, and information about their hobbies and preferences to the system. The server analyzes this information and generates an AI model of the deceased.
[1313] Next, the user puts on the smart glasses installed in the memorial space and selects "Talk to Grandma." The system then generates video and audio of the deceased in real time. When the user asks, "How was your day?", the server generates a response based on the deceased's hobbies and past data, such as, "Today I was tending to the flowers in the garden. It was a lot of fun."
[1314] Prompt Sentence Examples
[1315] "Generate a dialogue between me and my grandmother, sharing memories from the past."
[1316] "Build a realistic dialogue system that can answer questions about deceased family members' hobbies and preferences."
[1317] In this way, the present invention provides users with a very realistic interactive experience with the deceased, creating new value in the memorial space of physical stores.
[1318] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1319] Step 1:
[1320] User upload of information about deceased persons
[1321] Input: Photos of the deceased, audio files, and text information about their hobbies, preferences, and personality.
[1322] How it works: A user uploads information related to the deceased through an interface.
[1323] Output: Uploaded data is sent to the system.
[1324] Step 2:
[1325] Sending data by the device
[1326] Input: The deceased person's information uploaded in Step 1
[1327] Operation: The user terminal sends the collected data to the server.
[1328] Output: The data is received on the server and is ready for analysis.
[1329] Step 3:
[1330] Data analysis and model generation by the server
[1331] Input: Received photos, audio files, and text information of the deceased
[1332] Operation:
[1333] Extract facial features from photos using facial recognition software (e.g., OpenCV).
[1334] Analyze speech features using speech analysis software (e.g., Google Speech-to-Text API).
[1335] Use natural language processing libraries (e.g., Hugging Face Transformers) to model personality and speech patterns from text information.
[1336] Output: Appearance, voice, and personality models of the deceased are generated and integrated.
[1337] Step 4:
[1338] AI model generation by the server
[1339] Input: Appearance model, voice model, and personality model generated in Step 3
[1340] How it works: Each model of the deceased person is combined to generate an AI model capable of real-time interaction.
[1341] Output: The completed AI model is saved and used for future dialogue generation.
[1342] Step 5:
[1343] User initiated interaction
[1344] Input: Access to smart glasses or head-mounted displays installed within the memorial space
[1345] Action: The user puts on the device, launches the application and selects an interaction mode.
[1346] Output: A conversation initiation request is sent to the server.
[1347] Step 6:
[1348] Server-generated dialogue
[1349] Input: User voice input (real-time conversation)
[1350] Operation:
[1351] It uses speech recognition technology to convert the user's speech into text.
[1352] It uses natural language processing techniques to analyze the text and generate appropriate responses.
[1353] A speech synthesizer is used to convert the generated text response into speech.
[1354] Use an image generation library (e.g., OpenCV) to render a video that matches the video of the deceased.
[1355] Output: Audio and visual responses of the deceased model are generated.
[1356] Step 7:
[1357] Display and playback of terminal interactions
[1358] Input: Audio and video data generated in step 6
[1359] Operation:
[1360] The terminal plays the received audio data and displays the video data.
[1361] The deceased's response is output to the user in the memorial space.
[1362] Output: The user experiences an interactive dialogue with the deceased person.
[1363] Step 8:
[1364] Continuously updating the model with the server
[1365] Input: Latest trend data and interaction history
[1366] Operation:
[1367] The server periodically updates the deceased person model, incorporating new trend information and past interaction data.
[1368] Improve the accuracy of your model based on updated data.
[1369] Output: The AI model always reflects the latest information, maintaining the quality of real-time interactions.
[1370] Through the above steps, the present invention provides the user with a realistic interaction experience with the deceased.
[1371] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1372] The embodiment of this invention is a system that allows a user to upload information about a deceased person, generates an AI model based on that information, provides real-time dialogue with the deceased, and dynamically adjusts the content of the dialogue based on the user's emotions. This will be described in detail below with specific examples.
[1373] System Overview
[1374] The system provides an interface for users to upload information about the deceased, including photos, audio files, hobbies, preferences, and personality. This information is sent to a server, which then generates an appearance model, voice model, and personality model of the deceased. The system also incorporates an emotion engine that recognizes the user's emotions and dynamically changes the content of the dialogue.
[1375] Initial Setup
[1376] Users first log in to the system and provide information about the deceased. On the registration screen, users upload photos and audio files of the deceased, and enter text information about their hobbies, preferences, and personality.
[1377] The device sends the uploaded information to a server, which then stores the deceased person's information in a database.
[1378] Data analysis and model generation
[1379] The server extracts facial features from uploaded photos to generate facial recognition data, analyzes voice characteristics from audio files to generate a voice model, and uses natural language processing technology to analyze text information and create a model of the deceased's personality and speech patterns.
[1380] The server then builds a model of the deceased person's appearance, voice, and personality, and combines these to generate an AI model that serves as the basis for real-time interaction with the deceased person, with their appearance and voice appropriate to their current age.
[1381] Dialogue generation and execution
[1382] The user accesses the system using a terminal and selects the mode to start a conversation. When the request to start a conversation is sent to the server, the server generates video and audio of the deceased in real time based on the generated AI model, allowing the user to instantly experience a conversation with the deceased.
[1383] The device analyzes the user's voice input and converts it into text data, which is then sent to the server, which uses a dialogue generation algorithm to generate a response, which is then sent back to the device and played back as audio and video.
[1384] Emotion recognition and dialogue adjustment
[1385] The device analyzes the user's voice and video and recognizes the user's emotions using an emotion engine, which extracts emotions from the user's tone of voice and facial expressions.
[1386] The server dynamically adjusts the dialogue content based on the user's emotional data recognized by the emotion engine, thereby generating more appropriate responses to the user and improving the realism of the dialogue.
[1387] Continuous model updates
[1388] The server regularly collects current trend data and keeps the AI model up to date. This ensures that conversations with the deceased always contain the latest information, maintaining realism. The emotion engine also continually learns, improving the accuracy of user emotion recognition.
[1389] Specific examples
[1390] For example, consider a user who wants to reunite with their deceased grandmother. The user uploads photos of the grandmother, recordings of her voice, and information about her favorite activities and hobbies. The server analyzes this information and creates a model of the grandmother's current appearance, voice, and personality.
[1391] Next, the user launches the app and selects "Talk to Grandma." This causes the device to display the generated video and audio of the grandmother. When the user asks, "How was your day?", the server generates a response based on the grandmother's hobbies and past data, such as, "I was taking care of the flowers in the garden today. It was a lot of fun." Furthermore, the emotion engine recognizes the user's emotions, and if the user is excited, it adjusts the content of the conversation, such as, "Can I see those flowers?"
[1392] In this way, the present invention is a system that allows users to have daily conversations with deceased loved ones who are important to them, and provides emotionally sensitive responses, thereby achieving greater mental comfort.
[1393] The processing flow will be explained below.
[1394] Step 1:
[1395] Users log into the system and upload information about the deceased.
[1396] Users select and upload photos and audio files of the deceased on the registration screen.
[1397] The user inputs the deceased's hobbies, preferences, and personality information in text format.
[1398] Step 2:
[1399] The terminal transmits the uploaded data to the server.
[1400] The terminal transmits the photos, audio files, and text information to the server.
[1401] Step 3:
[1402] The server stores the received data in a database.
[1403] The server stores the photo data in an image database.
[1404] The server stores the voice data in a voice database.
[1405] The server stores the text information in a text database.
[1406] Step 4:
[1407] The server generates face authentication data based on the photograph data.
[1408] The server uses image analysis algorithms to extract facial feature points.
[1409] The server generates a face recognition model based on the feature points.
[1410] Step 5:
[1411] The server analyzes the speech data to generate a speech model.
[1412] The server uses a voice analysis algorithm to extract voice characteristics.
[1413] The server configures the voice synthesizer based on the extracted features.
[1414] Step 6:
[1415] The server generates a personality model based on the text information.
[1416] The server analyzes the text information using natural language processing techniques.
[1417] The server models the deceased's speaking style and personality based on the analysis results.
[1418] Step 7:
[1419] The server integrates the generated appearance model, voice model, and personality model to generate an AI model.
[1420] The server combines the data from each model to create an integrated model.
[1421] Step 8:
[1422] The user initiates the interaction using a dedicated application.
[1423] The user presses a button within the application to start an interaction.
[1424] The terminal sends a request to start a conversation to the server.
[1425] Step 9:
[1426] The server generates video and audio of the deceased in real time based on an AI model.
[1427] The server uses an image generation algorithm to generate an appearance of the deceased person according to their current age.
[1428] The server generates the audio using a voice synthesizer.
[1429] The server transmits the generated video and audio to the terminal.
[1430] Step 10:
[1431] The terminal plays back the received video and audio and presents them to the user.
[1432] The device uses playback software to play back the video and audio in sync.
[1433] Step 11:
[1434] The user speaks to the deceased person.
[1435] The user speaks about the specific content of the interaction (e.g., "How was your day?").
[1436] Step 12:
[1437] The terminal converts the user's voice input into text data.
[1438] The device uses voice recognition software to convert the speech into text.
[1439] The terminal transmits the text data to the server.
[1440] Step 13:
[1441] The server executes a dialogue generation algorithm based on the received text data.
[1442] The server parses the text data and generates an appropriate response.
[1443] The server converts the generated response into text data and audio data.
[1444] Step 14:
[1445] The server transmits the generated text data and voice data to the terminal.
[1446] Step 15:
[1447] The terminal reproduces audio and video based on the transmitted data and presents a response to the user.
[1448] The terminal uses playback software to play back the audio and video in sync.
[1449] Step 16:
[1450] The device analyzes the user's voice and video and recognizes the user's emotions using an emotion engine.
[1451] The device uses voice analysis and facial recognition to extract the user's emotions.
[1452] Step 17:
[1453] The server dynamically adjusts the content of the dialogue based on the user's emotion data recognized by the emotion engine.
[1454] The server adjusts the response to be brighter if the user's sentiment is positive.
[1455] The server adjusts its response to tone down the user's emotions if they are negative.
[1456] Through these steps, users can have a conversation with the deceased and have a realistic, emotional experience.
[1457] Example 2
[1458] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1459] In modern society, there is a demand for technology that can recreate memories of deceased loved ones. However, conventional methods have difficulty in providing a real-time conversational experience with the deceased, and they are also insufficient in adjusting the conversation to take the user's emotions into account. Therefore, there is a need for a system that provides real-time, emotionally sensitive conversations based on information about the deceased.
[1460] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1461] In this invention, the server includes means for uploading information about the deceased, means for analyzing the uploaded information to generate appearance, voice, and personality models, means for generating a real-time dialogue with the deceased based on the generated model, means for converting the user's voice into text data and generating a response using a dialogue generation algorithm, means for transmitting the generated dialogue to the user's terminal for display and playback, and means for recognizing the user's emotions and dynamically adjusting the content of the dialogue, thereby enabling a real-time dialogue experience with the deceased and providing appropriate responses that are in tune with the user's emotions.
[1462] "Information about the deceased" refers to digital data such as photographs, audio files, hobbies, preferences, and personality information of the deceased.
[1463] "Means for uploading" refers to the interface and technology that allows a user to enter information about a deceased person into the system and transmit it to the server.
[1464] "Means of analyzing and generating appearance, voice, and personality models" refers to technology that allows a server to use uploaded data to recreate the appearance, voice, and personality of the deceased as a digital model.
[1465] "Means for generating real-time interactions" refers to the functionality and algorithms that use the generated model to create interactions with the deceased person in real time.
[1466] "Means for converting into text data" refers to speech recognition technology for converting a user's voice input into text form.
[1467] "Dialogue generation algorithm" refers to an algorithm for generating appropriate responses based on input from a user.
[1468] "Means for displaying and playing" refers to the technology for displaying the response sent from the server as audio or video on the user's device.
[1469] "Means for recognizing emotions and dynamically adjusting dialogue content" refers to technology that analyzes a user's emotions and changes the dialogue content and responses to match those emotions.
[1470] "Trend data" refers to data that refers to current trends and the latest information.
[1471] "Speech synthesizer" refers to technology for converting text data into voice data and playing it back.
[1472] The embodiment of this invention is a system that allows a user to upload information about a deceased person, generates an AI model based on that information, provides real-time dialogue with the deceased, and dynamically adjusts the content of the dialogue based on the user's emotions. This will be described in detail below with specific examples.
[1473] System Overview
[1474] The system provides an interface for users to upload information about the deceased, including photos, audio files, hobbies, preferences, and personality. This information is sent to a server, which then generates an appearance model, voice model, and personality model of the deceased. The system also incorporates an emotion engine that recognizes the user's emotions and dynamically changes the content of the dialogue.
[1475] Initial Setup
[1476] The user first logs in to the system and provides information about the deceased. On the registration screen, the user uploads photos and audio files of the deceased, and enters text information about their hobbies, preferences, and personality. For example, the user might upload a photo of their grandmother working in the garden or an audio recording of their grandmother's life. After receiving this information, the device sends it to the server, which then stores the information in a database.
[1477] Data analysis and model generation
[1478] The server uses an image analysis engine to extract the facial features of the deceased from uploaded photos. For example, it uses OpenCV and facial recognition APIs to recognize features such as the position of the eyes and mouth and skin color. The server then uses a voice analysis engine to analyze the voice features of the uploaded audio file. For example, it uses voice recognition technology to extract specific patterns and voice qualities from the voice. It also uses a natural language processing engine (e.g., GPT-3) to model the personality and speaking patterns of the deceased from text information. The server combines the results of these analyses to create three models: an appearance model, a voice model, and a personality model of the deceased, and generates an AI model based on this information.
[1479] Dialogue generation and execution
[1480] To start a conversation, the user accesses the system and selects the corresponding mode. For example, they select "Dialogue with Grandma" from the app menu. The device sends this request to the server as an API request. The server uses the generated AI model to generate video and audio of the deceased person in real time to display to the user and sends them to the device. The device receives the user's voice input, converts it into text using speech recognition technology, and sends this text data to the server. The server uses a dialogue generation engine based on the text data to generate an appropriate response. For example, in response to the question, "How was your day?", the server generates a response such as, "I was taking care of the flowers in the garden today. It was a lot of fun." The device receives the response from the server, generates audio data using speech synthesis technology, and plays it back to the user as audio and video.
[1481] Emotion recognition and dialogue adjustment
[1482] The device analyzes the user's voice and video using an emotion engine (e.g., Amazon Rekognition or Microsoft Azure Emotion API) to extract the user's emotional data. The server then adjusts the content of the dialogue based on this emotional data. For example, if the user is excited, the server generates a response that reflects their emotion, such as, "Can I see that flower?"
[1483] Continuous model updates
[1484] The server periodically collects current trend data and updates the AI model, ensuring that conversations with the deceased reflect the latest information and maintain realism. The emotion engine also continuously learns, improving the accuracy of user emotion recognition.
[1485] Example operation
[1486] If a user wants to reunite with their deceased grandmother, they upload photos and audio files of the grandmother to the system, along with information about her hobbies and preferences, such as "she liked gardening" and "she was good at cooking." The server analyzes this data and generates an AI model of the grandmother. When the user selects "Interact with Grandmother," the device displays real-time video and audio of the grandmother, and the conversation begins. When the user asks, "How was your day?", the server generates a response based on the grandmother's hobbies, such as, "I was tending to the flowers in the garden today. It was so much fun." Furthermore, if the user is excited, the server dynamically adjusts the content of the conversation, such as, "Can I see those flowers?" This system strives to create a realistic conversation with a deceased person who is important to the user, providing responses that are in tune with their emotions.
[1487] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1488] Step 1: Upload your information
[1489] Users log into the system and upload photos, audio files, hobbies, preferences, and personality information about the deceased.
[1490] Input: photos, audio files, hobbies, preferences, personality information
[1491] Output: Information of the deceased person sent to the server
[1492] Specific operation: The user selects various files from the interface and clicks the "Upload" button, and the device issues an API request to send this data to the server.
[1493] Step 2: Receiving and storing data
[1494] The terminal transmits the uploaded information to the server.
[1495] Input: User-uploaded information data about the deceased
[1496] Output: Information data of the deceased person stored on the server
[1497] Specific operation: The terminal splits the data received from the user and sends it to the server's API endpoint in the appropriate format. The server records the received data in the database and returns a confirmation response.
[1498] Step 3: Appearance model generation
[1499] The server analyzes the received photo and extracts facial features.
[1500] Input: Photo data
[1501] Output: Appearance model including facial recognition data
[1502] Specific operation: The server uses an image analysis engine (e.g., OpenCV) to extract features such as eyes, mouth, and skin color from the photo. Based on these features, it generates an appearance model.
[1503] Step 4: Generate a voice model
[1504] The server analyzes the uploaded audio file and extracts audio characteristics.
[1505] Input: Audio file
[1506] Output: Audio model
[1507] Specific operation: The server uses a voice analysis engine to analyze characteristics such as pitch, tone, and speed from the audio file, and generates a voice model based on this.
[1508] Step 5: Generate a personality model
[1509] The server analyzes the text information using natural language processing technology and models the personality and speech patterns of the deceased.
[1510] Input: Text data of hobbies, preferences, and personality information
[1511] Output: personality model
[1512] Specific operation: The server uses a natural language processing engine (e.g., GPT-3) to analyze text data and extract characteristic phrases and contexts, thereby generating a personality model.
[1513] Step 6: Integrating the AI model
[1514] The server integrates the appearance model, voice model, and personality model to generate an AI model.
[1515] Input: Appearance model, voice model, personality model
[1516] Output: Integrated AI model
[1517] Specific operation: The server combines the data from each model to generate a unified AI model with a consistent personality.
[1518] Step 7: Creating and Executing Real-Time Interactions
[1519] A user accesses the system using a terminal and selects an interaction mode.
[1520] Input: User's interaction initiation request
[1521] Output: Real-time dialogue display by an AI model of the deceased person
[1522] Specific operation: The device sends a request to the server, which then generates video and audio of the deceased in real time based on the generated AI model and sends them to the device. The server then receives the user's dialogue input, converts it into text using voice recognition technology, and generates a dialogue response.
[1523] Step 8: Emotion recognition and dialogue adjustment
[1524] The device analyzes the user's voice and video to recognize the user's emotions.
[1525] Input: User's audio and video data
[1526] Output: User emotion data
[1527] How it works: The device uses an emotion engine (e.g., Amazon Rekognition) to analyze the user's tone of voice and facial expressions to extract emotional data. The server then dynamically adjusts its response based on this data.
[1528] Step 9: Continuously updating the model
[1529] The server periodically collects trend data and updates the AI model.
[1530] Input: Latest trend data
[1531] Output: Updated AI model
[1532] How it works: The server collects the latest news and trend information from the internet and reflects it in the AI model. The emotion engine also continuously learns and improves the accuracy of user emotion recognition.
[1533] These are the processing steps of the program for this system. Each step works together to provide the user with a real-time conversational experience with the deceased, generating emotionally sensitive responses.
[1534] (Application example 2)
[1535] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1536] While existing systems exist that simulate conversations with deceased loved ones, they lack realism and are unable to provide conversations that are sensitive to the user's emotions. Furthermore, the deceased's model is fixed and not continuously updated, resulting in inconsistent conversation content. Furthermore, the system lacks the ability to properly analyze the user's voice input and generate dialogue responses, leaving a need for improved user experience.
[1537] The identification processing by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for uploading information about the deceased, means for analyzing the uploaded information to generate appearance, voice, and personality models, means for generating dialogue with the deceased in real time based on the generated models, means for recognizing the user's emotions and dynamically adjusting the dialogue content based on the emotions, and means for analyzing the user's voice input and generating appropriate dialogue responses based on the analysis results. This enables highly realistic dialogue that is in tune with the user's emotions, providing an experience that more vividly revives memories of the deceased.
[1538] The "means for uploading information about the deceased" refers to a means for providing an interface for a user to send information about the deceased, such as photos, audio, and text, to the system.
[1539] "Means for analyzing uploaded information to generate appearance, voice, and personality models" refers to means for generating digital models to reproduce the face, voice, and personality of the deceased based on uploaded information about the deceased.
[1540] "Means for generating dialogue with the deceased in real time based on the generated model" refers to means for generating dialogue in real time using a model of the appearance, voice, and personality of the deceased.
[1541] "Means for transmitting the generated dialogue to the user's device and displaying / playing it back" refers to means for transmitting the generated dialogue with the deceased to the user's device, such as a smartphone or PC, and playing it back as audio and video.
[1542] "Means for recognizing the user's emotions and dynamically adjusting the dialogue content based on those emotions" refers to means for detecting the user's emotions from the tone of voice and facial expression, and appropriately changing the dialogue content in accordance with those emotions.
[1543] The "means for analyzing uploaded voice data and extracting voice characteristics" refers to a means for extracting voice tones and speaking style characteristics from uploaded voice data.
[1544] "Means for generating realistic voice using a voice synthesizer" refers to means for generating realistic voice using voice synthesis technology based on extracted voice characteristics.
[1545] "Means for analyzing a user's voice input and generating an appropriate dialogue response based on the analysis results" refers to means for analyzing the content of what the user has said and generating an optimal response based on that content.
[1546] The system for implementing this invention is configured by combining the following means. The system allows the user to upload information about the deceased, generates an AI model based on that information, and engages in real-time dialogue with the deceased. It also has the ability to recognize the user's emotions and dynamically adjust the content of the dialogue. This allows the system to provide the user with a highly realistic dialogue experience.
[1547] System configuration
[1548] server:
[1549] Store information about the deceased in a database.
[1550] Photographs, audio, and text information of the deceased are analyzed to generate appearance, voice, and personality models.
[1551] The generated models are integrated to create a generative AI model.
[1552] To analyze a user's emotions and dynamically adjust dialogue content based on the emotions.
[1553] Real-time video and audio of the deceased are generated and transmitted to the user's device.
[1554] Device:
[1555] A user interface is provided and a means is provided for uploading information about the deceased.
[1556] It captures the user's audio and video inputs and sends them to a server for sentiment analysis.
[1557] Display and play back the dialogue sent from the server.
[1558] Implementation details
[1559] Users first log in to the system and upload photos, audio, and text information about the deceased, such as their hobbies, preferences, and personality. This information is then sent to the server and stored in a database.
[1560] The server analyzes the uploaded information to generate facial recognition data, voice feature data, and personality data. Based on this data, it builds a model of the deceased person's appearance, voice, and personality, thereby completing the generative AI model.
[1561] When a user selects the dialogue mode and sends a request to start dialogue to the server, the server generates video and audio of the deceased in real time based on the generative AI model and sends them to the user's device.
[1562] The device analyzes the user's voice input and sends the text data to the server, which uses a dialogue generation algorithm to generate a response and sends it back to the device, where it is played back as audio and video.
[1563] The device analyzes the user's voice and facial expressions and uses an emotion engine to recognize the user's emotions. This recognized emotion data is sent to the server, which then dynamically adjusts the content of the conversation based on that data.
[1564] Specific examples
[1565] For example, consider a case where a user wants to reunite with a deceased parent. The user uploads photos of the parent, a recording of their voice, and information about their favorite activities and hobbies to the system. The server analyzes this information to create a model of the parent's current appearance, voice, and personality. Based on this model, the user can experience a real-time conversation with the parent. If the user says, "I was busy at work today," the system might generate a response such as, "Don't work too hard." Furthermore, if the user becomes emotional, the system will adjust the conversation based on that emotion to provide a more appropriate response.
[1566] Prompt Sentence Examples
[1567] User input: "I had a busy day at work today."
[1568] Emotion: "Fatigue" (user looks tired)
[1569] Generates response: "Don't try too hard."
[1570] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1571] Step 1:
[1572] The user logs in to the system and uploads information about the deceased. In this step, the user provides the system with photos, audio, and text information about the deceased, such as their hobbies, preferences, and personality. The input is the digital data uploaded by the user, and the output is that this data is sent to the server and stored in the database.
[1573] Step 2:
[1574] The server analyzes uploaded photos to generate facial recognition data. It also analyzes uploaded audio files to extract voice feature data and create a personality model of the deceased person based on text information. The input is the deceased person's photo, audio, and text information, and the output is facial recognition data, voice feature data, and the generation of a personality model. Specific operations include using a facial recognition algorithm to extract facial features and using voice analysis software to extract voice features.
[1575] Step 3:
[1576] The server integrates the generated facial recognition data, voice feature data, and personality model to build a generative AI model of the deceased. This generative AI model is used to comprehensively recreate the appearance, voice, and personality of the deceased. The input is various analytical data, and the output is the generation of an integrated generative AI model. Specifically, an integration algorithm is used to combine each element of the model into one.
[1577] Step 4:
[1578] The user selects a dialogue mode and sends a dialogue start request to the server. The input is the dialogue start request by the user's operation, and the output is that this request is sent to the server.
[1579] Step 5:
[1580] The server generates video and audio of the deceased in real time based on the generative AI model and transmits them to the user's device. The input is a request to start a dialogue between the generative AI model and the user, and the output is the generation and transmission of video and audio data of the deceased. Specific operations include synthesizing the voice using a voice synthesizer and generating video in real time using CG technology.
[1581] Step 6:
[1582] The device analyzes the user's voice input and sends the resulting text data to the server. The input is the user's voice, and the output is the conversion of that voice into text and transmission to the server. Specifically, the process of converting voice into text is carried out using voice recognition software.
[1583] Step 7:
[1584] The server generates a response based on the user's voice input using a dialogue generation algorithm and sends it back to the device. The input is the user's voice text and a generative AI model, and the output is the generated response text and voice data. Specifically, this involves the process of generating appropriate dialogue content using a generative AI model and natural language processing technology.
[1585] Step 8:
[1586] The terminal plays the generated response as audio and video. The input is the response data sent from the server, and the output is the display and playback of the dialogue for the user. Specifically, the operation involves playing back audio data and displaying video.
[1587] Step 9:
[1588] The device analyzes the user's voice and facial expressions and uses an emotion engine to recognize the user's emotions. This emotion data is sent to a server. The input is real-time audio and video data, and the output is the emotion recognition results. Specifically, emotion analysis software is used to detect emotions from voice tone and facial expressions.
[1589] Step 10:
[1590] The server dynamically adjusts the dialogue content based on the user's emotional data. The input is the emotional data and the current dialogue content, and the output is the adjusted dialogue content. Specifically, the server selects dialogue content that matches the emotion and updates the generative AI model.
[1591] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1592] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1593] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1594] [Fourth embodiment]
[1595] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1596] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1597] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1598] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1599] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1600] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1601] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1602] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1603] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1604] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1605] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1606] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1607] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1608] The embodiment of this invention is a system that allows users to upload information about the deceased, generates an AI model based on that information, and provides real-time dialogue with the deceased. This will be described in detail below with specific examples.
[1609] System Overview
[1610] The system provides an interface for users to upload information about the deceased, including photos, audio files, hobbies, preferences, and personality information. This information is sent to a server, which then generates appearance, voice, and personality models of the deceased.
[1611] Initial Setup
[1612] Users first log in to the system and provide information about the deceased. On the registration screen, users upload photos and audio files of the deceased, and enter text information about their hobbies, preferences, and personality.
[1613] The device sends the uploaded information to a server, which then stores the deceased person's information in a database.
[1614] Data analysis and model generation
[1615] The server extracts facial features from uploaded photos to generate facial recognition data, analyzes voice characteristics from audio files to generate a voice model, and uses natural language processing technology to analyze text information and create a model of the deceased's personality and speech patterns.
[1616] The server then builds a model of the deceased person's appearance, voice, and personality, and combines these to generate an AI model that serves as the basis for real-time interaction with the deceased person, with their appearance and voice appropriate to their current age.
[1617] Dialogue generation and execution
[1618] The user accesses the system using a terminal and selects the mode to start a conversation. When the request to start a conversation is sent to the server, the server generates video and audio of the deceased in real time based on the generated AI model, allowing the user to instantly experience a conversation with the deceased.
[1619] The device analyzes the user's voice input and converts it into text data, which is then sent to the server, which uses a dialogue generation algorithm to generate a response, which is then sent back to the device and played back as audio and video.
[1620] Continuous model updates
[1621] The server regularly collects current trend data and keeps the AI model up to date, ensuring that conversations with the deceased always contain the latest information and maintain realism.
[1622] Specific examples
[1623] For example, consider a user who wants to reunite with their deceased grandmother. The user uploads photos of the grandmother, recordings of her voice, and information about her favorite activities and hobbies. The server analyzes this information and creates a model of the grandmother's current appearance, voice, and personality.
[1624] Next, the user launches the app and selects "Interact with Grandmother." This causes the device to display the generated video and audio of the grandmother. When the user asks, "How was your day?", the server generates a response based on the grandmother's hobbies and past data, such as, "I was tending to the flowers in the garden today. It was a lot of fun." The user can enjoy the interaction as if the deceased were still alive.
[1625] In this way, the present invention is a system that enables users to have daily conversations with deceased loved ones who are important to them, and provides them with mental comfort.
[1626] The processing flow will be explained below.
[1627] Step 1:
[1628] Users log into the system and upload information about the deceased.
[1629] Users upload photos and audio files of the deceased on the registration screen.
[1630] The user inputs text information about the deceased's hobbies, preferences, and personality.
[1631] Step 2:
[1632] The terminal transmits the uploaded data to the server.
[1633] The device sends photos, audio files, and text information to the server.
[1634] Step 3:
[1635] The server stores the received data in a database.
[1636] The server stores the photo data in an image database.
[1637] The server stores the voice data in a voice database.
[1638] The server stores the text information in a text database.
[1639] Step 4:
[1640] The server generates face authentication data based on the photograph data.
[1641] The server uses image analysis algorithms to extract facial feature points.
[1642] The server generates a face recognition model based on the feature points.
[1643] Step 5:
[1644] The server analyzes the speech data to generate a speech model.
[1645] The server uses a voice analysis algorithm to extract voice characteristics.
[1646] The server configures the voice synthesizer based on the extracted features.
[1647] Step 6:
[1648] The server generates a personality model based on the text information.
[1649] The server analyzes the text information using natural language processing techniques.
[1650] The server models the deceased's speaking style and personality based on the analysis results.
[1651] Step 7:
[1652] The server integrates the generated appearance model, voice model, and personality model to generate an AI model.
[1653] The server combines the data from each model to create a single integrated model.
[1654] Step 8:
[1655] The user initiates the interaction using a dedicated application.
[1656] The user presses a button within the application to start an interaction.
[1657] The terminal sends a request to start a conversation to the server.
[1658] Step 9:
[1659] The server generates video and audio of the deceased in real time based on an AI model.
[1660] The server uses an image generation algorithm to generate an appearance of the deceased person according to their current age.
[1661] The server generates the audio using a voice synthesizer.
[1662] The server transmits the generated video and audio to the terminal.
[1663] Step 10:
[1664] The terminal plays back the received video and audio and presents them to the user.
[1665] The device uses playback software to play back the video and audio in sync.
[1666] Step 11:
[1667] The user speaks to the deceased person.
[1668] The user speaks about the specific content of the interaction (e.g., "How was your day?").
[1669] Step 12:
[1670] The terminal converts the user's voice input into text data.
[1671] The device uses voice recognition software to convert the speech into text.
[1672] The terminal transmits the text data to the server.
[1673] Step 13:
[1674] The server executes a dialogue generation algorithm based on the received text data.
[1675] The server parses the text data and generates an appropriate response.
[1676] The server converts the generated response into text data and audio data.
[1677] Step 14:
[1678] The server transmits the generated text data and voice data to the terminal.
[1679] Step 15:
[1680] The terminal reproduces audio and video based on the transmitted data and presents a response to the user.
[1681] The terminal uses playback software to play back the audio and video in sync.
[1682] Through these steps, users can have a conversation with the deceased and have a realistic experience.
[1683] Example 1
[1684] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1685] Current technology does not yet provide a method for communicating with the deceased in real time. Therefore, there is a need to realize a dialogue with the deceased and provide users with peace of mind. There is also a need to enable the deceased to have information based on current trends. There is a need for technology that can integrate a wide range of information, such as audio, video, and personality, to enable natural dialogue.
[1686] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1687] In this invention, the server comprises: a means for a user to upload information about the deceased;
[1688] A means for the server to receive the information uploaded by the user and store it in a database;
[1689] A means for the server to analyze the uploaded photo data using image processing technology and generate facial recognition data;
[1690] A means for the server to analyze the uploaded voice data using a voice analysis technology and generate a voice model;
[1691] A server analyzes the uploaded text data using natural language processing technology to generate a personality model;
[1692] A means of integrating the generated models to build an AI model;
[1693] means for transmitting a request to a server instructing a user to initiate a dialogue using a terminal;
[1694] A means for the server to generate video and audio of the deceased in real time based on the generated AI model;
[1695] A means for converting the user's voice into text data using voice recognition technology on the user's terminal and transmitting the text data to the server;
[1696] A means for the server to analyze the text data and generate a response using a dialogue generation algorithm;
[1697] means for transmitting the generated response to a user's terminal and playing it as audio and video;
[1698] A means for the server to continuously collect trend data based on the appearance model, voice model, and personality model of the deceased and update the AI model based on the collected trend data;
[1699] A means for the server to analyze the uploaded voice data and extract voice characteristics;
[1700] The system also includes a means for the server to generate realistic voice using speech synthesis technology based on the extracted voice characteristics. This allows the server to conduct natural dialogue in real time using information about the deceased, providing the user with peace of mind. Furthermore, continuous model updates ensure that the dialogue contains the latest information, maintaining greater realism.
[1701] "User" refers to the individual or entity who accesses the system, provides information about the deceased, and initiates the interaction.
[1702] "Server" refers to the computer system that receives, stores, analyzes information about the deceased sent by users and generates an AI model.
[1703] "Terminal" refers to a device through which a user accesses the system, initiates a dialogue, provides voice input, and plays back responses from the server.
[1704] A "database" refers to a system for systematically storing and managing information about the deceased (photographs, audio data, text information, etc.).
[1705] "Image processing technology" refers to the technology that analyzes uploaded photo data and generates facial recognition data.
[1706] "Voice analysis technology" refers to technology that analyzes uploaded voice data and generates a voice model.
[1707] "Natural language processing technology" refers to technology that analyzes uploaded text data and generates a personality model.
[1708] "AI model" refers to an artificial intelligence model that integrates the appearance, voice, and personality models of the deceased, allowing the deceased to interact in real time.
[1709] "Speech recognition technology" refers to technology that converts a user's voice into text data.
[1710] A "dialogue generation algorithm" refers to an algorithm for generating an appropriate response based on input text from a user.
[1711] "Speech synthesis technology" refers to the technology that generates realistic speech based on text data.
[1712] "Trend data" refers to the latest data that reflects current information, trends, and user preferences.
[1713] This invention is a system that allows users to upload information about the deceased, generates an AI model based on that information, and provides real-time dialogue with the deceased. Specific embodiments of this system are described in detail below.
[1714] Initial Setup
[1715] Users first log in to the system and provide information about the deceased. Through the registration screen, users upload or enter information such as:
[1716] Photos (JPEG, PNG, etc.)
[1717] Audio files (MP3, WAV, etc.)
[1718] Text information about hobbies, preferences, and personality
[1719] This information is sent from the user's terminal to the server.
[1720] Receiving and storing data
[1721] The server receives the deceased person's information sent from the device and stores it securely in a database, including:
[1722] Photo data
[1723] Audio data
[1724] Text data
[1725] Data analysis and model generation
[1726] The server uses the following specific software techniques to analyze the received data:
[1727] Image processing: Using image processing libraries such as OpenCV, facial features are extracted from photos and facial recognition data is generated.
[1728] Speech analysis: Use speech analysis tools such as Google Cloud Speech-to-Text API and IBM Watson to extract speech characteristics from audio files and generate a speech model.
[1729] Text analysis: Using natural language processing libraries such as spaCy and GPT-3, text information is analyzed to model the personality and speech patterns of the deceased.
[1730] By combining this data, the server generates an AI model that serves as the basis for real-time interaction with the deceased person, with their current age-appropriate appearance and voice.
[1731] Dialogue generation and execution
[1732] A user accesses the system using his / her terminal and selects a mode to start a conversation, which sends a request to start a conversation to the server.
[1733] In response to this request, the server begins generating video and audio of the deceased person in real time based on the generated AI model. Specifically, the process involves the following steps:
[1734] Video Generation: Using 3D modeling tools and animation software, real-time video of the deceased is generated.
[1735] Voice generation: Generate the voice of the deceased using text-to-speech technology (TTS, e.g., Google Cloud Text-to-Speech).
[1736] The device analyzes the user's voice input and converts it into text data using speech recognition technology (e.g., Google Cloud Speech-to-Text), which is then sent to the server.
[1737] The server analyzes the received user text data and generates a response using a dialogue generation algorithm (e.g., GPT-3). The generated response text is converted back into speech and sent to the device.
[1738] Continuous model updates
[1739] The server periodically collects user interaction history and current trend data to update the AI model. This ensures that conversations with the deceased always contain the latest information, maintaining realism. For example, the deceased can talk naturally about their current hobbies or the latest news.
[1740] Specific examples
[1741] For example, consider a user who wants to reunite with their deceased grandmother. The user uploads photos of the grandmother, recordings of her voice, and information about her favorite activities and hobbies. The server analyzes this information and creates a model of the grandmother's current appearance, voice, and personality.
[1742] Next, the user launches the app and selects "Interact with Grandmother." This causes the device to display the generated video and audio of the grandmother. When the user asks, "How was your day?", the server generates a response based on the grandmother's hobbies and past data: "I was tending to the flowers in the garden today. It was a lot of fun." Through this interaction, the user can enjoy the experience as if the deceased were still alive.
[1743] This system allows users to have daily conversations with deceased loved ones who are important to them, providing them with spiritual comfort.
[1744] Prompt Sentence Examples
[1745] For example, when generating a dialogue using GPT-3, you can enter a prompt like this:
[1746] "The user begins a conversation with their grandmother, who likes to spend time in her garden tending to her flowers. The user asks, 'How was your day?' The grandmother talks about her day's events."
[1747] By using this prompt, the AI can generate natural dialogue that meets the user's expectations.
[1748] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1749] Step 1:
[1750] Users log in to the system and provide information about the deceased. Specifically, they access a registration screen and upload photos (JPEG, PNG, etc.), audio files (MP3, WAV, etc.), and text information about the deceased's hobbies, preferences, and personality.
[1751] Input: Photos, audio files, and text information from the user
[1752] Output: Uploaded data
[1753] Step 2:
[1754] The terminal receives the information uploaded by the user and transmits it to the server.
[1755] Input: Photos, audio files, and text information from the user
[1756] Output: Data transferred to the server
[1757] Step 3:
[1758] The server receives the data sent from the terminal and stores it in a database.
[1759] Input: Uploaded information (photos, audio files, text information)
[1760] Output: Information stored in the database
[1761] Step 4:
[1762] The server analyzes the stored photo data using image processing technology (e.g., OpenCV) and generates facial recognition data.
[1763] Input: Photo data
[1764] Output: Facial recognition data
[1765] Specific operation: Use OpenCV to extract facial feature points and convert them into data.
[1766] Step 5:
[1767] The server analyzes the stored voice data using voice analysis technology (such as Google Cloud Speech-to-Text) and generates a voice model.
[1768] Input: Audio data
[1769] Output: Audio model
[1770] Specific operation: Analyzes audio files, extracts audio characteristics, and models them.
[1771] Step 6:
[1772] The server analyzes the stored text data using natural language processing technology (such as spaCy or GPT-3) to model the personality and speaking style of the deceased.
[1773] Input: Text data
[1774] Output: personality model
[1775] Specific behavior: Analyze text data and model writing style and word usage.
[1776] Step 7:
[1777] The server combines facial recognition data, voice models, and personality models to generate an AI model.
[1778] Input: Face recognition data, voice model, personality model
[1779] Output: AI model
[1780] Specific operation: Each model is integrated and generated as a single AI model.
[1781] Step 8:
[1782] The user sends a request to start a dialogue from the terminal to the server.
[1783] Input: Dialogue-initiating request
[1784] Output: Request sent to the server
[1785] Step 9:
[1786] When the server receives a request to start a dialogue, it generates video and audio of the deceased in real time based on the generated AI model.
[1787] Input: Dialogue start request, AI model
[1788] Output: Generated video and audio
[1789] How it works: Images are generated using 3D modeling tools and animation software, and audio is generated using TTS technology.
[1790] Step 10:
[1791] The device receives the user's voice input, converts it into text data using voice recognition technology (e.g., Google Cloud Speech-to-Text), and sends it to the server.
[1792] Input: User voice input
[1793] Output: Text data
[1794] Specific operation: Converts speech into text and sends it to the server.
[1795] Step 11:
[1796] The server analyzes the text data, generates a response using a dialogue generation algorithm (e.g., GPT-3), converts the response text back into speech, and sends it to the terminal.
[1797] Input: Text data
[1798] Output: Generated response (audio and text)
[1799] Specific operation: GPT-3 is used to analyze text data, generate responses, and convert them into speech.
[1800] Step 12:
[1801] The terminal receives the response sent from the server and plays it back as audio and video.
[1802] Input: Generated response (audio and video)
[1803] Output: Replayed response
[1804] Specific operation: Plays back received audio and video in real time.
[1805] Step 13:
[1806] The server periodically collects dialogue history and current trend data to update the AI model.
[1807] Input: Dialogue history, trend data
[1808] Output: Updated AI model
[1809] What it does: Rebuild and update the model based on new data.
[1810] The above are the specific processing steps and operations of the program for this system.
[1811] (Application example 1)
[1812] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1813] Conventional memorials and memorial services have only been able to preserve memories of the deceased in static forms such as photographs and videos, and have been unable to provide an interactive dialogue experience. Furthermore, there has been a lack of systems that allow people to find spiritual comfort through dialogue with the deceased. The present invention aims to enable real-time dialogue with the deceased, providing a deeper emotional experience, especially in memorial spaces in brick-and-mortar stores.
[1814] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1815] In this invention, the server includes a means for uploading information about the deceased, a means for analyzing the uploaded information to generate appearance, voice, and personality models, and a means for generating dialogue with the deceased in real time based on the generated models, thereby enabling an interactive dialogue experience to be provided on a device installed in the memorial space of a physical store.
[1816] "Means for uploading information about the deceased" refers to an interface that allows users to submit online photos, audio files, and text information about the deceased's personality and hobbies.
[1817] "Means of analyzing uploaded information to generate appearance, voice, and personality models" refers to technology in which the server extracts facial features from photographs of the deceased, analyzes voice characteristics from audio files, and models personality patterns from text information.
[1818] "Means for generating dialogue with the deceased in real time based on the generated model" refers to a technology that combines appearance, voice, and personality models generated by the server to generate natural dialogue in response to user input and provides it in real time.
[1819] "Means for displaying and playing on devices installed within the memorial space, including display and playback means for provision in physical stores" refers to technology that displays conversations with the deceased on dedicated devices within physical stores and plays them back as audio.
[1820] "Means for continuously collecting trend data and updating the model based on that data" refers to technology in which the server periodically retrieves the latest information and dynamically updates the deceased person model to improve its accuracy.
[1821] "Means for analyzing uploaded voice data and extracting voice characteristics" refers to technology that identifies unique voice patterns from audio files uploaded by users.
[1822] "Means for generating realistic voice using a voice synthesizer" refers to technology that generates artificial voice based on the characteristics of the deceased's voice.
[1823] "Means for providing an interactive dialogue experience in a memorial space within a physical store" refers to devices or systems installed to allow visitors to have real-time conversations with models of the deceased.
[1824] The embodiments for carrying out the present invention are as follows.
[1825] System Program
[1826] The system primarily consists of a user terminal, a server, and devices installed within the memorial space. The user first uploads information about the deceased, including photos, audio files, hobbies, preferences, and personality information. The server analyzes this information and generates appearance, voice, and personality models of the deceased. These models are integrated to create an AI model that enables real-time interaction with the deceased.
[1827] Processing Description
[1828] Uploading and initial settings from the user's device
[1829] Users first log in to the system through a dedicated interface and upload photos, audio files, and written information about the deceased. The user's device then sends this information to the server.
[1830] Server-based analysis and model generation
[1831] The server extracts facial and vocal features from uploaded photos and audio files. It also uses natural language processing technology to analyze the deceased's personality and dialogue patterns from text information. Specifically, it uses facial recognition software (e.g., OpenCV), voice analysis software (e.g., Google Speech-to-Text API), and natural language processing libraries (e.g., Hugging Face Transformers). Using these technologies, the server builds and integrates appearance, voice, and personality models of the deceased to generate an AI model.
[1832] Real-time dialogue generation and display
[1833] When a user begins a dialogue with the deceased using smart glasses or a head-mounted display installed in the memorial space, the terminal analyzes the user's voice input and converts it into text data. The server uses a dialogue generation algorithm based on this text data to generate a response, which is then sent back to the terminal. The terminal then plays back the response as audio and displays a video of the deceased. A video generation library (e.g., OpenCV) is used to render the video.
[1834] Continuous model updates
[1835] The server periodically collects the latest trend data and updates the deceased person model. This ensures that interactions with the deceased always reflect the latest information, maintaining realism. Trend data collection and analysis utilizes the latest databases and cloud computing technology.
[1836] Specific examples
[1837] For example, consider a case where a user uses the system to reminisce about a deceased family member. The user uploads photos of the deceased, audio recordings, and information about their hobbies and preferences to the system. The server analyzes this information and generates an AI model of the deceased.
[1838] Next, the user puts on the smart glasses installed in the memorial space and selects "Talk to Grandma." The system then generates video and audio of the deceased in real time. When the user asks, "How was your day?", the server generates a response based on the deceased's hobbies and past data, such as, "Today I was tending to the flowers in the garden. It was a lot of fun."
[1839] Prompt Sentence Examples
[1840] "Generate a dialogue between me and my grandmother, sharing memories from the past."
[1841] "Build a realistic dialogue system that can answer questions about deceased family members' hobbies and preferences."
[1842] In this way, the present invention provides users with a very realistic interactive experience with the deceased, creating new value in the memorial space of physical stores.
[1843] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1844] Step 1:
[1845] User upload of information about deceased persons
[1846] Input: Photos of the deceased, audio files, and text information about their hobbies, preferences, and personality.
[1847] How it works: A user uploads information related to the deceased through an interface.
[1848] Output: Uploaded data is sent to the system.
[1849] Step 2:
[1850] Sending data by the device
[1851] Input: The deceased person's information uploaded in Step 1
[1852] Operation: The user terminal sends the collected data to the server.
[1853] Output: The data is received on the server and is ready for analysis.
[1854] Step 3:
[1855] Data analysis and model generation by the server
[1856] Input: Received photos, audio files, and text information of the deceased
[1857] Operation:
[1858] Extract facial features from photos using facial recognition software (e.g., OpenCV).
[1859] Analyze speech features using speech analysis software (e.g., Google Speech-to-Text API).
[1860] Use natural language processing libraries (e.g., Hugging Face Transformers) to model personality and speech patterns from text information.
[1861] Output: Appearance, voice, and personality models of the deceased are generated and integrated.
[1862] Step 4:
[1863] AI model generation by the server
[1864] Input: Appearance model, voice model, and personality model generated in Step 3
[1865] How it works: Each model of the deceased person is combined to generate an AI model capable of real-time interaction.
[1866] Output: The completed AI model is saved and used for future dialogue generation.
[1867] Step 5:
[1868] User initiated interaction
[1869] Input: Access to smart glasses or head-mounted displays installed within the memorial space
[1870] Action: The user puts on the device, launches the application and selects an interaction mode.
[1871] Output: A conversation initiation request is sent to the server.
[1872] Step 6:
[1873] Server-generated dialogue
[1874] Input: User voice input (real-time conversation)
[1875] Operation:
[1876] It uses speech recognition technology to convert the user's speech into text.
[1877] It uses natural language processing techniques to analyze the text and generate appropriate responses.
[1878] A speech synthesizer is used to convert the generated text response into speech.
[1879] Use an image generation library (e.g., OpenCV) to render a video that matches the video of the deceased.
[1880] Output: Audio and visual responses of the deceased model are generated.
[1881] Step 7:
[1882] Display and playback of terminal interactions
[1883] Input: Audio and video data generated in step 6
[1884] Operation:
[1885] The terminal plays the received audio data and displays the video data.
[1886] The deceased's response is output to the user in the memorial space.
[1887] Output: The user experiences an interactive dialogue with the deceased person.
[1888] Step 8:
[1889] Continuously updating the model with the server
[1890] Input: Latest trend data and interaction history
[1891] Operation:
[1892] The server periodically updates the deceased person model, incorporating new trend information and past interaction data.
[1893] Improve the accuracy of your model based on updated data.
[1894] Output: The AI model always reflects the latest information, maintaining the quality of real-time interactions.
[1895] Through the above steps, the present invention provides the user with a realistic interaction experience with the deceased.
[1896] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1897] The embodiment of this invention is a system that allows a user to upload information about a deceased person, generates an AI model based on that information, provides real-time dialogue with the deceased, and dynamically adjusts the content of the dialogue based on the user's emotions. This will be described in detail below with specific examples.
[1898] System Overview
[1899] The system provides an interface for users to upload information about the deceased, including photos, audio files, hobbies, preferences, and personality. This information is sent to a server, which then generates an appearance model, voice model, and personality model of the deceased. The system also incorporates an emotion engine that recognizes the user's emotions and dynamically changes the content of the dialogue.
[1900] Initial Setup
[1901] Users first log in to the system and provide information about the deceased. On the registration screen, users upload photos and audio files of the deceased, and enter text information about their hobbies, preferences, and personality.
[1902] The device sends the uploaded information to a server, which then stores the deceased person's information in a database.
[1903] Data analysis and model generation
[1904] The server extracts facial features from uploaded photos to generate facial recognition data, analyzes voice characteristics from audio files to generate a voice model, and uses natural language processing technology to analyze text information and create a model of the deceased's personality and speech patterns.
[1905] The server then builds a model of the deceased person's appearance, voice, and personality, and combines these to generate an AI model that serves as the basis for real-time interaction with the deceased person, with their appearance and voice appropriate to their current age.
[1906] Dialogue generation and execution
[1907] The user accesses the system using a terminal and selects the mode to start a conversation. When the request to start a conversation is sent to the server, the server generates video and audio of the deceased in real time based on the generated AI model, allowing the user to instantly experience a conversation with the deceased.
[1908] The device analyzes the user's voice input and converts it into text data, which is then sent to the server, which uses a dialogue generation algorithm to generate a response, which is then sent back to the device and played back as audio and video.
[1909] Emotion recognition and dialogue adjustment
[1910] The device analyzes the user's voice and video and recognizes the user's emotions using an emotion engine, which extracts emotions from the user's tone of voice and facial expressions.
[1911] The server dynamically adjusts the dialogue content based on the user's emotional data recognized by the emotion engine, thereby generating more appropriate responses to the user and improving the realism of the dialogue.
[1912] Continuous model updates
[1913] The server regularly collects current trend data and keeps the AI model up to date. This ensures that conversations with the deceased always contain the latest information, maintaining realism. The emotion engine also continually learns, improving the accuracy of user emotion recognition.
[1914] Specific examples
[1915] For example, consider a user who wants to reunite with their deceased grandmother. The user uploads photos of the grandmother, recordings of her voice, and information about her favorite activities and hobbies. The server analyzes this information and creates a model of the grandmother's current appearance, voice, and personality.
[1916] Next, the user launches the app and selects "Talk to Grandma." This causes the device to display the generated video and audio of the grandmother. When the user asks, "How was your day?", the server generates a response based on the grandmother's hobbies and past data, such as, "I was taking care of the flowers in the garden today. It was a lot of fun." Furthermore, the emotion engine recognizes the user's emotions, and if the user is excited, it adjusts the content of the conversation, such as, "Can I see those flowers?"
[1917] In this way, the present invention is a system that allows users to have daily conversations with deceased loved ones who are important to them, and provides emotionally sensitive responses, thereby achieving greater mental comfort.
[1918] The processing flow will be explained below.
[1919] Step 1:
[1920] Users log into the system and upload information about the deceased.
[1921] Users select and upload photos and audio files of the deceased on the registration screen.
[1922] The user inputs the deceased's hobbies, preferences, and personality information in text format.
[1923] Step 2:
[1924] The terminal transmits the uploaded data to the server.
[1925] The terminal transmits the photos, audio files, and text information to the server.
[1926] Step 3:
[1927] The server stores the received data in a database.
[1928] The server stores the photo data in an image database.
[1929] The server stores the voice data in a voice database.
[1930] The server stores the text information in a text database.
[1931] Step 4:
[1932] The server generates face authentication data based on the photograph data.
[1933] The server uses image analysis algorithms to extract facial feature points.
[1934] The server generates a face recognition model based on the feature points.
[1935] Step 5:
[1936] The server analyzes the speech data to generate a speech model.
[1937] The server uses a voice analysis algorithm to extract voice characteristics.
[1938] The server configures the voice synthesizer based on the extracted features.
[1939] Step 6:
[1940] The server generates a personality model based on the text information.
[1941] The server analyzes the text information using natural language processing techniques.
[1942] The server models the deceased's speaking style and personality based on the analysis results.
[1943] Step 7:
[1944] The server integrates the generated appearance model, voice model, and personality model to generate an AI model.
[1945] The server combines the data from each model to create an integrated model.
[1946] Step 8:
[1947] The user initiates the interaction using a dedicated application.
[1948] The user presses a button within the application to start an interaction.
[1949] The terminal sends a request to start a conversation to the server.
[1950] Step 9:
[1951] The server generates video and audio of the deceased in real time based on an AI model.
[1952] The server uses an image generation algorithm to generate an appearance of the deceased person according to their current age.
[1953] The server generates the audio using a voice synthesizer.
[1954] The server transmits the generated video and audio to the terminal.
[1955] Step 10:
[1956] The terminal plays back the received video and audio and presents them to the user.
[1957] The device uses playback software to play back the video and audio in sync.
[1958] Step 11:
[1959] The user speaks to the deceased person.
[1960] The user speaks about the specific content of the interaction (e.g., "How was your day?").
[1961] Step 12:
[1962] The terminal converts the user's voice input into text data.
[1963] The device uses voice recognition software to convert the speech into text.
[1964] The terminal transmits the text data to the server.
[1965] Step 13:
[1966] The server executes a dialogue generation algorithm based on the received text data.
[1967] The server parses the text data and generates an appropriate response.
[1968] The server converts the generated response into text data and audio data.
[1969] Step 14:
[1970] The server transmits the generated text data and voice data to the terminal.
[1971] Step 15:
[1972] The terminal reproduces audio and video based on the transmitted data and presents a response to the user.
[1973] The terminal uses playback software to play back the audio and video in sync.
[1974] Step 16:
[1975] The device analyzes the user's voice and video and recognizes the user's emotions using an emotion engine.
[1976] The device uses voice analysis and facial recognition to extract the user's emotions.
[1977] Step 17:
[1978] The server dynamically adjusts the content of the dialogue based on the user's emotion data recognized by the emotion engine.
[1979] The server adjusts the response to be brighter if the user's sentiment is positive.
[1980] The server adjusts its response to tone down the user's emotions if they are negative.
[1981] Through these steps, users can have a conversation with the deceased and have a realistic, emotional experience.
[1982] Example 2
[1983] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1984] In modern society, there is a demand for technology that can recreate memories of deceased loved ones. However, conventional methods have difficulty in providing a real-time conversational experience with the deceased, and they are also insufficient in adjusting the conversation to take the user's emotions into account. Therefore, there is a need for a system that provides real-time, emotionally sensitive conversations based on information about the deceased.
[1985] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1986] In this invention, the server includes means for uploading information about the deceased, means for analyzing the uploaded information to generate appearance, voice, and personality models, means for generating a real-time dialogue with the deceased based on the generated model, means for converting the user's voice into text data and generating a response using a dialogue generation algorithm, means for transmitting the generated dialogue to the user's terminal for display and playback, and means for recognizing the user's emotions and dynamically adjusting the content of the dialogue, thereby enabling a real-time dialogue experience with the deceased and providing appropriate responses that are in tune with the user's emotions.
[1987] "Information about the deceased" refers to digital data such as photographs, audio files, hobbies, preferences, and personality information of the deceased.
[1988] "Means for uploading" refers to the interface and technology that allows a user to enter information about a deceased person into the system and transmit it to the server.
[1989] "Means of analyzing and generating appearance, voice, and personality models" refers to technology that allows a server to use uploaded data to recreate the appearance, voice, and personality of the deceased as a digital model.
[1990] "Means for generating real-time interactions" refers to the functionality and algorithms that use the generated model to create interactions with the deceased person in real time.
[1991] "Means for converting into text data" refers to speech recognition technology for converting a user's voice input into text form.
[1992] "Dialogue generation algorithm" refers to an algorithm for generating appropriate responses based on input from a user.
[1993] "Means for displaying and playing" refers to the technology for displaying the response sent from the server as audio or video on the user's device.
[1994] "Means for recognizing emotions and dynamically adjusting dialogue content" refers to technology that analyzes a user's emotions and changes the dialogue content and responses to match those emotions.
[1995] "Trend data" refers to data that refers to current trends and the latest information.
[1996] "Speech synthesizer" refers to technology for converting text data into voice data and playing it back.
[1997] The embodiment of this invention is a system that allows a user to upload information about a deceased person, generates an AI model based on that information, provides real-time dialogue with the deceased, and dynamically adjusts the content of the dialogue based on the user's emotions. This will be described in detail below with specific examples.
[1998] System Overview
[1999] The system provides an interface for users to upload information about the deceased, including photos, audio files, hobbies, preferences, and personality. This information is sent to a server, which then generates an appearance model, voice model, and personality model of the deceased. The system also incorporates an emotion engine that recognizes the user's emotions and dynamically changes the content of the dialogue.
[2000] Initial Setup
[2001] The user first logs in to the system and provides information about the deceased. On the registration screen, the user uploads photos and audio files of the deceased, and enters text information about their hobbies, preferences, and personality. For example, the user might upload a photo of their grandmother working in the garden or an audio recording of their grandmother's life. After receiving this information, the device sends it to the server, which then stores the information in a database.
[2002] Data analysis and model generation
[2003] The server uses an image analysis engine to extract the facial features of the deceased from uploaded photos. For example, it uses OpenCV and facial recognition APIs to recognize features such as the position of the eyes and mouth and skin color. The server then uses a voice analysis engine to analyze the voice features of the uploaded audio file. For example, it uses voice recognition technology to extract specific patterns and voice qualities from the voice. It also uses a natural language processing engine (e.g., GPT-3) to model the personality and speaking patterns of the deceased from text information. The server combines the results of these analyses to create three models: an appearance model, a voice model, and a personality model of the deceased, and generates an AI model based on this information.
[2004] Dialogue generation and execution
[2005] To start a conversation, the user accesses the system and selects the corresponding mode. For example, they select "Dialogue with Grandma" from the app menu. The device sends this request to the server as an API request. The server uses the generated AI model to generate video and audio of the deceased person in real time to display to the user and sends them to the device. The device receives the user's voice input, converts it into text using speech recognition technology, and sends this text data to the server. The server uses a dialogue generation engine based on the text data to generate an appropriate response. For example, in response to the question, "How was your day?", the server generates a response such as, "I was taking care of the flowers in the garden today. It was a lot of fun." The device receives the response from the server, generates audio data using speech synthesis technology, and plays it back to the user as audio and video.
[2006] Emotion recognition and dialogue adjustment
[2007] The device analyzes the user's voice and video using an emotion engine (e.g., Amazon Rekognition or Microsoft Azure Emotion API) to extract the user's emotional data. The server then adjusts the content of the dialogue based on this emotional data. For example, if the user is excited, the server generates a response that reflects their emotion, such as, "Can I see that flower?"
[2008] Continuous model updates
[2009] The server periodically collects current trend data and updates the AI model, ensuring that conversations with the deceased reflect the latest information and maintain realism. The emotion engine also continuously learns, improving the accuracy of user emotion recognition.
[2010] Example operation
[2011] If a user wants to reunite with their deceased grandmother, they upload photos and audio files of the grandmother to the system, along with information about her hobbies and preferences, such as "she liked gardening" and "she was good at cooking." The server analyzes this data and generates an AI model of the grandmother. When the user selects "Interact with Grandmother," the device displays real-time video and audio of the grandmother, and the conversation begins. When the user asks, "How was your day?", the server generates a response based on the grandmother's hobbies, such as, "I was tending to the flowers in the garden today. It was so much fun." Furthermore, if the user is excited, the server dynamically adjusts the content of the conversation, such as, "Can I see those flowers?" This system strives to create a realistic conversation with a deceased person who is important to the user, providing responses that are in tune with their emotions.
[2012] The flow of the identification process in the second embodiment will be described with reference to FIG.
[2013] Step 1: Upload your information
[2014] Users log into the system and upload photos, audio files, hobbies, preferences, and personality information about the deceased.
[2015] Input: photos, audio files, hobbies, preferences, personality information
[2016] Output: Information of the deceased person sent to the server
[2017] Specific operation: The user selects various files from the interface and clicks the "Upload" button, and the device issues an API request to send this data to the server.
[2018] Step 2: Receiving and storing data
[2019] The terminal transmits the uploaded information to the server.
[2020] Input: User-uploaded information data about the deceased
[2021] Output: Information data of the deceased person stored on the server
[2022] Specific operation: The terminal splits the data received from the user and sends it to the server's API endpoint in the appropriate format. The server records the received data in the database and returns a confirmation response.
[2023] Step 3: Appearance model generation
[2024] The server analyzes the received photo and extracts facial features.
[2025] Input: Photo data
[2026] Output: Appearance model including facial recognition data
[2027] Specific operation: The server uses an image analysis engine (e.g., OpenCV) to extract features such as eyes, mouth, and skin color from the photo. Based on these features, it generates an appearance model.
[2028] Step 4: Generate a voice model
[2029] The server analyzes the uploaded audio file and extracts audio characteristics.
[2030] Input: Audio file
[2031] Output: Audio model
[2032] Specific operation: The server uses a voice analysis engine to analyze characteristics such as pitch, tone, and speed from the audio file, and generates a voice model based on this.
[2033] Step 5: Generate a personality model
[2034] The server analyzes the text information using natural language processing technology and models the personality and speech patterns of the deceased.
[2035] Input: Text data of hobbies, preferences, and personality information
[2036] Output: personality model
[2037] Specific operation: The server uses a natural language processing engine (e.g., GPT-3) to analyze text data and extract characteristic phrases and contexts, thereby generating a personality model.
[2038] Step 6: Integrating the AI model
[2039] The server integrates the appearance model, voice model, and personality model to generate an AI model.
[2040] Input: Appearance model, voice model, personality model
[2041] Output: Integrated AI model
[2042] Specific operation: The server combines the data from each model to generate a unified AI model with a consistent personality.
[2043] Step 7: Creating and Executing Real-Time Interactions
[2044] A user accesses the system using a terminal and selects an interaction mode.
[2045] Input: User's interaction initiation request
[2046] Output: Real-time dialogue display by an AI model of the deceased person
[2047] Specific operation: The device sends a request to the server, which then generates video and audio of the deceased in real time based on the generated AI model and sends them to the device. The server then receives the user's dialogue input, converts it into text using voice recognition technology, and generates a dialogue response.
[2048] Step 8: Emotion recognition and dialogue adjustment
[2049] The device analyzes the user's voice and video to recognize the user's emotions.
[2050] Input: User's audio and video data
[2051] Output: User emotion data
[2052] How it works: The device uses an emotion engine (e.g., Amazon Rekognition) to analyze the user's tone of voice and facial expressions to extract emotional data. The server then dynamically adjusts its response based on this data.
[2053] Step 9: Continuously updating the model
[2054] The server periodically collects trend data and updates the AI model.
[2055] Input: Latest trend data
[2056] Output: Updated AI model
[2057] How it works: The server collects the latest news and trend information from the internet and reflects it in the AI model. The emotion engine also continuously learns and improves the accuracy of user emotion recognition.
[2058] These are the processing steps of the program for this system. Each step works together to provide the user with a real-time conversational experience with the deceased, generating emotionally sensitive responses.
[2059] (Application example 2)
[2060] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[2061] While existing systems exist that simulate conversations with deceased loved ones, they lack realism and are unable to provide conversations that are sensitive to the user's emotions. Furthermore, the deceased's model is fixed and not continuously updated, resulting in inconsistent conversation content. Furthermore, the system lacks the ability to properly analyze the user's voice input and generate dialogue responses, leaving a need for improved user experience.
[2062] The identification processing by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for uploading information about the deceased, means for analyzing the uploaded information to generate appearance, voice, and personality models, means for generating dialogue with the deceased in real time based on the generated models, means for recognizing the user's emotions and dynamically adjusting the dialogue content based on the emotions, and means for analyzing the user's voice input and generating appropriate dialogue responses based on the analysis results. This enables highly realistic dialogue that is in tune with the user's emotions, providing an experience that more vividly revives memories of the deceased.
[2063] The "means for uploading information about the deceased" refers to a means for providing an interface for a user to send information about the deceased, such as photos, audio, and text, to the system.
[2064] "Means for analyzing uploaded information to generate appearance, voice, and personality models" refers to means for generating digital models to reproduce the face, voice, and personality of the deceased based on uploaded information about the deceased.
[2065] "Means for generating dialogue with the deceased in real time based on the generated model" refers to means for generating dialogue in real time using a model of the appearance, voice, and personality of the deceased.
[2066] "Means for transmitting the generated dialogue to the user's device and displaying / playing it back" refers to means for transmitting the generated dialogue with the deceased to the user's device, such as a smartphone or PC, and playing it back as audio and video.
[2067] "Means for recognizing the user's emotions and dynamically adjusting the dialogue content based on those emotions" refers to means for detecting the user's emotions from the tone of voice and facial expression, and appropriately changing the dialogue content in accordance with those emotions.
[2068] The "means for analyzing uploaded voice data and extracting voice characteristics" refers to a means for extracting voice tones and speaking style characteristics from uploaded voice data.
[2069] "Means for generating realistic voice using a voice synthesizer" refers to means for generating realistic voice using voice synthesis technology based on extracted voice characteristics.
[2070] "Means for analyzing a user's voice input and generating an appropriate dialogue response based on the analysis results" refers to means for analyzing the content of what the user has said and generating an optimal response based on that content.
[2071] The system for implementing this invention is configured by combining the following means. The system allows the user to upload information about the deceased, generates an AI model based on that information, and engages in real-time dialogue with the deceased. It also has the ability to recognize the user's emotions and dynamically adjust the content of the dialogue. This allows the system to provide the user with a highly realistic dialogue experience.
[2072] System configuration
[2073] server:
[2074] Store information about the deceased in a database.
[2075] Photographs, audio, and text information of the deceased are analyzed to generate appearance, voice, and personality models.
[2076] The generated models are integrated to create a generative AI model.
[2077] To analyze a user's emotions and dynamically adjust dialogue content based on the emotions.
[2078] Real-time video and audio of the deceased are generated and transmitted to the user's device.
[2079] Device:
[2080] A user interface is provided and a means is provided for uploading information about the deceased.
[2081] It captures the user's audio and video inputs and sends them to a server for sentiment analysis.
[2082] Display and play back the dialogue sent from the server.
[2083] Implementation details
[2084] Users first log in to the system and upload photos, audio, and text information about the deceased, such as their hobbies, preferences, and personality. This information is then sent to the server and stored in a database.
[2085] The server analyzes the uploaded information to generate facial recognition data, voice feature data, and personality data. Based on this data, it builds a model of the deceased person's appearance, voice, and personality, thereby completing the generative AI model.
[2086] When a user selects the dialogue mode and sends a request to start dialogue to the server, the server generates video and audio of the deceased in real time based on the generative AI model and sends them to the user's device.
[2087] The device analyzes the user's voice input and sends the text data to the server, which uses a dialogue generation algorithm to generate a response and sends it back to the device, where it is played back as audio and video.
[2088] The device analyzes the user's voice and facial expressions and uses an emotion engine to recognize the user's emotions. This recognized emotion data is sent to the server, which then dynamically adjusts the content of the conversation based on that data.
[2089] Specific examples
[2090] For example, consider a case where a user wants to reunite with a deceased parent. The user uploads photos of the parent, a recording of their voice, and information about their favorite activities and hobbies to the system. The server analyzes this information to create a model of the parent's current appearance, voice, and personality. Based on this model, the user can experience a real-time conversation with the parent. If the user says, "I was busy at work today," the system might generate a response such as, "Don't work too hard." Furthermore, if the user becomes emotional, the system will adjust the conversation based on that emotion to provide a more appropriate response.
[2091] Prompt Sentence Examples
[2092] User input: "I had a busy day at work today."
[2093] Emotion: "Fatigue" (user looks tired)
[2094] Generates response: "Don't try too hard."
[2095] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[2096] Step 1:
[2097] The user logs in to the system and uploads information about the deceased. In this step, the user provides the system with photos, audio, and text information about the deceased, such as their hobbies, preferences, and personality. The input is the digital data uploaded by the user, and the output is that this data is sent to the server and stored in the database.
[2098] Step 2:
[2099] The server analyzes uploaded photos to generate facial recognition data. It also analyzes uploaded audio files to extract voice feature data and create a personality model of the deceased person based on text information. The input is the deceased person's photo, audio, and text information, and the output is facial recognition data, voice feature data, and the generation of a personality model. Specific operations include using a facial recognition algorithm to extract facial features and using voice analysis software to extract voice features.
[2100] Step 3:
[2101] The server integrates the generated facial recognition data, voice feature data, and personality model to build a generative AI model of the deceased. This generative AI model is used to comprehensively recreate the appearance, voice, and personality of the deceased. The input is various analytical data, and the output is the generation of an integrated generative AI model. Specifically, an integration algorithm is used to combine each element of the model into one.
[2102] Step 4:
[2103] The user selects a dialogue mode and sends a dialogue start request to the server. The input is the dialogue start request by the user's operation, and the output is that this request is sent to the server.
[2104] Step 5:
[2105] The server generates video and audio of the deceased in real time based on the generative AI model and transmits them to the user's device. The input is a request to start a dialogue between the generative AI model and the user, and the output is the generation and transmission of video and audio data of the deceased. Specific operations include synthesizing the voice using a voice synthesizer and generating video in real time using CG technology.
[2106] Step 6:
[2107] The device analyzes the user's voice input and sends the resulting text data to the server. The input is the user's voice, and the output is the conversion of that voice into text and transmission to the server. Specifically, the process of converting voice into text is carried out using voice recognition software.
[2108] Step 7:
[2109] The server generates a response based on the user's voice input using a dialogue generation algorithm and sends it back to the device. The input is the user's voice text and a generative AI model, and the output is the generated response text and voice data. Specifically, this involves the process of generating appropriate dialogue content using a generative AI model and natural language processing technology.
[2110] Step 8:
[2111] The terminal plays the generated response as audio and video. The input is the response data sent from the server, and the output is the display and playback of the dialogue for the user. Specifically, the operation involves playing back audio data and displaying video.
[2112] Step 9:
[2113] The device analyzes the user's voice and facial expressions and uses an emotion engine to recognize the user's emotions. This emotion data is sent to a server. The input is real-time audio and video data, and the output is the emotion recognition results. Specifically, emotion analysis software is used to detect emotions from voice tone and facial expressions.
[2114] Step 10:
[2115] The server dynamically adjusts the dialogue content based ...
Claims
1. A means to upload information about the deceased; means for analyzing the uploaded information to generate appearance, voice, and personality models; A means for generating a dialogue with the deceased in real time based on the generated model; means for transmitting the generated dialogue to a user's terminal and displaying and playing it back; A system including:
2. 10. The system of claim 1, further comprising means for continually collecting trend data based on the deceased's appearance, voice, and personality model and updating the model based thereon.
3. A means for analyzing the uploaded voice data and extracting voice characteristics; a means for generating realistic voice using a voice synthesizer based on the extracted voice characteristics; The system of claim 1 further comprising:
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A