System

A system that collects and analyzes personal data to generate a 3D virtual model for natural interactions in a metaverse space addresses the lack of means to recreate deceased loved ones, allowing emotional relief through conversations.

JP2026033959APending Publication Date: 2026-02-27SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024137080
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-08-16
Publication Date
2026-02-27

AI Technical Summary

Technical Problem

Existing technologies lack the means to recreate the voice, manner of speaking, and appearance of deceased loved ones, making it difficult for individuals to share memories and conversations, thereby failing to ease emotional pain.

Method used

A system that collects personal data during a person's lifetime, analyzes it to generate a 3D virtual model, and uses interactive AI in a metaverse space for natural interactions, allowing users to converse with the deceased through a dedicated client application.

Benefits of technology

Enables natural conversations with deceased loved ones, providing emotional relief by recreating their voice, manner of speaking, and appearance in a virtual environment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026033959000001_ABST
    Figure 2026033959000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system, comprising: means for collecting personal data during a lifetime; means for analyzing the collected personal data to extract personal features; means for generating a 3D virtual model using the features; means for providing a 3D space in which a user can interact with the generated Metaverse virtual model; and means for using an interactive AI to make interactions in the Metaverse space natural.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] To fulfill the desire of many people to be able to converse with their deceased loved ones again, it is necessary to provide a system that allows reunions after death by recreating the voice, manner of speaking, and appearance of the deceased. In particular, there is a need to solve the problem of a lack of means to ease the emotional pain of those who lose their loved ones without being able to share what they wanted to say. [Means for solving the problem]

[0005] To solve this problem, the present invention provides a system that includes a means for collecting personal data during a person's lifetime, analyzing the collected data to extract personal characteristics, a means for generating a 3D virtual model using the characteristics, a means for providing a metaverse space in which a user can interact with the generated 3D virtual model, and a means for using an interactive AI to ensure natural interactions within the metaverse space. The system also includes a means for storing the collected episode data in a database and for the interactive AI to generate responses by referencing the episodes, and a means for the user to access the metaverse space using a dedicated client application. This allows for a natural interaction experience with the deceased, easing emotional pain.

[0006] "Personal data" is information about a specific individual collected during their lifetime, including their appearance, facial expressions, voice, speaking style, and anecdotes.

[0007] "Features" are information extracted from collected data to identify individual personalities, including facial features, vocal spectrum, speaking rhythm and tone, etc.

[0008] A "3D virtual model" is a three-dimensional digital avatar generated based on the characteristics of the deceased, and is a model that reproduces the appearance and movements of the deceased.

[0009] The "metaverse space" is a digital virtual world constructed using virtual reality technology, and is an area that users can interact with.

[0010] "Conversational AI" is artificial intelligence used to generate natural conversations and has the ability to generate appropriate responses based on user input.

[0011] "Episode data" is information related to an individual's past events or hobbies, and is data that conversational AI references to generate natural dialogue.

[0012] A "dedicated client application" is software that a user uses to access the metaverse space, and is a tool that provides the necessary interface and functions. [Brief explanation of the drawings]

[0013] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14]FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0014] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0015] First, the terms used in the following description will be explained.

[0016] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0017] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0018] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0019] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0020] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0021] [First embodiment]

[0022] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0023] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0024] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0025] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0026] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0027] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0028] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0029] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0030] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0031] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0032] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0033] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0034] This invention is a system that collects personal data from people while they are alive, analyzes the collected data to generate a 3D virtual model, and uses interactive AI in a metaverse space to allow users to experience natural conversations with the deceased.

[0035] System program generation and implementation method

[0036] Data collection

[0037] 1. A user requests data collection for a deceased person. For example, a user might say, "I want to create a virtual human of my father, so please collect his data."

[0038] 2. The device interviews and photographs the deceased. Using a camera and microphone, it records the father's facial expressions, movements, voice, and speaking style for approximately seven hours. The collected data is categorized into visual, audio, and linguistic information.

[0039] 3. The device encodes the collected data and sends it to the server using a secure protocol.

[0040] Data analysis and training of generative AI models

[0041] 1. The server analyzes the received data and uses deep learning models to extract features such as facial features, voice spectrum, speaking rhythm and tone.

[0042] 2. The server performs 3D modeling, generating a 3D virtual model that reproduces natural facial expressions and movements based on the extracted features.

[0043] 3. The server trains the conversational AI, learning from the collected voice and language data to recreate the speaking style and tone of the deceased.

[0044] Storing episode data

[0045] 1. The device collects information such as episodes, hobbies, etc. For example, it asks, "What is your favorite food?" and the father answers, "I love sushi."

[0046] 2. The device organizes this information according to a format and sends it to the server.

[0047] 3. The server stores the episode data in a database and indexes it for quick reference by the conversational AI.

[0048] Expanding into the Metaverse

[0049] 1. A user logs into the metaverse space using a dedicated client application.

[0050] 2. The server verifies the user's credentials and grants access.

[0051] 3. The server creates a 3D virtual model of the deceased person in the Metaverse space and adjusts the position of the virtual model based on the user's location information.

[0052] Specific examples

[0053] The user logs into the metaverse space and says, "Dad, it's been a while."

[0054] The device transmits this audio data to the server in real time.

[0055] The server uses conversational AI to analyze the voice data, generate an appropriate response such as, "It's been a really long time. How have you been lately?" and send it to the device.

[0056] The terminal returns the generated response to the user, allowing the user to experience a natural conversation with the deceased.

[0057] This allows users to enjoy natural conversations with their deceased loved ones within the metaverse space, providing a means to ease emotional pain.

[0058] The processing flow will be explained below.

[0059] Step 1:

[0060] The user requests the collection of data on the deceased. The request is recorded and the collection is prepared.

[0061] Step 2:

[0062] The device interviews and photographs the deceased. Using a camera and microphone, it records the deceased's facial expressions, movements, voice, and speaking style for approximately seven hours. Each piece of data is saved with a timestamp.

[0063] Step 3:

[0064] The device encodes the collected data and sends it to the server using a secure protocol, along with a checksum to verify the data's integrity and completeness.

[0065] Step 4:

[0066] The server analyzes the received data, first extracting features such as facial features, voice spectrum, and speaking rhythm and tone using deep learning algorithms.

[0067] Step 5:

[0068] The server performs 3D modeling based on the feature data, and generates a 3D virtual model using the extracted features to reproduce the natural facial expressions and movements of the deceased.

[0069] Step 6:

[0070] The server trains the conversational AI, using the collected voice and language data to train the AI ​​model to reproduce the speech style and tone of the deceased, specifically understanding the voice patterns and flow of the conversation.

[0071] Step 7:

[0072] The device separately collects information such as anecdotes and hobbies. For example, during an interview, it might ask, "What is your favorite food?" and record the answer.

[0073] Step 8:

[0074] The device sends episode data to the server, where the data is categorized into episodes, hobbies, favorite things, etc. and organized according to a format.

[0075] Step 9:

[0076] The server stores the episode data in a database, where it is indexed and made available for quick reference by the conversational AI.

[0077] Step 10:

[0078] A user logs into the metaverse space using a dedicated client application, and the login information is sent to the server via authentication means.

[0079] Step 11:

[0080] The server verifies the user's credentials and grants access. If authentication is successful, access is established within the metaverse space.

[0081] Step 12:

[0082] The server creates a 3D virtual model of the deceased person in the Metaverse space based on the user's location information, and places the model in an appropriate location, taking into account the user's line of sight and location.

[0083] Step 13:

[0084] The user talks to the deceased, for example, saying, "Dad, how are you?"

[0085] Step 14:

[0086] The device transmits the user's voice in real time to a server, where the voice data is recorded and then stored for analysis.

[0087] Step 15:

[0088] The server uses conversational AI to analyze the voice data and generate an appropriate response, such as "I've been doing well. How have you been lately?"

[0089] Step 16:

[0090] The server sends the generated response to the terminal, which then outputs the response to the user.

[0091] This allows users to experience natural interactions with the deceased within the metaverse space.

[0092] Example 1

[0093] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0094] There is a need for a system that allows users to communicate with deceased individuals and ease their emotional pain by recreating memories and conversations with them. However, existing technologies are insufficient in recreating the deceased, making it difficult to achieve natural conversations. To solve this problem, a system that accurately recreates the characteristics of the deceased and allows users to experience natural conversations is needed.

[0095] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0096] In this invention, the server includes: a means for a user to request data collection of the deceased; a means for recording the deceased's facial expressions, movements, voice, and speaking style using a camera and microphone; a means for encoding the recorded data and transmitting it to the server using a secure protocol; a means for the server to analyze the received data using a deep learning model and extract features; a means for generating a 3D virtual model based on the extracted features; a means for training an interactive AI using voice data and language data; a means for collecting information such as episodes and hobbies and storing it in a database; a means for a user to access the metaverse space using a dedicated client application; a means for the server to create a 3D virtual model of the deceased in the metaverse space and adjust it based on the user's location information; and a means for realizing natural conversation using the interactive AI, allowing users to enjoy natural conversations with the deceased in the metaverse space.

[0097] "User" refers to an individual who utilizes the system to collect data about or interact with a deceased person.

[0098] "Means for requesting data collection" refers to a function that allows a user to request data collection of a deceased person from the system.

[0099] "Camera and microphone" refers to visual and audio input devices used to record the facial expressions, movements, voice, and speaking style of the deceased.

[0100] "Means of recording" refers to the ability to use cameras and microphones to digitally record the facial expressions, movements, voice, and speaking style of the deceased.

[0101] "Encoding" refers to the process of compressing and converting recorded data into a format that can be transmitted efficiently.

[0102] "Secure protocol" refers to a communication protocol for ensuring the security of data communications, and includes SSL and TLS.

[0103] A "deep learning model" refers to an artificial intelligence technology that uses neural networks to analyze data and extract features.

[0104] "Feature extraction" refers to finding important patterns or attributes in data and using them for analysis.

[0105] A "3D virtual model" refers to a digital model that reproduces the features of the deceased and displays them in three-dimensional space.

[0106] "Conversational AI" refers to artificial intelligence technology that can mimic natural human conversation.

[0107] "Anecdotes and hobbies" refers to information that reveals the deceased's personal experiences and interests.

[0108] A "database" refers to an information management system that stores data in a structured format that makes it easy to search and reference.

[0109] "Client Application" refers to software that a user uses to access the metaverse space.

[0110] "Metaverse space" refers to a virtual reality environment constructed digitally, including a space that users can interact with.

[0111] "Location information" refers to data indicating the current coordinates or location of the user and virtual model.

[0112] This invention is a system that collects personal data from people while they are alive, analyzes the collected data to generate a 3D virtual model, and allows users to experience natural conversations with the deceased using interactive AI in a metaverse space.

[0113] Data collection

[0114] 1. A user requests data collection for a deceased person. For example, they might request, "I would like to create a virtual model of my father, so please collect his data." This request is made via a dedicated web portal or application.

[0115] 2. The device interviews and films the deceased. It uses a camera and microphone to record the father's facial expressions, movements, voice, and speaking style for approximately seven hours. The specific hardware used includes a high-resolution camera and a directional microphone. The collected data is categorized into visual, audio, and language information.

[0116] 3. The device encodes the collected data and sends it to the server using a secure protocol. A data encoder is used for this encoding, and the SSL / TLS protocol is used to ensure secure data transmission. Encoding is done using a format such as H.264 or FLAC.

[0117] Data analysis and training of generative AI models

[0118] 1. The server analyzes the received data and uses a deep learning model (e.g., TENSORFLOW (registered trademark) or PyTorch) to extract features such as facial features, voice spectrum, speaking rhythm and tone.

[0119] 2. The server performs 3D modeling, using Blender, Maya, or other software to generate a 3D virtual model that reproduces natural facial expressions and movements based on the extracted features.

[0120] 3. The server trains the conversational AI, learning from the collected voice and language data and using language models such as GPT-3 (registered trademark) and BERT to reproduce the speaking style and tone of the deceased, enabling natural conversation.

[0121] Storing episode data

[0122] 1. The device collects information such as episodes, hobbies, etc. For example, it asks, "What is your favorite food?" and the father answers, "I like sushi."

[0123] 2. The terminal organizes the collected episode information according to a format and sends it to the server. This organization is done using a formatting tool.

[0124] 3. The server stores the episode data in a database and indexes it for quick reference by the conversational AI. Database systems used include MySQL® and MongoDB.

[0125] Expanding into the Metaverse

[0126] 1. A user logs into the Metaverse space using a dedicated client application, including the client application for Oculus Rift.

[0127] 2. The server verifies the user's authentication information and grants access. The authentication system used is OAuth 2.0 or similar.

[0128] 3. The server creates a 3D virtual model of the deceased person in the Metaverse space, and adjusts the virtual model's position appropriately using Unity or Unreal Engine based on the user's location information.

[0129] Specific examples

[0130] 1. The user logs into the metaverse space and says, "Dad, it's been a while."

[0131] 2. The device sends this audio data to the server in real time.

[0132] 3. The server uses conversational AI to analyze the voice data, generate an appropriate response such as, "It's been a really long time. How have you been lately?" and send it to the device.

[0133] 4. The device returns the generated response to the user, allowing the user to experience a natural conversation with the deceased.

[0134] An example of a specific prompt is, "Dad, what did you do today?"

[0135] This allows users to enjoy natural conversations with their deceased loved ones within the metaverse space, providing a means to ease emotional pain.

[0136] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0137] System program processing flow

[0138] Data Collection Steps

[0139] Step 1:

[0140] A user requests data collection of a deceased person using a dedicated web portal or application. As input, the user enters details such as the name of the deceased, their relationship, and the information they wish to collect. As output, a request form is sent to the server.

[0141] Step 2:

[0142] The server receives the user's request and prepares the data collection. The user's request information is transmitted to the server as input. The collection plan is sent to the terminal as output.

[0143] Step 3:

[0144] The device uses a camera and microphone to record the facial expressions, movements, voice, and speaking style of the deceased. The camera and microphone capture the deceased's visual, audio, and linguistic information as input. The captured data is then stored on the device as output. Specifically, the camera captures high-resolution video and the microphone records audio.

[0145] Step 4:

[0146] The data recorded by the device is encoded and sent to the server using a secure protocol. Visual, audio, and language information collected as input is encoded. Encrypted data is sent to the server as output. Specifically, the data is encoded into H.264 or FLAC format and sent using the SSL / TLS protocol.

[0147] Steps for Data Analysis and Generative AI Model Training

[0148] Step 5:

[0149] The server analyzes the received data using a deep learning model. Encoded data securely transmitted to the server is provided as input. Feature-extracted data is generated as output. Specifically, TensorFlow and PyTorch are used to analyze features such as facial features, voice spectrum, speaking rhythm and tone.

[0150] Step 6:

[0151] The server generates a 3D virtual model based on the extracted features. The analyzed feature data is given as input. The 3D virtual model is generated as output. Specifically, the model is constructed using Blender or Maya, and skeletal animation is implemented.

[0152] Step 7:

[0153] The server trains the conversational AI. The collected speech and language data is used as input. The output is a trained conversational AI model that improves its ability to reproduce the speech style and tone of the deceased. Specifically, training is performed using GPT-3 and BERT to enhance the model's ability to mimic natural conversation.

[0154] Steps for storing episode data

[0155] Step 8:

[0156] The device collects information from the deceased, such as episodes and hobbies. The user or interviewer asks questions as input, and the deceased's responses are collected as voice data. The output is the collected episode information organized. Specifically, the system works by collecting a situation in which the interviewer asks, "What is your favorite food?" and the deceased answers, "I like sushi."

[0157] Step 9:

[0158] The episode information collected by the device is organized according to a format and sent to the server. Unorganized episode information is given as input, and formatted data is sent to the server as output. Specifically, the data is organized into JSON or XML format using a formatting tool and sent to the server.

[0159] Step 10:

[0160] The server stores the episode data in a database and indexes it for quick reference by the conversational AI. Formatted episode data is given as input. The output is episode information stored in a database and indexed. Specifically, data is organized and indexed using MySQL or MongoDB.

[0161] Steps for expanding into the metaverse space

[0162] Step 11:

[0163] A user logs into the metaverse space using a dedicated client application. The user's authentication information (username and password) is used as input. The output is the initiation of a login session. Specifically, the client application takes the user's credentials as input and sends them to the authentication system.

[0164] Step 12:

[0165] The server verifies the user's authentication information and grants access. User credentials are provided as input. An authentication result is generated as output. Specifically, an authentication system such as OAuth 2.0 is used, and a session begins if authentication is successful.

[0166] Step 13:

[0167] The server creates a 3D virtual model of the deceased person in the Metaverse space and adjusts it based on the user's location information. The user's current location information and 3D model information are provided as input. The output is a 3D model that has been adjusted to correspond to the user. Specifically, Unity or Unreal Engine is used to place the virtual model at a specified location in the Metaverse space.

[0168] Specific examples

[0169] 1. A user logs in to the metaverse space and says, "Dad, it's been a while." The user's voice is captured by a microphone as input. The voice data is sent to the server as output.

[0170] 2. The device sends this audio data in real time to the server. The input is the captured audio data. The output is the audio data that is transferred to the server.

[0171] 3. The server uses conversational AI to analyze the voice data, generate an appropriate response such as "It's been a really long time. How have you been lately?" and send it to the device. The user's voice data is given as input, and the response voice data is generated as output.

[0172] 4. The device returns the generated response to the user. The response audio data sent from the server is given as input. The response is played back to the user through the speaker as output.

[0173] An example of a specific prompt is, "Dad, what did you do today?"

[0174] (Application example 1)

[0175] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0176] Although existing technologies exist for systems that collect personal data from the deceased and recreate natural interactions with the deceased, they lack specific environments and memory-based interactions to improve the user's emotional satisfaction.This invention aims to provide a virtual memorial service that allows users to feel a deeper emotional connection by interacting with the deceased in a specific memorial store.

[0177] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0178] In this invention, the server includes a means for collecting personal data of the deceased while they are alive, a means for analyzing the collected data to extract personal characteristics, a means for generating a 3D virtual model using the characteristics, a means for providing a virtual space in which the user can interact with the generated 3D virtual model, a means for using interactive AI to make the interaction in the virtual space natural, a means for placing a specific memorial store in the virtual space in which the user can move, and a means for placing objects and music related to the deceased in the memorial store and realizing interaction based on memorable episodes. This allows the user to interact with the deceased in a specific environment, further strengthening the emotional connection.

[0179] "Personal data" refers to all information, including audio, video, and linguistic information, captured during a person's lifetime.

[0180] "Features" are information extracted from collected personal data, such as facial features, voice spectrum, speaking rhythm and tone.

[0181] A "3D virtual model" is a three-dimensional human model generated in a virtual space based on collected features.

[0182] A "virtual space" is a digitally recreated environment that a user can access.

[0183] "Conversational AI" is artificial intelligence that is trained to engage in natural conversations using collected voice and language data.

[0184] A "memorial store" is a specific area within the virtual space where objects, music, and stories related to the deceased are interactively arranged.

[0185] "Object" refers to any object or item placed in a virtual space.

[0186] "Music" refers to melodies and songs related to the deceased, and is sound information played within the virtual space.

[0187] "Memorable episodes" are information that provides conversations based on the deceased's hobbies, favorite things, and past events.

[0188] A "dedicated client application" is a specially designed software application used by a user to access a virtual space.

[0189] System program generation

[0190] To realize this invention, we start by collecting data on individuals during their lifetime. This process uses hardware such as a camera (e.g., OpenCV) and an audio recorder. The collected data is categorized into visual, audio, and language information.

[0191] The server receives the collected data and uses a deep learning model (e.g., TensorFlow) to extract features such as facial features, voice spectrum, speaking rhythm, and tone. Based on this, a 3D virtual model is generated. It also trains a conversational AI model based on the episode data and develops algorithms to achieve natural conversations. This model is then used during subsequent conversations.

[0192] Meanwhile, a specific memorial store will be set up in the virtual space, and will be stocked with objects and music related to the deceased. This setting is important for users to feel a deep emotional connection with the deceased. The memorial store will be rendered using Unity.

[0193] Processing procedure description

[0194] Data collection:

[0195] The server interviews and films the deceased. It uses a camera and microphone to record the deceased's facial expressions, movements, voice, and speaking style. For example, if a user requests, "I want to create a virtual human of my father, so please collect data," the device encodes the collected data and sends it to the server using a secure protocol.

[0196] Data analysis and generative AI model training:

[0197] The server analyzes the received data and uses deep learning models to extract features, which are then used to generate a 3D virtual model that reproduces natural facial expressions and movements. Furthermore, the conversational AI is trained to reproduce the speech style and tone of the deceased.

[0198] Episode data storage:

[0199] To ensure a more natural interaction when users visit the memorial store, the device collects information about the deceased's anecdotes and hobbies and stores it in a database, which the conversational AI can quickly reference.

[0200] Virtual Space Deployment:

[0201] A user logs into the virtual space using a dedicated client application, and after the server verifies the user's authentication information, a 3D virtual model of the deceased person appears in the virtual space. For example, a user can log into a "Memorial Cafe" and enjoy interacting with the 3D model of the deceased. The cafe is furnished with objects and music based on the preferences of the deceased.

[0202] Examples of concrete examples and prompts

[0203] As a concrete example, if a user logs in to Memorial Cafe and says, "Mom, you liked this cafe, didn't you?", the conversational AI will respond, "Yes, I really liked it. I love the atmosphere."

[0204] Example prompt sentence:

[0205] User: Mom, do you like this cafe?

[0206] Deceased Model: Yes, I love it. I love the atmosphere here.

[0207] This allows users to experience natural conversation with the deceased within the memorial store, further strengthening their emotional connection.

[0208] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0209] Step 1:

[0210] Collecting personal data before death.

[0211] Input: User requests data collection

[0212] How it works: The device uses a camera and microphone to record the facial expressions, movements, voice, and speaking style of the deceased. For example, based on a user's request to "create a virtual model of my father," the device collects data for approximately seven hours.

[0213] Output: A dataset containing visual, audio, and linguistic information

[0214] Step 2:

[0215] The collected data is sent to a server using a secure protocol.

[0216] Input: Dataset collected in step 1

[0217] Specific operation: The device encodes the data and sends it to the server using a secure protocol (e.g., SSL / TLS).

[0218] Output: The encoded data reaches the server

[0219] Step 3:

[0220] The server parses the data it receives.

[0221] Input: Encoded data

[0222] How it works: The server uses a deep learning model (e.g., TensorFlow) to decode the data and extract features such as facial features, voice spectrum, speaking rhythm and tone.

[0223] Output: A dataset of extracted facial feature points, voice spectrum, speaking rhythm and tone.

[0224] Step 4:

[0225] Generate a 3D virtual model.

[0226] Input: Feature data extracted in step 3

[0227] Specific operation: The server uses Unity to generate a 3D virtual model based on the extracted feature data, reproducing natural facial expressions and movements.

[0228] Output: Generated 3D virtual model

[0229] Step 5:

[0230] Train a conversational AI model.

[0231] Input: Collected speech and language data

[0232] How it works: The server uses collected voice and language data to train the conversational AI so that it can converse naturally, and uses deep learning algorithms to build a model that reproduces the speech style and tone of the deceased.

[0233] Output: A trained conversational AI model

[0234] Step 6:

[0235] Episode data is stored and indexed in a database.

[0236] Input: Information about episodes and hobbies collected from users

[0237] Specific operation: The device organizes information about episodes and hobbies, stores it in a database, and indexes it for quick reference by the conversational AI.

[0238] Output: Indexed episode data

[0239] Step 7:

[0240] A user logs into the virtual space using a dedicated client application.

[0241] Input: User credentials

[0242] How it works: A user logs into a virtual space using a dedicated client application. The server verifies the authentication information and grants the user access.

[0243] Output: The user can access the virtual space.

[0244] Step 8:

[0245] A memorial store will be placed in the virtual space, and a 3D virtual model of the deceased person will appear.

[0246] Input: User's location, 3D virtual model generated in step 4

[0247] Specific operation: The server places a memorial store in the virtual space and makes a 3D virtual model of the deceased person appear based on the user's location information.

[0248] Output: 3D virtual model of the deceased placed in a memorial store

[0249] Step 9:

[0250] The user engages in natural dialogue with conversational AI.

[0251] Input: User voice input, the conversational AI model trained in step 5, and episode data from step 6

[0252] How it works: When a user speaks to a model of the deceased person, the device sends the voice data in real time to the server. The server uses a conversational AI model to generate a response and sends it to the device. The device then returns the response to the user, realizing a natural dialogue.

[0253] Output: A natural dialogue experience between the user and conversational AI

[0254] Specific examples

[0255] For example, if a user logs in to Memorial Cafe and says, "Mom, do you like this cafe?" the conversational AI will respond, "Yes, I love it. I love the atmosphere here."

[0256] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0257] This invention combines a system that collects personal data from a person's lifetime, analyzes the collected data to generate a 3D virtual model, and enables users to experience natural conversations with the deceased using interactive AI within the metaverse space, with an emotion engine that recognizes the user's emotions.

[0258] System program generation and implementation method

[0259] Generate 3D virtual models from data collection

[0260] 1. A user requests the collection of data on a deceased person. The request is recorded and preparations for collection are made.

[0261] 2. The device interviews and photographs the deceased. Using a camera and microphone, it records the deceased's facial expressions, movements, voice, and speaking style for approximately seven hours. Each piece of data is saved with a timestamp.

[0262] 3. The device encodes the collected data and sends it to the server using a secure protocol, along with a checksum to verify the data's integrity and completeness.

[0263] 4. The server analyzes the received data, using deep learning models to extract features such as facial features, voice spectrum, and speaking rhythm and tone.

[0264] 5. The server performs 3D modeling based on the feature data. The extracted features are used to generate a 3D virtual model that reproduces the natural facial expressions and movements of the deceased.

[0265] 6. The server trains the conversational AI. Using the collected voice and language data, the AI ​​model is trained to reproduce the speaking style and tone of the deceased, specifically understanding the deceased's voice patterns and conversational flow.

[0266] 7. The device collects information about episodes, hobbies, etc. For example, during an interview, it might ask, "What is your favorite food?" and record the answer.

[0267] 8. The device sends the episode data to the server. The data is categorized into episodes, hobbies, favorite things, etc. and organized according to a format.

[0268] 9. The server stores the episode data in a database, where it is indexed and made available for quick reference by the conversational AI.

[0269] Expanding into the Metaverse

[0270] 1. A user logs into the metaverse space using a dedicated client application. The login information is sent to the server via authentication methods.

[0271] 2. The server verifies the user's credentials and grants access. If authentication is successful, access is established within the metaverse space.

[0272] 3. The server creates a 3D virtual model of the deceased person in the Metaverse space based on the user's location information, placing it in an appropriate position taking into account the user's line of sight and location.

[0273] 4. The user talks to the deceased, for example, "Dad, how are you?"

[0274] 5. The device transmits the user's voice in real time to the server, where it is recorded and stored for analysis.

[0275] 6. The server uses conversational AI to analyze the voice data and generate an appropriate response, such as "I've been doing well. How have you been lately?"

[0276] 7. The server sends the generated response to the terminal, which then outputs the response to the user.

[0277] Incorporating an emotion engine

[0278] 1. The device captures the user's facial expressions and tone of voice in real time using a camera and microphone.

[0279] 2. The device sends the captured data to the emotion engine, which then recognizes the user's emotion based on this data. For example, if a smile is detected, it will recognize that the user is feeling "joy."

[0280] 3. The server receives the emotion data sent from the emotion engine and reflects it in the conversational AI. For example, if the user is recognized as sad, the conversational AI generates a response such as "Cheer up."

[0281] Specific examples

[0282] For example, a user logs into the metaverse space and says, "Dad, it's been a while."

[0283] The device transmits this audio data to the server in real time.

[0284] The emotion engine analyzes and recognizes that the user's tone of voice is sad.

[0285] The server uses conversational AI to analyze the voice data and, referring to the analysis results of the emotion engine, generates an appropriate response such as, "It's been a while. What's wrong? You seem down."

[0286] The server sends the generated response to the terminal, which then speaks the response to the user.

[0287] This allows users to have more emotionally enriching interactions with deceased loved ones within the metaverse space.

[0288] The processing flow will be explained below.

[0289] Step 1:

[0290] The user requests the collection of data on the deceased. The request is confirmed and preparations for collection are made.

[0291] Step 2:

[0292] The device interviews and photographs the deceased. Using a camera and microphone, it records the deceased's facial expressions, movements, voice, and speaking style for approximately seven hours. The collected data is then saved with a timestamp.

[0293] Step 3:

[0294] The device encodes the collected data and transmits it to the server using a secure protocol, along with a checksum to ensure data integrity and completeness.

[0295] Step 4:

[0296] The server analyzes the received data and uses deep learning models to extract features such as facial features, voice spectrum, speaking rhythm and tone.

[0297] Step 5:

[0298] The server performs 3D modeling based on the feature data, and generates a 3D virtual model using the extracted features to reproduce the natural facial expressions and movements of the deceased.

[0299] Step 6:

[0300] The server trains the conversational AI. Using the collected voice and language data, the AI ​​model is trained to reproduce the speaking style and tone of the deceased. Specifically, it analyzes and learns from the deceased's voice patterns and conversational flow.

[0301] Step 7:

[0302] The device collects information such as episodes and hobbies. For example, it asks, "What is your favorite food?" and records the answer.

[0303] Step 8:

[0304] The device sends episode data to the server, which organizes the data into categories such as episodes, hobbies, and favorite things.

[0305] Step 9:

[0306] The server stores the episode data in a database, where it is indexed and prepared for fast and efficient reference by the conversational AI.

[0307] Step 10:

[0308] A user logs into the metaverse space using a dedicated client application, and the login information is sent to the server via authentication means.

[0309] Step 11:

[0310] The server verifies the user's credentials and grants access. If authentication is successful, access is established within the metaverse space.

[0311] Step 12:

[0312] The server creates a 3D virtual model of the deceased person in the Metaverse space based on the user's location information, and adjusts the position of the virtual model based on the user's line of sight and location.

[0313] Step 13:

[0314] The user talks to the deceased, for example, saying, "Dad, it's been a long time."

[0315] Step 14:

[0316] The device transmits the user's voice in real time to a server, where the voice data is recorded and stored for analysis.

[0317] Step 15:

[0318] The server uses conversational AI to analyze the voice data and generate an appropriate response, such as "It's been a while, how have you been?"

[0319] Step 16:

[0320] The device captures the user's facial expressions and tone of voice in real time using a camera and microphone.

[0321] Step 17:

[0322] The device captures facial and voice data and sends it to the emotion engine, which analyzes this data and recognizes the user's emotions.

[0323] Step 18:

[0324] The server receives the emotion data sent from the emotion engine and reflects it in the conversational AI. For example, if the user looks sad, the conversational AI will generate a response such as "Cheer up, what happened recently?"

[0325] Step 19:

[0326] The server sends the generated response to the terminal, which then provides an audio output to the user.

[0327] This allows users to interact with deceased loved ones in a more emotionally rich and natural way within the metaverse space.

[0328] Example 2

[0329] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0330] Conventional systems have had difficulty using data from deceased people to recreate conversations with them in a virtual space. Furthermore, they lacked the ability to generate responses based on the user's emotions, making the conversations less emotionally rich and natural.

[0331] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for collecting personal data before death, means for analyzing the collected data to extract personal characteristics, means for generating a 3D virtual model using the characteristics, means for providing a virtual space in which the user can interact with the generated 3D virtual model, means for using interactive artificial intelligence to make the interaction in the virtual space natural, and means for recognizing the user's emotions and flexibly adjusting the response content. This allows the user to experience an emotionally rich and natural interaction with the deceased person in the virtual space.

[0332] "Personal data" is information about a person that is collected during their lifetime, including photographs, videos, audio files, and written material.

[0333] "Feature extraction means" refers to a process or technique for identifying and analyzing individual characteristics or features from collected data.

[0334] A "3D virtual model" is a three-dimensional digital representation of an individual's appearance, gestures, movements, etc., based on collected data and extracted features.

[0335] "Virtual space" refers to a virtual environment that users can access through a computer or dedicated client application, including the metaverse and VR space.

[0336] "Conversational artificial intelligence" refers to a computer program or system that can use natural language analysis and generation techniques to engage in natural conversations with humans.

[0337] "Means for recognizing emotions and flexibly adjusting response content" refers to a function that analyzes the user's facial expressions, tone of voice, and other emotional indicators in real time and dynamically changes the response content based on the results.

[0338] "Episode data" refers to information based on specific events, hobbies, or specific memories related to the deceased.

[0339] "Dedicated client application" refers to software specifically designed to access virtual spaces and interface with interactive AI and 3D virtual models.

[0340] MODE FOR CARRYING OUT THE INVENTION

[0341] This invention is a system that collects personal data from a person's life and analyzes it to generate a 3D virtual model. This system uses conversational artificial intelligence to allow users to experience natural interactions with the deceased in the metaverse space. It also combines an emotion engine that recognizes the user's emotions to realize more emotionally rich interactions.

[0342] Hardware and software used

[0343] Hardware:

[0344] Camera: Used to record the facial expressions and movements of the deceased.

[0345] Microphone: Used to record the voice and speaking style of the deceased.

[0346] Devices (e.g., computers, smartphones): used for data collection and transmission.

[0347] Server: Used for data processing, analysis, storage, AI training, and virtual space management.

[0348] software:

[0349] Dedicated client application: Software that allows users to access the virtual space.

[0350] Data collection software: collects data from the camera and microphone, encodes it, and sends it to a server.

[0351] Deep learning models: Algorithms for data analysis and feature extraction.

[0352] 3D modeling tools (e.g. Blender): Generate 3D virtual models based on collected feature data.

[0353] Conversational Artificial Intelligence (AI): Enables natural dialogue with users.

[0354] Emotion engine: Analyzes user emotions and reflects them in the dialogue AI.

[0355] Examples of concrete examples and prompts

[0356] For example, suppose a user logs into the metaverse space using a dedicated client application and says, "Dad, it's been a while." The specific processing at this time is as follows.

[0357] 1. The device sends this voice data to the server in real time.

[0358] 2. The server uses conversational artificial intelligence to analyze the voice data and generates an appropriate response by referring to the analysis results of the emotion engine.

[0359] 3. For example, the emotion engine analyzes and recognizes that the user's tone of voice is sad, and the server generates a response such as, "It's been a while, what's wrong? You seem depressed."

[0360] 4. The server sends the generated response to the terminal, which then speaks the response to the user.

[0361] In this way, users can have emotionally enriching interactions with deceased loved ones within the metaverse space.

[0362] Example prompt sentence:

[0363] "Dad, how have you been?"

[0364] "How are you feeling today?"

[0365] "I want to talk about recent events."

[0366] By inputting these prompt sentences, the conversational AI can understand the user's intentions and emotions and generate the most appropriate response.

[0367] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0368] Step 1: (Data collection request)

[0369] The user requests the collection of data on the deceased, specifying specific items to be collected, such as photos, videos, and audio files of the deceased.

[0370] The device records the request and begins preparations for collection, preparing the camera and microphone, and confirming the data format to be collected.

[0371] Input: User request.

[0372] Output: Collection readiness status.

[0373] Step 2: (Interview and filming)

[0374] The device uses a camera and microphone to interview and film the deceased, and the data collected includes facial expressions, movements, voice, and speaking style.

[0375] The device stores all data with a timestamp, splitting it into video clips every 10 minutes, for example.

[0376] Input: Real-time data from camera and microphone.

[0377] Output: Collected data with timestamp.

[0378] Step 3: (Data encoding and transmission)

[0379] The device encodes the collected data and transmits it to a server using a secure protocol (e.g., HTTPS).

[0380] The terminal generates a checksum to verify the integrity and completeness of the data and also sends this to the server.

[0381] Input: Collected data before encoding.

[0382] Output: The encoded data and a checksum.

[0383] Step 4: (Data Analysis)

[0384] The server then analyzes the received data using a deep learning model that extracts features such as facial features, voice spectrum, and speaking rhythm and tone.

[0385] The server specifically analyzes the deceased person's unique movements and expressions and stores them in a database.

[0386] Input: The encoded data.

[0387] Output: Extracted feature data.

[0388] Step 5: (3D modeling)

[0389] The server generates a 3D virtual model based on the extracted feature data, carefully adjusting it to reproduce the natural facial expressions and movements of the deceased.

[0390] The server uses a 3D modeling tool (e.g. Blender) to see in real time what needs to be fixed.

[0391] Input: Extracted feature data.

[0392] Output: 3D virtual model.

[0393] Step 6: (Training the conversational AI)

[0394] The server uses the collected speech and language data to train the conversational AI, specifically learning patterns to replicate the speech and narration of the deceased.

[0395] The server mimics voice patterns and adjusts them to create a natural conversational flow.

[0396] Input: Audio and language data.

[0397] Output: A trained conversational AI model.

[0398] Step 7: (Collect episode information)

[0399] Users provide information about the deceased, such as anecdotes, hobbies, and interests, in the form of an interview.

[0400] The device records this information and also saves it as text data.

[0401] Input: The answer given by the user.

[0402] Output: Episode data.

[0403] Step 8: (Submit episode data)

[0404] The episode data collected by the device is organized according to a format and sent to the server.

[0405] The device classifies the data (episodes, hobbies, interests) before sending it.

[0406] Input: Episode data.

[0407] Output: Organized episode data.

[0408] Step 9: (Storing episode data)

[0409] The server indexes and stores the received episode data in a database.

[0410] The server applies a search algorithm to allow quick access to the data.

[0411] Input: Organized episode data.

[0412] Output: Indexed episode data in a database.

[0413] Step 10: (Login and Authentication)

[0414] A user logs into the virtual space using a dedicated client application, and the login information is sent from the client to the server.

[0415] The server verifies the user's credentials and grants access.

[0416] Input: User login information.

[0417] Output: Authentication check and access granted.

[0418] Step 11: (Placing the virtual model)

[0419] The server creates a 3D virtual model of the deceased person in a virtual space based on the user's location information.

[0420] The server places the virtual model in an appropriate position, taking into account the user's line of sight and position.

[0421] Input: User's location.

[0422] Output: Positioning information of the virtual model in the virtual space.

[0423] Step 12: (User interaction)

[0424] The user talks to the deceased, for example, asking, "Dad, how are you?"

[0425] The device transmits this audio data to the server in real time.

[0426] The server uses conversational artificial intelligence to analyze the voice data and generates an appropriate response by referring to the analysis results of the emotion engine.

[0427] Input: User's question, voice data.

[0428] Output: A response from the conversational AI.

[0429] Step 13: (Capturing the user's facial expressions and voice)

[0430] The device captures the user's facial expressions and tone of voice in real time using a camera and microphone.

[0431] The device transmits the captured data to the emotion engine in real time.

[0432] Input: User's facial expression data and tone of voice.

[0433] Output: Emotion data sent to the emotion engine.

[0434] Step 14: (Analyze emotion data)

[0435] The server analyzes the user's emotions based on the data sent from the emotion engine. For example, it recognizes "joy" by detecting a smile.

[0436] Input: User emotion data from the emotion engine.

[0437] Output: Analyzed user emotion information.

[0438] Step 15: (Generating and reflecting dialogue responses)

[0439] The server generates a response from the conversational AI based on the emotional data. For example, if the server recognizes that the user is sad, it generates a response such as "Cheer up."

[0440] The server sends the generated response to the terminal, which then outputs it to the user.

[0441] Input: Parsed user emotion information.

[0442] Output: AI response based on emotion.

[0443] (Application example 2)

[0444] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0445] In modern society, people often seek emotional connections through conversations with deceased loved ones, but systems that can provide such experiences are not yet widespread. Furthermore, there are no systems incorporating conversational AI that can recognize and reflect emotions in real time, making it difficult to understand the user's emotions and provide natural conversations accordingly. Furthermore, for users to naturally converse with deceased loved ones in virtual spaces, it is necessary to manage and analyze episode data and emotional data, and an efficient system for this purpose is needed.

[0446] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0447] In this invention, the server includes means for collecting personal data before death, means for analyzing the collected data to extract personal characteristics, means for generating a 3D virtual model using the characteristics, means for providing a virtual space in which a user can interact with the generated 3D virtual model, means for using interactive AI to make the interaction in the virtual space natural, means for capturing and analyzing the user's voice and emotions in real time, and means for generating a response that reflects the analyzed emotional data, thereby enabling a user to have an emotional experience through interaction with the deceased.

[0448] "Personal data" refers to information provided by an individual during their lifetime, such as their voice, facial expressions, movements, and speaking style.

[0449] "Features" refers to identifying information such as facial features, voice spectrum, speaking rhythm and tone extracted from collected personal data.

[0450] "3D Virtual Model" means a three-dimensional model of an individual displayed in a virtual space, generated based on collected characteristics.

[0451] "Virtual space" refers to a virtual reality environment accessible to a user.

[0452] "Conversational AI" refers to an artificial intelligence system that uses natural language processing to engage in dialogue with users.

[0453] "Means for capturing and analyzing voice and emotions in real time" refers to the function of collecting the user's voice and facial expressions in real time using a voice input device or camera, and analyzing them using an emotion analysis engine.

[0454] "Emotion data" refers to information about emotions analyzed from the user's tone of voice, facial expressions, etc.

[0455] "Means for generating a response" refers to an appropriate response that corresponds to the user's emotions, which is generated by the conversational AI based on the analyzed emotional data.

[0456] This invention is a system that allows a user to have natural conversations with a deceased person in a virtual space based on personal data collected by the user while he or she was alive. An embodiment of this system will be described in detail below.

[0457] First, the user uses a device to collect data on the deceased. Data such as the deceased's facial expressions, movements, voice, and speaking style is recorded using a camera and microphone for approximately seven hours. This collected data is then sent to a server using a secure protocol. The server then uses a deep learning model to analyze the data and extract individual features. Based on the extracted features, a realistic and natural-looking 3D virtual model is then generated.

[0458] The generated 3D virtual model is displayed in a virtual space accessible to users. Users access this virtual space using a dedicated client application on their smartphone or head-mounted display. When the user speaks to the deceased, their voice is transmitted in real time to a server, where conversational AI analyzes the voice and generates an appropriate response.

[0459] The system also incorporates an emotion engine that captures and analyzes the user's voice and facial expressions in real time. The emotion engine analyzes emotions from the user's tone of voice and facial expressions and reflects them in the responses generated by the conversational AI. This allows the user to experience a more natural and emotional dialogue. For example, if a user says, "Dad, it's been a while," the emotion engine will analyze that the user's tone of voice sounds sad, and the conversational AI will generate a response such as, "It's been a while. What's wrong? You seem down."

[0460] The hardware used includes a smartphone, head-mounted display, camera, and microphone, while the software used includes Unity (a 3D game engine), TensorFlow (a deep learning framework for emotion analysis), Google® Cloud Speech-to-Text API, and Azure® Cognitive Services (conversational AI).

[0461] An example prompt is:

[0462] "User prompt"

[0463] We would like to request the development of an application that provides an emotional experience through interaction with the deceased, such as:

[0464] The user collects data about the deceased and conducts an interview using a camera and microphone. The collected data is sent to a server, and a 3D virtual model is generated using deep learning. The user logs in to the virtual space and begins a dialogue. The emotion engine analyzes the user's emotions in real time, and the conversational AI generates a response based on the user's emotions. Please help us develop an application that allows users to have an emotionally rich experience through natural dialogue.

[0465] In this way, a system is realized that allows users to naturally engage in emotional dialogue with the deceased in a virtual space.

[0466] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0467] Step 1:

[0468] The device collects data about the deceased.

[0469] Specifically, the user uses a device to record the deceased's facial expressions, movements, voice, speaking style, etc. for approximately seven hours using a camera and microphone. This data is time-stamped, and the audio data is saved in WAV format, and the video data in MP4 format. The input is the audio and video data of the deceased, and the output is the saved multimedia data.

[0470] Step 2:

[0471] The terminal transmits the collected data to the server.

[0472] The collected data is sent to the server using a secure protocol (e.g. HTTPS), along with a checksum to verify the integrity and completeness of the data. The input is the stored multimedia data, and the output is the data received by the server.

[0473] Step 3:

[0474] The server analyzes the data and extracts personal characteristics.

[0475] A deep learning model (for example, a model using TensorFlow) is used to analyze facial feature points, voice spectrum, speaking rhythm and tone, etc. The analysis results are saved as feature data. The input is the multimedia data received by the server, and the output is the extracted feature data.

[0476] Step 4:

[0477] The server generates a 3D virtual model based on the feature data.

[0478] Using the extracted features, 3D modeling software (e.g., Unity) is used to generate a 3D virtual model that reproduces the deceased's natural facial expressions and movements. The input is the feature data, and the output is the 3D virtual model.

[0479] Step 5:

[0480] Users access the virtual space using a dedicated client application.

[0481] Users log in to the Metaverse space using a smartphone or head-mounted display. The login information is sent to the server via authentication means. The input is the user's authentication information, and the output is the access permission granted upon successful authentication.

[0482] Step 6:

[0483] The server makes a 3D virtual model of the deceased appear in the virtual space based on the user's location information.

[0484] The 3D virtual model is placed in an appropriate position taking into account the user's line of sight and position information. The input is the user's position information, and the output is the position of the 3D virtual model in the metaverse space.

[0485] Step 7:

[0486] The terminal transmits the user's voice to the server in real time.

[0487] When a user speaks to the deceased, the audio data is captured in real time and sent to a server, where it is converted to text using the Google Cloud Speech-to-Text API. The input is the user's audio data, and the output is the converted audio data.

[0488] Step 8:

[0489] The server uses conversational AI to analyze the voice data and generate an appropriate response.

[0490] The server uses conversational AI to generate appropriate responses based on user input. It uses Azure Cognitive Services. The input is the converted voice data, and the output is the generated response text.

[0491] Step 9:

[0492] The terminal then vocalizes the generated response to the user.

[0493] The response text from the server is converted into audio format and played back to the user in real time. The input is the generated response text, and the output is the audio output.

[0494] Step 10:

[0495] The device captures the user's facial expressions and tone of voice in real time and sends them to the emotion engine.

[0496] It uses a camera and microphone to capture the user's facial expressions and tone of voice and sends them to the emotion engine, which then analyzes them using a deep learning model. The input is the user's facial and voice data, and the output is emotion data.

[0497] Step 11:

[0498] The server reflects the emotional data received from the emotion engine in the conversational AI.

[0499] The conversational AI adjusts the response it generates depending on the user's emotions. For example, if the user is recognized as sad, the response may be "Cheer up." The input is emotional data, and the output is the adjusted response text.

[0500] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0501] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (registered trademark) (Internet search engine).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0502] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0503] [Second embodiment]

[0504] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0505] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0506] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0507] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0508] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0509] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0510] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0511] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0512] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0513] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0514] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0515] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0516] This invention is a system that collects personal data from people while they are alive, analyzes the collected data to generate a 3D virtual model, and uses interactive AI in a metaverse space to allow users to experience natural conversations with the deceased.

[0517] System program generation and implementation method

[0518] Data collection

[0519] 1. A user requests data collection for a deceased person. For example, a user might say, "I want to create a virtual human of my father, so please collect his data."

[0520] 2. The device interviews and photographs the deceased. Using a camera and microphone, it records the father's facial expressions, movements, voice, and speaking style for approximately seven hours. The collected data is categorized into visual, audio, and linguistic information.

[0521] 3. The device encodes the collected data and sends it to the server using a secure protocol.

[0522] Data analysis and training of generative AI models

[0523] 1. The server analyzes the received data and uses deep learning models to extract features such as facial features, voice spectrum, speaking rhythm and tone.

[0524] 2. The server performs 3D modeling, generating a 3D virtual model that reproduces natural facial expressions and movements based on the extracted features.

[0525] 3. The server trains the conversational AI, learning from the collected voice and language data to recreate the speaking style and tone of the deceased.

[0526] Storing episode data

[0527] 1. The device collects information such as episodes, hobbies, etc. For example, it asks, "What is your favorite food?" and the father answers, "I love sushi."

[0528] 2. The device organizes this information according to a format and sends it to the server.

[0529] 3. The server stores the episode data in a database and indexes it for quick reference by the conversational AI.

[0530] Expanding into the Metaverse

[0531] 1. A user logs into the metaverse space using a dedicated client application.

[0532] 2. The server verifies the user's credentials and grants access.

[0533] 3. The server creates a 3D virtual model of the deceased person in the Metaverse space and adjusts the position of the virtual model based on the user's location information.

[0534] Specific examples

[0535] The user logs into the metaverse space and says, "Dad, it's been a while."

[0536] The device transmits this audio data to the server in real time.

[0537] The server uses conversational AI to analyze the voice data, generate an appropriate response such as, "It's been a really long time. How have you been lately?" and send it to the device.

[0538] The terminal returns the generated response to the user, allowing the user to experience a natural conversation with the deceased.

[0539] This allows users to enjoy natural conversations with their deceased loved ones within the metaverse space, providing a means to ease emotional pain.

[0540] The processing flow will be explained below.

[0541] Step 1:

[0542] The user requests the collection of data on the deceased. The request is recorded and the collection is prepared.

[0543] Step 2:

[0544] The device interviews and photographs the deceased. Using a camera and microphone, it records the deceased's facial expressions, movements, voice, and speaking style for approximately seven hours. Each piece of data is saved with a timestamp.

[0545] Step 3:

[0546] The device encodes the collected data and sends it to the server using a secure protocol, along with a checksum to verify the data's integrity and completeness.

[0547] Step 4:

[0548] The server analyzes the received data, first extracting features such as facial features, voice spectrum, and speaking rhythm and tone using deep learning algorithms.

[0549] Step 5:

[0550] The server performs 3D modeling based on the feature data, and generates a 3D virtual model using the extracted features to reproduce the natural facial expressions and movements of the deceased.

[0551] Step 6:

[0552] The server trains the conversational AI, using the collected voice and language data to train the AI ​​model to reproduce the speech style and tone of the deceased, specifically understanding the voice patterns and flow of the conversation.

[0553] Step 7:

[0554] The device separately collects information such as anecdotes and hobbies. For example, during an interview, it might ask, "What is your favorite food?" and record the answer.

[0555] Step 8:

[0556] The device sends episode data to the server, where the data is categorized into episodes, hobbies, favorite things, etc. and organized according to a format.

[0557] Step 9:

[0558] The server stores the episode data in a database, where it is indexed and made available for quick reference by the conversational AI.

[0559] Step 10:

[0560] A user logs into the metaverse space using a dedicated client application, and the login information is sent to the server via authentication means.

[0561] Step 11:

[0562] The server verifies the user's credentials and grants access. If authentication is successful, access is established within the metaverse space.

[0563] Step 12:

[0564] The server creates a 3D virtual model of the deceased person in the Metaverse space based on the user's location information, and places the model in an appropriate location, taking into account the user's line of sight and location.

[0565] Step 13:

[0566] The user talks to the deceased, for example, saying, "Dad, how are you?"

[0567] Step 14:

[0568] The device transmits the user's voice in real time to a server, where the voice data is recorded and then stored for analysis.

[0569] Step 15:

[0570] The server uses conversational AI to analyze the voice data and generate an appropriate response, such as "I've been doing well. How have you been lately?"

[0571] Step 16:

[0572] The server sends the generated response to the terminal, which then outputs the response to the user.

[0573] This allows users to experience natural interactions with the deceased within the metaverse space.

[0574] Example 1

[0575] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0576] There is a need for a system that allows users to communicate with deceased individuals and ease their emotional pain by recreating memories and conversations with them. However, existing technologies are insufficient in recreating the deceased, making it difficult to achieve natural conversations. To solve this problem, a system that accurately recreates the characteristics of the deceased and allows users to experience natural conversations is needed.

[0577] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0578] In this invention, the server includes: a means for a user to request data collection of the deceased; a means for recording the deceased's facial expressions, movements, voice, and speaking style using a camera and microphone; a means for encoding the recorded data and transmitting it to the server using a secure protocol; a means for the server to analyze the received data using a deep learning model and extract features; a means for generating a 3D virtual model based on the extracted features; a means for training an interactive AI using voice data and language data; a means for collecting information such as episodes and hobbies and storing it in a database; a means for a user to access the metaverse space using a dedicated client application; a means for the server to create a 3D virtual model of the deceased in the metaverse space and adjust it based on the user's location information; and a means for realizing natural conversation using the interactive AI, allowing users to enjoy natural conversations with the deceased in the metaverse space.

[0579] "User" refers to an individual who utilizes the system to collect data about or interact with a deceased person.

[0580] "Means for requesting data collection" refers to a function that allows a user to request data collection of a deceased person from the system.

[0581] "Camera and microphone" refers to visual and audio input devices used to record the facial expressions, movements, voice, and speaking style of the deceased.

[0582] "Means of recording" refers to the ability to use cameras and microphones to digitally record the facial expressions, movements, voice, and speaking style of the deceased.

[0583] "Encoding" refers to the process of compressing and converting recorded data into a format that can be transmitted efficiently.

[0584] "Secure protocol" refers to a communication protocol for ensuring the security of data communications, and includes SSL and TLS.

[0585] A "deep learning model" refers to an artificial intelligence technology that uses neural networks to analyze data and extract features.

[0586] "Feature extraction" refers to finding important patterns or attributes in data and using them for analysis.

[0587] A "3D virtual model" refers to a digital model that reproduces the features of the deceased and displays them in three-dimensional space.

[0588] "Conversational AI" refers to artificial intelligence technology that can mimic natural human conversation.

[0589] "Anecdotes and hobbies" refers to information that reveals the deceased's personal experiences and interests.

[0590] A "database" refers to an information management system that stores data in a structured format that makes it easy to search and reference.

[0591] "Client Application" refers to software that a user uses to access the metaverse space.

[0592] "Metaverse space" refers to a virtual reality environment constructed digitally, including a space that users can interact with.

[0593] "Location information" refers to data indicating the current coordinates or location of the user and virtual model.

[0594] This invention is a system that collects personal data from people while they are alive, analyzes the collected data to generate a 3D virtual model, and allows users to experience natural conversations with the deceased using interactive AI in a metaverse space.

[0595] Data collection

[0596] 1. A user requests data collection for a deceased person. For example, they might request, "I would like to create a virtual model of my father, so please collect his data." This request is made via a dedicated web portal or application.

[0597] 2. The device interviews and films the deceased. It uses a camera and microphone to record the father's facial expressions, movements, voice, and speaking style for approximately seven hours. The specific hardware used includes a high-resolution camera and a directional microphone. The collected data is categorized into visual, audio, and language information.

[0598] 3. The device encodes the collected data and sends it to the server using a secure protocol. A data encoder is used for this encoding, and the SSL / TLS protocol is used to ensure secure data transmission. Encoding is done using a format such as H.264 or FLAC.

[0599] Data analysis and training of generative AI models

[0600] 1. The server analyzes the received data and uses a deep learning model (e.g., TensorFlow or PyTorch) to extract features such as facial features, voice spectrum, speaking rhythm and tone.

[0601] 2. The server performs 3D modeling, using Blender, Maya, or other software to generate a 3D virtual model that reproduces natural facial expressions and movements based on the extracted features.

[0602] 3. The server trains the conversational AI, learning from the collected voice and language data and using language models such as GPT-3 and BERT to recreate the speaking style and tone of the deceased, enabling natural conversation.

[0603] Storing episode data

[0604] 1. The device collects information such as episodes, hobbies, etc. For example, it asks, "What is your favorite food?" and the father answers, "I like sushi."

[0605] 2. The terminal organizes the collected episode information according to a format and sends it to the server. This organization is done using a formatting tool.

[0606] 3. The server stores the episode data in a database and indexes it for quick reference by the conversational AI. Database systems used include MySQL and MongoDB.

[0607] Expanding into the Metaverse

[0608] 1. A user logs into the Metaverse space using a dedicated client application, including the client application for Oculus Rift.

[0609] 2. The server verifies the user's authentication information and grants access. The authentication system used is OAuth 2.0 or similar.

[0610] 3. The server creates a 3D virtual model of the deceased person in the Metaverse space, and adjusts the virtual model's position appropriately using Unity or Unreal Engine based on the user's location information.

[0611] Specific examples

[0612] 1. The user logs into the metaverse space and says, "Dad, it's been a while."

[0613] 2. The device sends this audio data to the server in real time.

[0614] 3. The server uses conversational AI to analyze the voice data, generate an appropriate response such as, "It's been a really long time. How have you been lately?" and send it to the device.

[0615] 4. The device returns the generated response to the user, allowing the user to experience a natural conversation with the deceased.

[0616] An example of a specific prompt is, "Dad, what did you do today?"

[0617] This allows users to enjoy natural conversations with their deceased loved ones within the metaverse space, providing a means to ease emotional pain.

[0618] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0619] System program processing flow

[0620] Data Collection Steps

[0621] Step 1:

[0622] A user requests data collection of a deceased person using a dedicated web portal or application. As input, the user enters details such as the name of the deceased, their relationship, and the information they wish to collect. As output, a request form is sent to the server.

[0623] Step 2:

[0624] The server receives the user's request and prepares the data collection. The user's request information is transmitted to the server as input. The collection plan is sent to the terminal as output.

[0625] Step 3:

[0626] The device uses a camera and microphone to record the facial expressions, movements, voice, and speaking style of the deceased. The camera and microphone capture the deceased's visual, audio, and linguistic information as input. The captured data is then stored on the device as output. Specifically, the camera captures high-resolution video and the microphone records audio.

[0627] Step 4:

[0628] The data recorded by the device is encoded and sent to the server using a secure protocol. Visual, audio, and language information collected as input is encoded. Encrypted data is sent to the server as output. Specifically, the data is encoded into H.264 or FLAC format and sent using the SSL / TLS protocol.

[0629] Steps for Data Analysis and Generative AI Model Training

[0630] Step 5:

[0631] The server analyzes the received data using a deep learning model. Encoded data securely transmitted to the server is provided as input. Feature-extracted data is generated as output. Specifically, TensorFlow and PyTorch are used to analyze features such as facial features, voice spectrum, speaking rhythm and tone.

[0632] Step 6:

[0633] The server generates a 3D virtual model based on the extracted features. The analyzed feature data is given as input. The 3D virtual model is generated as output. Specifically, the model is constructed using Blender or Maya, and skeletal animation is implemented.

[0634] Step 7:

[0635] The server trains the conversational AI. The collected speech and language data is used as input. The output is a trained conversational AI model that improves its ability to reproduce the speech style and tone of the deceased. Specifically, training is performed using GPT-3 and BERT to enhance the model's ability to mimic natural conversation.

[0636] Steps for storing episode data

[0637] Step 8:

[0638] The device collects information from the deceased, such as episodes and hobbies. The user or interviewer asks questions as input, and the deceased's responses are collected as voice data. The output is the collected episode information organized. Specifically, the system works by collecting a situation in which the interviewer asks, "What is your favorite food?" and the deceased answers, "I like sushi."

[0639] Step 9:

[0640] The episode information collected by the device is organized according to a format and sent to the server. Unorganized episode information is given as input, and formatted data is sent to the server as output. Specifically, the data is organized into JSON or XML format using a formatting tool and sent to the server.

[0641] Step 10:

[0642] The server stores the episode data in a database and indexes it for quick reference by the conversational AI. Formatted episode data is given as input. The output is episode information stored in a database and indexed. Specifically, data is organized and indexed using MySQL or MongoDB.

[0643] Steps for expanding into the metaverse space

[0644] Step 11:

[0645] A user logs into the metaverse space using a dedicated client application. The user's authentication information (username and password) is used as input. The output is the initiation of a login session. Specifically, the client application takes the user's credentials as input and sends them to the authentication system.

[0646] Step 12:

[0647] The server verifies the user's authentication information and grants access. User credentials are provided as input. An authentication result is generated as output. Specifically, an authentication system such as OAuth 2.0 is used, and a session begins if authentication is successful.

[0648] Step 13:

[0649] The server creates a 3D virtual model of the deceased person in the Metaverse space and adjusts it based on the user's location information. The user's current location information and 3D model information are provided as input. The output is a 3D model that has been adjusted to correspond to the user. Specifically, Unity or Unreal Engine is used to place the virtual model at a specified location in the Metaverse space.

[0650] Specific examples

[0651] 1. A user logs in to the metaverse space and says, "Dad, it's been a while." The user's voice is captured by a microphone as input. The voice data is sent to the server as output.

[0652] 2. The device sends this audio data in real time to the server. The input is the captured audio data. The output is the audio data that is transferred to the server.

[0653] 3. The server uses conversational AI to analyze the voice data, generate an appropriate response such as "It's been a really long time. How have you been lately?" and send it to the device. The user's voice data is given as input, and the response voice data is generated as output.

[0654] 4. The device returns the generated response to the user. The response audio data sent from the server is given as input. The response is played back to the user through the speaker as output.

[0655] An example of a specific prompt is, "Dad, what did you do today?"

[0656] (Application example 1)

[0657] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0658] Although existing technologies exist for systems that collect personal data from the deceased and recreate natural interactions with the deceased, they lack specific environments and memory-based interactions to improve the user's emotional satisfaction.This invention aims to provide a virtual memorial service that allows users to feel a deeper emotional connection by interacting with the deceased in a specific memorial store.

[0659] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0660] In this invention, the server includes a means for collecting personal data of the deceased while they are alive, a means for analyzing the collected data to extract personal characteristics, a means for generating a 3D virtual model using the characteristics, a means for providing a virtual space in which the user can interact with the generated 3D virtual model, a means for using interactive AI to make the interaction in the virtual space natural, a means for placing a specific memorial store in the virtual space in which the user can move, and a means for placing objects and music related to the deceased in the memorial store and realizing interaction based on memorable episodes. This allows the user to interact with the deceased in a specific environment, further strengthening the emotional connection.

[0661] "Personal data" refers to all information, including audio, video, and linguistic information, captured during a person's lifetime.

[0662] "Features" are information extracted from collected personal data, such as facial features, voice spectrum, speaking rhythm and tone.

[0663] A "3D virtual model" is a three-dimensional human model generated in a virtual space based on collected features.

[0664] A "virtual space" is a digitally recreated environment that a user can access.

[0665] "Conversational AI" is artificial intelligence that is trained to engage in natural conversations using collected voice and language data.

[0666] A "memorial store" is a specific area within the virtual space where objects, music, and stories related to the deceased are interactively arranged.

[0667] "Object" refers to any object or item placed in a virtual space.

[0668] "Music" refers to melodies and songs related to the deceased, and is sound information played within the virtual space.

[0669] "Memorable episodes" are information that provides conversations based on the deceased's hobbies, favorite things, and past events.

[0670] A "dedicated client application" is a specially designed software application used by a user to access a virtual space.

[0671] System program generation

[0672] To realize this invention, we start by collecting data on individuals during their lifetime. This process uses hardware such as a camera (e.g., OpenCV) and an audio recorder. The collected data is categorized into visual, audio, and language information.

[0673] The server receives the collected data and uses a deep learning model (e.g., TensorFlow) to extract features such as facial features, voice spectrum, speaking rhythm, and tone. Based on this, a 3D virtual model is generated. It also trains a conversational AI model based on the episode data and develops algorithms to achieve natural conversations. This model is then used during subsequent conversations.

[0674] Meanwhile, a specific memorial store will be set up in the virtual space, and will be stocked with objects and music related to the deceased. This setting is important for users to feel a deep emotional connection with the deceased. The memorial store will be rendered using Unity.

[0675] Processing procedure description

[0676] Data collection:

[0677] The server interviews and films the deceased. It uses a camera and microphone to record the deceased's facial expressions, movements, voice, and speaking style. For example, if a user requests, "I want to create a virtual human of my father, so please collect data," the device encodes the collected data and sends it to the server using a secure protocol.

[0678] Data analysis and generative AI model training:

[0679] The server analyzes the received data and uses deep learning models to extract features, which are then used to generate a 3D virtual model that reproduces natural facial expressions and movements. Furthermore, the conversational AI is trained to reproduce the speech style and tone of the deceased.

[0680] Episode data storage:

[0681] To ensure a more natural interaction when users visit the memorial store, the device collects information about the deceased's anecdotes and hobbies and stores it in a database, which the conversational AI can quickly reference.

[0682] Virtual Space Deployment:

[0683] A user logs into the virtual space using a dedicated client application, and after the server verifies the user's authentication information, a 3D virtual model of the deceased person appears in the virtual space. For example, a user can log into a "Memorial Cafe" and enjoy interacting with the 3D model of the deceased. The cafe is furnished with objects and music based on the preferences of the deceased.

[0684] Examples of concrete examples and prompts

[0685] As a concrete example, if a user logs in to Memorial Cafe and says, "Mom, you liked this cafe, didn't you?", the conversational AI will respond, "Yes, I really liked it. I love the atmosphere."

[0686] Example prompt sentence:

[0687] User: Mom, do you like this cafe?

[0688] Deceased Model: Yes, I love it. I love the atmosphere here.

[0689] This allows users to experience natural conversation with the deceased within the memorial store, further strengthening their emotional connection.

[0690] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0691] Step 1:

[0692] Collecting personal data before death.

[0693] Input: User requests data collection

[0694] How it works: The device uses a camera and microphone to record the facial expressions, movements, voice, and speaking style of the deceased. For example, based on a user's request to "create a virtual model of my father," the device collects data for approximately seven hours.

[0695] Output: A dataset containing visual, audio, and linguistic information

[0696] Step 2:

[0697] The collected data is sent to a server using a secure protocol.

[0698] Input: Dataset collected in step 1

[0699] Specific operation: The device encodes the data and sends it to the server using a secure protocol (e.g., SSL / TLS).

[0700] Output: The encoded data reaches the server

[0701] Step 3:

[0702] The server parses the data it receives.

[0703] Input: Encoded data

[0704] How it works: The server uses a deep learning model (e.g., TensorFlow) to decode the data and extract features such as facial features, voice spectrum, speaking rhythm and tone.

[0705] Output: A dataset of extracted facial feature points, voice spectrum, speaking rhythm and tone.

[0706] Step 4:

[0707] Generate a 3D virtual model.

[0708] Input: Feature data extracted in step 3

[0709] Specific operation: The server uses Unity to generate a 3D virtual model based on the extracted feature data, reproducing natural facial expressions and movements.

[0710] Output: Generated 3D virtual model

[0711] Step 5:

[0712] Train a conversational AI model.

[0713] Input: Collected speech and language data

[0714] How it works: The server uses collected voice and language data to train the conversational AI so that it can converse naturally, and uses deep learning algorithms to build a model that reproduces the speech style and tone of the deceased.

[0715] Output: A trained conversational AI model

[0716] Step 6:

[0717] Episode data is stored and indexed in a database.

[0718] Input: Information about episodes and hobbies collected from users

[0719] Specific operation: The device organizes information about episodes and hobbies, stores it in a database, and indexes it for quick reference by the conversational AI.

[0720] Output: Indexed episode data

[0721] Step 7:

[0722] A user logs into the virtual space using a dedicated client application.

[0723] Input: User credentials

[0724] How it works: A user logs into a virtual space using a dedicated client application. The server verifies the authentication information and grants the user access.

[0725] Output: The user can access the virtual space.

[0726] Step 8:

[0727] A memorial store will be placed in the virtual space, and a 3D virtual model of the deceased person will appear.

[0728] Input: User's location, 3D virtual model generated in step 4

[0729] Specific operation: The server places a memorial store in the virtual space and makes a 3D virtual model of the deceased person appear based on the user's location information.

[0730] Output: 3D virtual model of the deceased placed in a memorial store

[0731] Step 9:

[0732] The user engages in natural dialogue with conversational AI.

[0733] Input: User voice input, the conversational AI model trained in step 5, and episode data from step 6

[0734] How it works: When a user speaks to a model of the deceased person, the device sends the voice data in real time to the server. The server uses a conversational AI model to generate a response and sends it to the device. The device then returns the response to the user, realizing a natural dialogue.

[0735] Output: A natural dialogue experience between the user and conversational AI

[0736] Specific examples

[0737] For example, if a user logs in to Memorial Cafe and says, "Mom, do you like this cafe?" the conversational AI will respond, "Yes, I love it. I love the atmosphere here."

[0738] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0739] This invention combines a system that collects personal data from a person's lifetime, analyzes the collected data to generate a 3D virtual model, and enables users to experience natural conversations with the deceased using interactive AI within the metaverse space, with an emotion engine that recognizes the user's emotions.

[0740] System program generation and implementation method

[0741] Generate 3D virtual models from data collection

[0742] 1. A user requests the collection of data on a deceased person. The request is recorded and preparations for collection are made.

[0743] 2. The device interviews and photographs the deceased. Using a camera and microphone, it records the deceased's facial expressions, movements, voice, and speaking style for approximately seven hours. Each piece of data is saved with a timestamp.

[0744] 3. The device encodes the collected data and sends it to the server using a secure protocol, along with a checksum to verify the data's integrity and completeness.

[0745] 4. The server analyzes the received data, using deep learning models to extract features such as facial features, voice spectrum, and speaking rhythm and tone.

[0746] 5. The server performs 3D modeling based on the feature data. The extracted features are used to generate a 3D virtual model that reproduces the natural facial expressions and movements of the deceased.

[0747] 6. The server trains the conversational AI. Using the collected voice and language data, the AI ​​model is trained to reproduce the speaking style and tone of the deceased, specifically understanding the deceased's voice patterns and conversational flow.

[0748] 7. The device collects information about episodes, hobbies, etc. For example, during an interview, it might ask, "What is your favorite food?" and record the answer.

[0749] 8. The device sends the episode data to the server. The data is categorized into episodes, hobbies, favorite things, etc. and organized according to a format.

[0750] 9. The server stores the episode data in a database, where it is indexed and made available for quick reference by the conversational AI.

[0751] Expanding into the Metaverse

[0752] 1. A user logs into the metaverse space using a dedicated client application. The login information is sent to the server via authentication methods.

[0753] 2. The server verifies the user's credentials and grants access. If authentication is successful, access is established within the metaverse space.

[0754] 3. The server creates a 3D virtual model of the deceased person in the Metaverse space based on the user's location information, placing it in an appropriate position taking into account the user's line of sight and location.

[0755] 4. The user talks to the deceased, for example, "Dad, how are you?"

[0756] 5. The device transmits the user's voice in real time to the server, where it is recorded and stored for analysis.

[0757] 6. The server uses conversational AI to analyze the voice data and generate an appropriate response, such as "I've been doing well. How have you been lately?"

[0758] 7. The server sends the generated response to the terminal, which then outputs the response to the user.

[0759] Incorporating an emotion engine

[0760] 1. The device captures the user's facial expressions and tone of voice in real time using a camera and microphone.

[0761] 2. The device sends the captured data to the emotion engine, which then recognizes the user's emotion based on this data. For example, if a smile is detected, it will recognize that the user is feeling "joy."

[0762] 3. The server receives the emotion data sent from the emotion engine and reflects it in the conversational AI. For example, if the user is recognized as sad, the conversational AI generates a response such as "Cheer up."

[0763] Specific examples

[0764] For example, a user logs into the metaverse space and says, "Dad, it's been a while."

[0765] The device transmits this audio data to the server in real time.

[0766] The emotion engine analyzes and recognizes that the user's tone of voice is sad.

[0767] The server uses conversational AI to analyze the voice data and, referring to the analysis results of the emotion engine, generates an appropriate response such as, "It's been a while. What's wrong? You seem down."

[0768] The server sends the generated response to the terminal, which then speaks the response to the user.

[0769] This allows users to have more emotionally enriching interactions with deceased loved ones within the metaverse space.

[0770] The processing flow will be explained below.

[0771] Step 1:

[0772] The user requests the collection of data on the deceased. The request is confirmed and preparations for collection are made.

[0773] Step 2:

[0774] The device interviews and photographs the deceased. Using a camera and microphone, it records the deceased's facial expressions, movements, voice, and speaking style for approximately seven hours. The collected data is then saved with a timestamp.

[0775] Step 3:

[0776] The device encodes the collected data and transmits it to the server using a secure protocol, along with a checksum to ensure data integrity and completeness.

[0777] Step 4:

[0778] The server analyzes the received data and uses deep learning models to extract features such as facial features, voice spectrum, speaking rhythm and tone.

[0779] Step 5:

[0780] The server performs 3D modeling based on the feature data, and generates a 3D virtual model using the extracted features to reproduce the natural facial expressions and movements of the deceased.

[0781] Step 6:

[0782] The server trains the conversational AI. Using the collected voice and language data, the AI ​​model is trained to reproduce the speaking style and tone of the deceased. Specifically, it analyzes and learns from the deceased's voice patterns and conversational flow.

[0783] Step 7:

[0784] The device collects information such as episodes and hobbies. For example, it asks, "What is your favorite food?" and records the answer.

[0785] Step 8:

[0786] The device sends episode data to the server, which organizes the data into categories such as episodes, hobbies, and favorite things.

[0787] Step 9:

[0788] The server stores the episode data in a database, where it is indexed and prepared for fast and efficient reference by the conversational AI.

[0789] Step 10:

[0790] A user logs into the metaverse space using a dedicated client application, and the login information is sent to the server via authentication means.

[0791] Step 11:

[0792] The server verifies the user's credentials and grants access. If authentication is successful, access is established within the metaverse space.

[0793] Step 12:

[0794] The server creates a 3D virtual model of the deceased person in the Metaverse space based on the user's location information, and adjusts the position of the virtual model based on the user's line of sight and location.

[0795] Step 13:

[0796] The user talks to the deceased, for example, saying, "Dad, it's been a long time."

[0797] Step 14:

[0798] The device transmits the user's voice in real time to a server, where the voice data is recorded and stored for analysis.

[0799] Step 15:

[0800] The server uses conversational AI to analyze the voice data and generate an appropriate response, such as "It's been a while, how have you been?"

[0801] Step 16:

[0802] The device captures the user's facial expressions and tone of voice in real time using a camera and microphone.

[0803] Step 17:

[0804] The device captures facial and voice data and sends it to the emotion engine, which analyzes this data and recognizes the user's emotions.

[0805] Step 18:

[0806] The server receives the emotion data sent from the emotion engine and reflects it in the conversational AI. For example, if the user looks sad, the conversational AI will generate a response such as "Cheer up, what happened recently?"

[0807] Step 19:

[0808] The server sends the generated response to the terminal, which then provides an audio output to the user.

[0809] This allows users to interact with deceased loved ones in a more emotionally rich and natural way within the metaverse space.

[0810] Example 2

[0811] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0812] Conventional systems have had difficulty using data from deceased people to recreate conversations with them in a virtual space. Furthermore, they lacked the ability to generate responses based on the user's emotions, making the conversations less emotionally rich and natural.

[0813] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for collecting personal data before death, means for analyzing the collected data to extract personal characteristics, means for generating a 3D virtual model using the characteristics, means for providing a virtual space in which the user can interact with the generated 3D virtual model, means for using interactive artificial intelligence to make the interaction in the virtual space natural, and means for recognizing the user's emotions and flexibly adjusting the response content. This allows the user to experience an emotionally rich and natural interaction with the deceased person in the virtual space.

[0814] "Personal data" is information about a person that is collected during their lifetime, including photographs, videos, audio files, and written material.

[0815] "Feature extraction means" refers to a process or technique for identifying and analyzing individual characteristics or features from collected data.

[0816] A "3D virtual model" is a three-dimensional digital representation of an individual's appearance, gestures, movements, etc., based on collected data and extracted features.

[0817] "Virtual space" refers to a virtual environment that users can access through a computer or dedicated client application, including the metaverse and VR space.

[0818] "Conversational artificial intelligence" refers to a computer program or system that can use natural language analysis and generation techniques to engage in natural conversations with humans.

[0819] "Means for recognizing emotions and flexibly adjusting response content" refers to a function that analyzes the user's facial expressions, tone of voice, and other emotional indicators in real time and dynamically changes the response content based on the results.

[0820] "Episode data" refers to information based on specific events, hobbies, or specific memories related to the deceased.

[0821] "Dedicated client application" refers to software specifically designed to access virtual spaces and interface with interactive AI and 3D virtual models.

[0822] MODE FOR CARRYING OUT THE INVENTION

[0823] This invention is a system that collects personal data from a person's life and analyzes it to generate a 3D virtual model. This system uses conversational artificial intelligence to allow users to experience natural interactions with the deceased in the metaverse space. It also combines an emotion engine that recognizes the user's emotions to realize more emotionally rich interactions.

[0824] Hardware and software used

[0825] Hardware:

[0826] Camera: Used to record the facial expressions and movements of the deceased.

[0827] Microphone: Used to record the voice and speaking style of the deceased.

[0828] Devices (e.g., computers, smartphones): used for data collection and transmission.

[0829] Server: Used for data processing, analysis, storage, AI training, and virtual space management.

[0830] software:

[0831] Dedicated client application: Software that allows users to access the virtual space.

[0832] Data collection software: collects data from the camera and microphone, encodes it, and sends it to a server.

[0833] Deep learning models: Algorithms for data analysis and feature extraction.

[0834] 3D modeling tools (e.g. Blender): Generate 3D virtual models based on collected feature data.

[0835] Conversational Artificial Intelligence (AI): Enables natural dialogue with users.

[0836] Emotion engine: Analyzes user emotions and reflects them in the dialogue AI.

[0837] Examples of concrete examples and prompts

[0838] For example, suppose a user logs into the metaverse space using a dedicated client application and says, "Dad, it's been a while." The specific processing at this time is as follows.

[0839] 1. The device sends this voice data to the server in real time.

[0840] 2. The server uses conversational artificial intelligence to analyze the voice data and generates an appropriate response by referring to the analysis results of the emotion engine.

[0841] 3. For example, the emotion engine analyzes and recognizes that the user's tone of voice is sad, and the server generates a response such as, "It's been a while, what's wrong? You seem depressed."

[0842] 4. The server sends the generated response to the terminal, which then speaks the response to the user.

[0843] In this way, users can have emotionally enriching interactions with deceased loved ones within the metaverse space.

[0844] Example prompt sentence:

[0845] "Dad, how have you been?"

[0846] "How are you feeling today?"

[0847] "I want to talk about recent events."

[0848] By inputting these prompt sentences, the conversational AI can understand the user's intentions and emotions and generate the most appropriate response.

[0849] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0850] Step 1: (Data collection request)

[0851] The user requests the collection of data on the deceased, specifying specific items to be collected, such as photos, videos, and audio files of the deceased.

[0852] The device records the request and begins preparations for collection, preparing the camera and microphone, and confirming the data format to be collected.

[0853] Input: User request.

[0854] Output: Collection readiness status.

[0855] Step 2: (Interview and filming)

[0856] The device uses a camera and microphone to interview and film the deceased, and the data collected includes facial expressions, movements, voice, and speaking style.

[0857] The device stores all data with a timestamp, splitting it into video clips every 10 minutes, for example.

[0858] Input: Real-time data from camera and microphone.

[0859] Output: Collected data with timestamp.

[0860] Step 3: (Data encoding and transmission)

[0861] The device encodes the collected data and transmits it to a server using a secure protocol (e.g., HTTPS).

[0862] The terminal generates a checksum to verify the integrity and completeness of the data and also sends this to the server.

[0863] Input: Collected data before encoding.

[0864] Output: The encoded data and a checksum.

[0865] Step 4: (Data Analysis)

[0866] The server then analyzes the received data using a deep learning model that extracts features such as facial features, voice spectrum, and speaking rhythm and tone.

[0867] The server specifically analyzes the deceased person's unique movements and expressions and stores them in a database.

[0868] Input: The encoded data.

[0869] Output: Extracted feature data.

[0870] Step 5: (3D modeling)

[0871] The server generates a 3D virtual model based on the extracted feature data, carefully adjusting it to reproduce the natural facial expressions and movements of the deceased.

[0872] The server uses a 3D modeling tool (e.g. Blender) to see in real time what needs to be fixed.

[0873] Input: Extracted feature data.

[0874] Output: 3D virtual model.

[0875] Step 6: (Training the conversational AI)

[0876] The server uses the collected speech and language data to train the conversational AI, specifically learning patterns to replicate the speech and narration of the deceased.

[0877] The server mimics voice patterns and adjusts them to create a natural conversational flow.

[0878] Input: Audio and language data.

[0879] Output: A trained conversational AI model.

[0880] Step 7: (Collect episode information)

[0881] Users provide information about the deceased, such as anecdotes, hobbies, and interests, in the form of an interview.

[0882] The device records this information and also saves it as text data.

[0883] Input: The answer given by the user.

[0884] Output: Episode data.

[0885] Step 8: (Submit episode data)

[0886] The episode data collected by the device is organized according to a format and sent to the server.

[0887] The device classifies the data (episodes, hobbies, interests) before sending it.

[0888] Input: Episode data.

[0889] Output: Organized episode data.

[0890] Step 9: (Storing episode data)

[0891] The server indexes and stores the received episode data in a database.

[0892] The server applies a search algorithm to allow quick access to the data.

[0893] Input: Organized episode data.

[0894] Output: Indexed episode data in a database.

[0895] Step 10: (Login and Authentication)

[0896] A user logs into the virtual space using a dedicated client application, and the login information is sent from the client to the server.

[0897] The server verifies the user's credentials and grants access.

[0898] Input: User login information.

[0899] Output: Authentication check and access granted.

[0900] Step 11: (Placing the virtual model)

[0901] The server creates a 3D virtual model of the deceased person in a virtual space based on the user's location information.

[0902] The server places the virtual model in an appropriate position, taking into account the user's line of sight and position.

[0903] Input: User's location.

[0904] Output: Positioning information of the virtual model in the virtual space.

[0905] Step 12: (User interaction)

[0906] The user talks to the deceased, for example, asking, "Dad, how are you?"

[0907] The device transmits this audio data to the server in real time.

[0908] The server uses conversational artificial intelligence to analyze the voice data and generates an appropriate response by referring to the analysis results of the emotion engine.

[0909] Input: User's question, voice data.

[0910] Output: A response from the conversational AI.

[0911] Step 13: (Capturing the user's facial expressions and voice)

[0912] The device captures the user's facial expressions and tone of voice in real time using a camera and microphone.

[0913] The device transmits the captured data to the emotion engine in real time.

[0914] Input: User's facial expression data and tone of voice.

[0915] Output: Emotion data sent to the emotion engine.

[0916] Step 14: (Analyze emotion data)

[0917] The server analyzes the user's emotions based on the data sent from the emotion engine. For example, it recognizes "joy" by detecting a smile.

[0918] Input: User emotion data from the emotion engine.

[0919] Output: Analyzed user emotion information.

[0920] Step 15: (Generating and reflecting dialogue responses)

[0921] The server generates a response from the conversational AI based on the emotional data. For example, if the server recognizes that the user is sad, it generates a response such as "Cheer up."

[0922] The server sends the generated response to the terminal, which then outputs it to the user.

[0923] Input: Parsed user emotion information.

[0924] Output: AI response based on emotion.

[0925] (Application example 2)

[0926] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0927] In modern society, people often seek emotional connections through conversations with deceased loved ones, but systems that can provide such experiences are not yet widespread. Furthermore, there are no systems incorporating conversational AI that can recognize and reflect emotions in real time, making it difficult to understand the user's emotions and provide natural conversations accordingly. Furthermore, for users to naturally converse with deceased loved ones in virtual spaces, it is necessary to manage and analyze episode data and emotional data, and an efficient system for this purpose is needed.

[0928] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0929] In this invention, the server includes means for collecting personal data before death, means for analyzing the collected data to extract personal characteristics, means for generating a 3D virtual model using the characteristics, means for providing a virtual space in which a user can interact with the generated 3D virtual model, means for using interactive AI to make the interaction in the virtual space natural, means for capturing and analyzing the user's voice and emotions in real time, and means for generating a response that reflects the analyzed emotional data, thereby enabling a user to have an emotional experience through interaction with the deceased.

[0930] "Personal data" refers to information provided by an individual during their lifetime, such as their voice, facial expressions, movements, and speaking style.

[0931] "Features" refers to identifying information such as facial features, voice spectrum, speaking rhythm and tone extracted from collected personal data.

[0932] "3D Virtual Model" means a three-dimensional model of an individual displayed in a virtual space, generated based on collected characteristics.

[0933] "Virtual space" refers to a virtual reality environment accessible to a user.

[0934] "Conversational AI" refers to an artificial intelligence system that uses natural language processing to engage in dialogue with users.

[0935] "Means for capturing and analyzing voice and emotions in real time" refers to the function of collecting the user's voice and facial expressions in real time using a voice input device or camera, and analyzing them using an emotion analysis engine.

[0936] "Emotion data" refers to information about emotions analyzed from the user's tone of voice, facial expressions, etc.

[0937] "Means for generating a response" refers to an appropriate response that corresponds to the user's emotions, which is generated by the conversational AI based on the analyzed emotional data.

[0938] This invention is a system that allows a user to have natural conversations with a deceased person in a virtual space based on personal data collected by the user while he or she was alive. An embodiment of this system will be described in detail below.

[0939] First, the user uses a device to collect data on the deceased. Data such as the deceased's facial expressions, movements, voice, and speaking style is recorded using a camera and microphone for approximately seven hours. This collected data is then sent to a server using a secure protocol. The server then uses a deep learning model to analyze the data and extract individual features. Based on the extracted features, a realistic and natural-looking 3D virtual model is then generated.

[0940] The generated 3D virtual model is displayed in a virtual space accessible to users. Users access this virtual space using a dedicated client application on their smartphone or head-mounted display. When the user speaks to the deceased, their voice is transmitted in real time to a server, where conversational AI analyzes the voice and generates an appropriate response.

[0941] The system also incorporates an emotion engine that captures and analyzes the user's voice and facial expressions in real time. The emotion engine analyzes emotions from the user's tone of voice and facial expressions and reflects them in the responses generated by the conversational AI. This allows the user to experience a more natural and emotional dialogue. For example, if a user says, "Dad, it's been a while," the emotion engine will analyze that the user's tone of voice sounds sad, and the conversational AI will generate a response such as, "It's been a while. What's wrong? You seem down."

[0942] The hardware used includes a smartphone, a head-mounted display, a camera, and a microphone, while the software used includes Unity (a 3D game engine), TensorFlow (a deep learning framework for sentiment analysis), Google Cloud Speech-to-Text API, and Azure Cognitive Services (conversational AI).

[0943] An example prompt is:

[0944] "User prompt"

[0945] We would like to request the development of an application that provides an emotional experience through interaction with the deceased, such as:

[0946] The user collects data about the deceased and conducts an interview using a camera and microphone. The collected data is sent to a server, and a 3D virtual model is generated using deep learning. The user logs in to the virtual space and begins a dialogue. The emotion engine analyzes the user's emotions in real time, and the conversational AI generates a response based on the user's emotions. Please help us develop an application that allows users to have an emotionally rich experience through natural dialogue.

[0947] In this way, a system is realized that allows users to naturally engage in emotional dialogue with the deceased in a virtual space.

[0948] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0949] Step 1:

[0950] The device collects data about the deceased.

[0951] Specifically, the user uses a device to record the deceased's facial expressions, movements, voice, speaking style, etc. for approximately seven hours using a camera and microphone. This data is time-stamped, and the audio data is saved in WAV format, and the video data in MP4 format. The input is the audio and video data of the deceased, and the output is the saved multimedia data.

[0952] Step 2:

[0953] The terminal transmits the collected data to the server.

[0954] The collected data is sent to the server using a secure protocol (e.g. HTTPS), along with a checksum to verify the integrity and completeness of the data. The input is the stored multimedia data, and the output is the data received by the server.

[0955] Step 3:

[0956] The server analyzes the data and extracts personal characteristics.

[0957] A deep learning model (for example, a model using TensorFlow) is used to analyze facial feature points, voice spectrum, speaking rhythm and tone, etc. The analysis results are saved as feature data. The input is the multimedia data received by the server, and the output is the extracted feature data.

[0958] Step 4:

[0959] The server generates a 3D virtual model based on the feature data.

[0960] Using the extracted features, 3D modeling software (e.g., Unity) is used to generate a 3D virtual model that reproduces the deceased's natural facial expressions and movements. The input is the feature data, and the output is the 3D virtual model.

[0961] Step 5:

[0962] Users access the virtual space using a dedicated client application.

[0963] Users log in to the Metaverse space using a smartphone or head-mounted display. The login information is sent to the server via authentication means. The input is the user's authentication information, and the output is the access permission granted upon successful authentication.

[0964] Step 6:

[0965] The server makes a 3D virtual model of the deceased appear in the virtual space based on the user's location information.

[0966] The 3D virtual model is placed in an appropriate position taking into account the user's line of sight and position information. The input is the user's position information, and the output is the position of the 3D virtual model in the metaverse space.

[0967] Step 7:

[0968] The terminal transmits the user's voice to the server in real time.

[0969] When a user speaks to the deceased, the audio data is captured in real time and sent to a server, where it is converted to text using the Google Cloud Speech-to-Text API. The input is the user's audio data, and the output is the converted audio data.

[0970] Step 8:

[0971] The server uses conversational AI to analyze the voice data and generate an appropriate response.

[0972] The server uses conversational AI to generate appropriate responses based on user input. It uses Azure Cognitive Services. The input is the converted voice data, and the output is the generated response text.

[0973] Step 9:

[0974] The terminal then vocalizes the generated response to the user.

[0975] The response text from the server is converted into audio format and played back to the user in real time. The input is the generated response text, and the output is the audio output.

[0976] Step 10:

[0977] The device captures the user's facial expressions and tone of voice in real time and sends them to the emotion engine.

[0978] It uses a camera and microphone to capture the user's facial expressions and tone of voice and sends them to the emotion engine, which then analyzes them using a deep learning model. The input is the user's facial and voice data, and the output is emotion data.

[0979] Step 11:

[0980] The server reflects the emotional data received from the emotion engine in the conversational AI.

[0981] The conversational AI adjusts the response it generates depending on the user's emotions. For example, if the user is recognized as sad, the response may be "Cheer up." The input is emotional data, and the output is the adjusted response text.

[0982] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0983] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0984] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0985] [Third embodiment]

[0986] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0987] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.

[0988] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0989] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0990] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0991] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0992] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0993] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0994] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0995] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0996] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0997] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0998] This invention is a system that collects personal data from people while they are alive, analyzes the collected data to generate a 3D virtual model, and uses interactive AI in a metaverse space to allow users to experience natural conversations with the deceased.

[0999] System program generation and implementation method

[1000] Data collection

[1001] 1. A user requests data collection for a deceased person. For example, a user might say, "I want to create a virtual human of my father, so please collect his data."

[1002] 2. The device interviews and photographs the deceased. Using a camera and microphone, it records the father's facial expressions, movements, voice, and speaking style for approximately seven hours. The collected data is categorized into visual, audio, and linguistic information.

[1003] 3. The device encodes the collected data and sends it to the server using a secure protocol.

[1004] Data analysis and training of generative AI models

[1005] 1. The server analyzes the received data and uses deep learning models to extract features such as facial features, voice spectrum, speaking rhythm and tone.

[1006] 2. The server performs 3D modeling, generating a 3D virtual model that reproduces natural facial expressions and movements based on the extracted features.

[1007] 3. The server trains the conversational AI, learning from the collected voice and language data to recreate the speaking style and tone of the deceased.

[1008] Storing episode data

[1009] 1. The device collects information such as episodes, hobbies, etc. For example, it asks, "What is your favorite food?" and the father answers, "I love sushi."

[1010] 2. The device organizes this information according to a format and sends it to the server.

[1011] 3. The server stores the episode data in a database and indexes it for quick reference by the conversational AI.

[1012] Expanding into the Metaverse

[1013] 1. A user logs into the metaverse space using a dedicated client application.

[1014] 2. The server verifies the user's credentials and grants access.

[1015] 3. The server creates a 3D virtual model of the deceased person in the Metaverse space and adjusts the position of the virtual model based on the user's location information.

[1016] Specific examples

[1017] The user logs into the metaverse space and says, "Dad, it's been a while."

[1018] The device transmits this audio data to the server in real time.

[1019] The server uses conversational AI to analyze the voice data, generate an appropriate response such as, "It's been a really long time. How have you been lately?" and send it to the device.

[1020] The terminal returns the generated response to the user, allowing the user to experience a natural conversation with the deceased.

[1021] This allows users to enjoy natural conversations with their deceased loved ones within the metaverse space, providing a means to ease emotional pain.

[1022] The processing flow will be explained below.

[1023] Step 1:

[1024] The user requests the collection of data on the deceased. The request is recorded and the collection is prepared.

[1025] Step 2:

[1026] The device interviews and photographs the deceased. Using a camera and microphone, it records the deceased's facial expressions, movements, voice, and speaking style for approximately seven hours. Each piece of data is saved with a timestamp.

[1027] Step 3:

[1028] The device encodes the collected data and sends it to the server using a secure protocol, along with a checksum to verify the data's integrity and completeness.

[1029] Step 4:

[1030] The server analyzes the received data, first extracting features such as facial features, voice spectrum, and speaking rhythm and tone using deep learning algorithms.

[1031] Step 5:

[1032] The server performs 3D modeling based on the feature data, and generates a 3D virtual model using the extracted features to reproduce the natural facial expressions and movements of the deceased.

[1033] Step 6:

[1034] The server trains the conversational AI, using the collected voice and language data to train the AI ​​model to reproduce the speech style and tone of the deceased, specifically understanding the voice patterns and flow of the conversation.

[1035] Step 7:

[1036] The device separately collects information such as anecdotes and hobbies. For example, during an interview, it might ask, "What is your favorite food?" and record the answer.

[1037] Step 8:

[1038] The device sends episode data to the server, where the data is categorized into episodes, hobbies, favorite things, etc. and organized according to a format.

[1039] Step 9:

[1040] The server stores the episode data in a database, where it is indexed and made available for quick reference by the conversational AI.

[1041] Step 10:

[1042] A user logs into the metaverse space using a dedicated client application, and the login information is sent to the server via authentication means.

[1043] Step 11:

[1044] The server verifies the user's credentials and grants access. If authentication is successful, access is established within the metaverse space.

[1045] Step 12:

[1046] The server creates a 3D virtual model of the deceased person in the Metaverse space based on the user's location information, and places the model in an appropriate location, taking into account the user's line of sight and location.

[1047] Step 13:

[1048] The user talks to the deceased, for example, saying, "Dad, how are you?"

[1049] Step 14:

[1050] The device transmits the user's voice in real time to a server, where the voice data is recorded and then stored for analysis.

[1051] Step 15:

[1052] The server uses conversational AI to analyze the voice data and generate an appropriate response, such as "I've been doing well. How have you been lately?"

[1053] Step 16:

[1054] The server sends the generated response to the terminal, which then outputs the response to the user.

[1055] This allows users to experience natural interactions with the deceased within the metaverse space.

[1056] Example 1

[1057] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1058] There is a need for a system that allows users to communicate with deceased individuals and ease their emotional pain by recreating memories and conversations with them. However, existing technologies are insufficient in recreating the deceased, making it difficult to achieve natural conversations. To solve this problem, a system that accurately recreates the characteristics of the deceased and allows users to experience natural conversations is needed.

[1059] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1060] In this invention, the server includes: a means for a user to request data collection of the deceased; a means for recording the deceased's facial expressions, movements, voice, and speaking style using a camera and microphone; a means for encoding the recorded data and transmitting it to the server using a secure protocol; a means for the server to analyze the received data using a deep learning model and extract features; a means for generating a 3D virtual model based on the extracted features; a means for training an interactive AI using voice data and language data; a means for collecting information such as episodes and hobbies and storing it in a database; a means for a user to access the metaverse space using a dedicated client application; a means for the server to create a 3D virtual model of the deceased in the metaverse space and adjust it based on the user's location information; and a means for realizing natural conversation using the interactive AI, allowing users to enjoy natural conversations with the deceased in the metaverse space.

[1061] "User" refers to an individual who utilizes the system to collect data about or interact with a deceased person.

[1062] "Means for requesting data collection" refers to a function that allows a user to request data collection of a deceased person from the system.

[1063] "Camera and microphone" refers to visual and audio input devices used to record the facial expressions, movements, voice, and speaking style of the deceased.

[1064] "Means of recording" refers to the ability to use cameras and microphones to digitally record the facial expressions, movements, voice, and speaking style of the deceased.

[1065] "Encoding" refers to the process of compressing and converting recorded data into a format that can be transmitted efficiently.

[1066] "Secure protocol" refers to a communication protocol for ensuring the security of data communications, and includes SSL and TLS.

[1067] A "deep learning model" refers to an artificial intelligence technology that uses neural networks to analyze data and extract features.

[1068] "Feature extraction" refers to finding important patterns or attributes in data and using them for analysis.

[1069] A "3D virtual model" refers to a digital model that reproduces the features of the deceased and displays them in three-dimensional space.

[1070] "Conversational AI" refers to artificial intelligence technology that can mimic natural human conversation.

[1071] "Anecdotes and hobbies" refers to information that reveals the deceased's personal experiences and interests.

[1072] A "database" refers to an information management system that stores data in a structured format that makes it easy to search and reference.

[1073] "Client Application" refers to software that a user uses to access the metaverse space.

[1074] "Metaverse space" refers to a virtual reality environment constructed digitally, including a space that users can interact with.

[1075] "Location information" refers to data indicating the current coordinates or location of the user and virtual model.

[1076] This invention is a system that collects personal data from people while they are alive, analyzes the collected data to generate a 3D virtual model, and allows users to experience natural conversations with the deceased using interactive AI in a metaverse space.

[1077] Data collection

[1078] 1. A user requests data collection for a deceased person. For example, they might request, "I would like to create a virtual model of my father, so please collect his data." This request is made via a dedicated web portal or application.

[1079] 2. The device interviews and films the deceased. It uses a camera and microphone to record the father's facial expressions, movements, voice, and speaking style for approximately seven hours. The specific hardware used includes a high-resolution camera and a directional microphone. The collected data is categorized into visual, audio, and language information.

[1080] 3. The device encodes the collected data and sends it to the server using a secure protocol. A data encoder is used for this encoding, and the SSL / TLS protocol is used to ensure secure data transmission. Encoding is done using a format such as H.264 or FLAC.

[1081] Data analysis and training of generative AI models

[1082] 1. The server analyzes the received data and uses a deep learning model (e.g., TensorFlow or PyTorch) to extract features such as facial features, voice spectrum, speaking rhythm and tone.

[1083] 2. The server performs 3D modeling, using Blender, Maya, or other software to generate a 3D virtual model that reproduces natural facial expressions and movements based on the extracted features.

[1084] 3. The server trains the conversational AI, learning from the collected voice and language data and using language models such as GPT-3 and BERT to recreate the speaking style and tone of the deceased, enabling natural conversation.

[1085] Storing episode data

[1086] 1. The device collects information such as episodes, hobbies, etc. For example, it asks, "What is your favorite food?" and the father answers, "I like sushi."

[1087] 2. The terminal organizes the collected episode information according to a format and sends it to the server. This organization is done using a formatting tool.

[1088] 3. The server stores the episode data in a database and indexes it for quick reference by the conversational AI. Database systems used include MySQL and MongoDB.

[1089] Expanding into the Metaverse

[1090] 1. A user logs into the Metaverse space using a dedicated client application, including the client application for Oculus Rift.

[1091] 2. The server verifies the user's authentication information and grants access. The authentication system used is OAuth 2.0 or similar.

[1092] 3. The server creates a 3D virtual model of the deceased person in the Metaverse space, and adjusts the virtual model's position appropriately using Unity or Unreal Engine based on the user's location information.

[1093] Specific examples

[1094] 1. The user logs into the metaverse space and says, "Dad, it's been a while."

[1095] 2. The device sends this audio data to the server in real time.

[1096] 3. The server uses conversational AI to analyze the voice data, generate an appropriate response such as, "It's been a really long time. How have you been lately?" and send it to the device.

[1097] 4. The device returns the generated response to the user, allowing the user to experience a natural conversation with the deceased.

[1098] An example of a specific prompt is, "Dad, what did you do today?"

[1099] This allows users to enjoy natural conversations with their deceased loved ones within the metaverse space, providing a means to ease emotional pain.

[1100] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1101] System program processing flow

[1102] Data Collection Steps

[1103] Step 1:

[1104] A user requests data collection of a deceased person using a dedicated web portal or application. As input, the user enters details such as the name of the deceased, their relationship, and the information they wish to collect. As output, a request form is sent to the server.

[1105] Step 2:

[1106] The server receives the user's request and prepares the data collection. The user's request information is transmitted to the server as input. The collection plan is sent to the terminal as output.

[1107] Step 3:

[1108] The device uses a camera and microphone to record the facial expressions, movements, voice, and speaking style of the deceased. The camera and microphone capture the deceased's visual, audio, and linguistic information as input. The captured data is then stored on the device as output. Specifically, the camera captures high-resolution video and the microphone records audio.

[1109] Step 4:

[1110] The data recorded by the device is encoded and sent to the server using a secure protocol. Visual, audio, and language information collected as input is encoded. Encrypted data is sent to the server as output. Specifically, the data is encoded into H.264 or FLAC format and sent using the SSL / TLS protocol.

[1111] Steps for Data Analysis and Generative AI Model Training

[1112] Step 5:

[1113] The server analyzes the received data using a deep learning model. Encoded data securely transmitted to the server is provided as input. Feature-extracted data is generated as output. Specifically, TensorFlow and PyTorch are used to analyze features such as facial features, voice spectrum, speaking rhythm and tone.

[1114] Step 6:

[1115] The server generates a 3D virtual model based on the extracted features. The analyzed feature data is given as input. The 3D virtual model is generated as output. Specifically, the model is constructed using Blender or Maya, and skeletal animation is implemented.

[1116] Step 7:

[1117] The server trains the conversational AI. The collected speech and language data is used as input. The output is a trained conversational AI model that improves its ability to reproduce the speech style and tone of the deceased. Specifically, training is performed using GPT-3 and BERT to enhance the model's ability to mimic natural conversation.

[1118] Steps for storing episode data

[1119] Step 8:

[1120] The device collects information from the deceased, such as episodes and hobbies. The user or interviewer asks questions as input, and the deceased's responses are collected as voice data. The output is the collected episode information organized. Specifically, the system works by collecting a situation in which the interviewer asks, "What is your favorite food?" and the deceased answers, "I like sushi."

[1121] Step 9:

[1122] The episode information collected by the device is organized according to a format and sent to the server. Unorganized episode information is given as input, and formatted data is sent to the server as output. Specifically, the data is organized into JSON or XML format using a formatting tool and sent to the server.

[1123] Step 10:

[1124] The server stores the episode data in a database and indexes it for quick reference by the conversational AI. Formatted episode data is given as input. The output is episode information stored in a database and indexed. Specifically, data is organized and indexed using MySQL or MongoDB.

[1125] Steps for expanding into the metaverse space

[1126] Step 11:

[1127] A user logs into the metaverse space using a dedicated client application. The user's authentication information (username and password) is used as input. The output is the initiation of a login session. Specifically, the client application takes the user's credentials as input and sends them to the authentication system.

[1128] Step 12:

[1129] The server verifies the user's authentication information and grants access. User credentials are provided as input. An authentication result is generated as output. Specifically, an authentication system such as OAuth 2.0 is used, and a session begins if authentication is successful.

[1130] Step 13:

[1131] The server creates a 3D virtual model of the deceased person in the Metaverse space and adjusts it based on the user's location information. The user's current location information and 3D model information are provided as input. The output is a 3D model that has been adjusted to correspond to the user. Specifically, Unity or Unreal Engine is used to place the virtual model at a specified location in the Metaverse space.

[1132] Specific examples

[1133] 1. A user logs in to the metaverse space and says, "Dad, it's been a while." The user's voice is captured by a microphone as input. The voice data is sent to the server as output.

[1134] 2. The device sends this audio data in real time to the server. The input is the captured audio data. The output is the audio data that is transferred to the server.

[1135] 3. The server uses conversational AI to analyze the voice data, generate an appropriate response such as "It's been a really long time. How have you been lately?" and send it to the device. The user's voice data is given as input, and the response voice data is generated as output.

[1136] 4. The device returns the generated response to the user. The response audio data sent from the server is given as input. The response is played back to the user through the speaker as output.

[1137] An example of a specific prompt is, "Dad, what did you do today?"

[1138] (Application example 1)

[1139] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1140] Although existing technologies exist for systems that collect personal data from the deceased and recreate natural interactions with the deceased, they lack specific environments and memory-based interactions to improve the user's emotional satisfaction.This invention aims to provide a virtual memorial service that allows users to feel a deeper emotional connection by interacting with the deceased in a specific memorial store.

[1141] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1142] In this invention, the server includes a means for collecting personal data of the deceased while they are alive, a means for analyzing the collected data to extract personal characteristics, a means for generating a 3D virtual model using the characteristics, a means for providing a virtual space in which the user can interact with the generated 3D virtual model, a means for using interactive AI to make the interaction in the virtual space natural, a means for placing a specific memorial store in the virtual space in which the user can move, and a means for placing objects and music related to the deceased in the memorial store and realizing interaction based on memorable episodes. This allows the user to interact with the deceased in a specific environment, further strengthening the emotional connection.

[1143] "Personal data" refers to all information, including audio, video, and linguistic information, captured during a person's lifetime.

[1144] "Features" are information extracted from collected personal data, such as facial features, voice spectrum, speaking rhythm and tone.

[1145] A "3D virtual model" is a three-dimensional human model generated in a virtual space based on collected features.

[1146] A "virtual space" is a digitally recreated environment that a user can access.

[1147] "Conversational AI" is artificial intelligence that is trained to engage in natural conversations using collected voice and language data.

[1148] A "memorial store" is a specific area within the virtual space where objects, music, and stories related to the deceased are interactively arranged.

[1149] "Object" refers to any object or item placed in a virtual space.

[1150] "Music" refers to melodies and songs related to the deceased, and is sound information played within the virtual space.

[1151] "Memorable episodes" are information that provides conversations based on the deceased's hobbies, favorite things, and past events.

[1152] A "dedicated client application" is a specially designed software application used by a user to access a virtual space.

[1153] System program generation

[1154] To realize this invention, we start by collecting data on individuals during their lifetime. This process uses hardware such as a camera (e.g., OpenCV) and an audio recorder. The collected data is categorized into visual, audio, and language information.

[1155] The server receives the collected data and uses a deep learning model (e.g., TensorFlow) to extract features such as facial features, voice spectrum, speaking rhythm, and tone. Based on this, a 3D virtual model is generated. It also trains a conversational AI model based on the episode data and develops algorithms to achieve natural conversations. This model is then used during subsequent conversations.

[1156] Meanwhile, a specific memorial store will be set up in the virtual space, and will be stocked with objects and music related to the deceased. This setting is important for users to feel a deep emotional connection with the deceased. The memorial store will be rendered using Unity.

[1157] Processing procedure description

[1158] Data collection:

[1159] The server interviews and films the deceased. It uses a camera and microphone to record the deceased's facial expressions, movements, voice, and speaking style. For example, if a user requests, "I want to create a virtual human of my father, so please collect data," the device encodes the collected data and sends it to the server using a secure protocol.

[1160] Data analysis and generative AI model training:

[1161] The server analyzes the received data and uses deep learning models to extract features, which are then used to generate a 3D virtual model that reproduces natural facial expressions and movements. Furthermore, the conversational AI is trained to reproduce the speech style and tone of the deceased.

[1162] Episode data storage:

[1163] To ensure a more natural interaction when users visit the memorial store, the device collects information about the deceased's anecdotes and hobbies and stores it in a database, which the conversational AI can quickly reference.

[1164] Virtual Space Deployment:

[1165] A user logs into the virtual space using a dedicated client application, and after the server verifies the user's authentication information, a 3D virtual model of the deceased person appears in the virtual space. For example, a user can log into a "Memorial Cafe" and enjoy interacting with the 3D model of the deceased. The cafe is furnished with objects and music based on the preferences of the deceased.

[1166] Examples of concrete examples and prompts

[1167] As a concrete example, if a user logs in to Memorial Cafe and says, "Mom, you liked this cafe, didn't you?", the conversational AI will respond, "Yes, I really liked it. I love the atmosphere."

[1168] Example prompt sentence:

[1169] User: Mom, do you like this cafe?

[1170] Deceased Model: Yes, I love it. I love the atmosphere here.

[1171] This allows users to experience natural conversation with the deceased within the memorial store, further strengthening their emotional connection.

[1172] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1173] Step 1:

[1174] Collecting personal data before death.

[1175] Input: User requests data collection

[1176] How it works: The device uses a camera and microphone to record the facial expressions, movements, voice, and speaking style of the deceased. For example, based on a user's request to "create a virtual model of my father," the device collects data for approximately seven hours.

[1177] Output: A dataset containing visual, audio, and linguistic information

[1178] Step 2:

[1179] The collected data is sent to a server using a secure protocol.

[1180] Input: Dataset collected in step 1

[1181] Specific operation: The device encodes the data and sends it to the server using a secure protocol (e.g., SSL / TLS).

[1182] Output: The encoded data reaches the server

[1183] Step 3:

[1184] The server parses the data it receives.

[1185] Input: Encoded data

[1186] How it works: The server uses a deep learning model (e.g., TensorFlow) to decode the data and extract features such as facial features, voice spectrum, speaking rhythm and tone.

[1187] Output: A dataset of extracted facial feature points, voice spectrum, speaking rhythm and tone.

[1188] Step 4:

[1189] Generate a 3D virtual model.

[1190] Input: Feature data extracted in step 3

[1191] Specific operation: The server uses Unity to generate a 3D virtual model based on the extracted feature data, reproducing natural facial expressions and movements.

[1192] Output: Generated 3D virtual model

[1193] Step 5:

[1194] Train a conversational AI model.

[1195] Input: Collected speech and language data

[1196] How it works: The server uses collected voice and language data to train the conversational AI so that it can converse naturally, and uses deep learning algorithms to build a model that reproduces the speech style and tone of the deceased.

[1197] Output: A trained conversational AI model

[1198] Step 6:

[1199] Episode data is stored and indexed in a database.

[1200] Input: Information about episodes and hobbies collected from users

[1201] Specific operation: The device organizes information about episodes and hobbies, stores it in a database, and indexes it for quick reference by the conversational AI.

[1202] Output: Indexed episode data

[1203] Step 7:

[1204] A user logs into the virtual space using a dedicated client application.

[1205] Input: User credentials

[1206] How it works: A user logs into a virtual space using a dedicated client application. The server verifies the authentication information and grants the user access.

[1207] Output: The user can access the virtual space.

[1208] Step 8:

[1209] A memorial store will be placed in the virtual space, and a 3D virtual model of the deceased person will appear.

[1210] Input: User's location, 3D virtual model generated in step 4

[1211] Specific operation: The server places a memorial store in the virtual space and makes a 3D virtual model of the deceased person appear based on the user's location information.

[1212] Output: 3D virtual model of the deceased placed in a memorial store

[1213] Step 9:

[1214] The user engages in natural dialogue with conversational AI.

[1215] Input: User voice input, the conversational AI model trained in step 5, and episode data from step 6

[1216] How it works: When a user speaks to a model of the deceased person, the device sends the voice data in real time to the server. The server uses a conversational AI model to generate a response and sends it to the device. The device then returns the response to the user, realizing a natural dialogue.

[1217] Output: A natural dialogue experience between the user and conversational AI

[1218] Specific examples

[1219] For example, if a user logs in to Memorial Cafe and says, "Mom, do you like this cafe?" the conversational AI will respond, "Yes, I love it. I love the atmosphere here."

[1220] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1221] This invention combines a system that collects personal data from a person's lifetime, analyzes the collected data to generate a 3D virtual model, and enables users to experience natural conversations with the deceased using interactive AI within the metaverse space, with an emotion engine that recognizes the user's emotions.

[1222] System program generation and implementation method

[1223] Generate 3D virtual models from data collection

[1224] 1. A user requests the collection of data on a deceased person. The request is recorded and preparations for collection are made.

[1225] 2. The device interviews and photographs the deceased. Using a camera and microphone, it records the deceased's facial expressions, movements, voice, and speaking style for approximately seven hours. Each piece of data is saved with a timestamp.

[1226] 3. The device encodes the collected data and sends it to the server using a secure protocol, along with a checksum to verify the data's integrity and completeness.

[1227] 4. The server analyzes the received data, using deep learning models to extract features such as facial features, voice spectrum, and speaking rhythm and tone.

[1228] 5. The server performs 3D modeling based on the feature data. The extracted features are used to generate a 3D virtual model that reproduces the natural facial expressions and movements of the deceased.

[1229] 6. The server trains the conversational AI. Using the collected voice and language data, the AI ​​model is trained to reproduce the speaking style and tone of the deceased, specifically understanding the deceased's voice patterns and conversational flow.

[1230] 7. The device collects information about episodes, hobbies, etc. For example, during an interview, it might ask, "What is your favorite food?" and record the answer.

[1231] 8. The device sends the episode data to the server. The data is categorized into episodes, hobbies, favorite things, etc. and organized according to a format.

[1232] 9. The server stores the episode data in a database, where it is indexed and made available for quick reference by the conversational AI.

[1233] Expanding into the Metaverse

[1234] 1. A user logs into the metaverse space using a dedicated client application. The login information is sent to the server via authentication methods.

[1235] 2. The server verifies the user's credentials and grants access. If authentication is successful, access is established within the metaverse space.

[1236] 3. The server creates a 3D virtual model of the deceased person in the Metaverse space based on the user's location information, placing it in an appropriate position taking into account the user's line of sight and location.

[1237] 4. The user talks to the deceased, for example, "Dad, how are you?"

[1238] 5. The device transmits the user's voice in real time to the server, where it is recorded and stored for analysis.

[1239] 6. The server uses conversational AI to analyze the voice data and generate an appropriate response, such as "I've been doing well. How have you been lately?"

[1240] 7. The server sends the generated response to the terminal, which then outputs the response to the user.

[1241] Incorporating an emotion engine

[1242] 1. The device captures the user's facial expressions and tone of voice in real time using a camera and microphone.

[1243] 2. The device sends the captured data to the emotion engine, which then recognizes the user's emotion based on this data. For example, if a smile is detected, it will recognize that the user is feeling "joy."

[1244] 3. The server receives the emotion data sent from the emotion engine and reflects it in the conversational AI. For example, if the user is recognized as sad, the conversational AI generates a response such as "Cheer up."

[1245] Specific examples

[1246] For example, a user logs into the metaverse space and says, "Dad, it's been a while."

[1247] The device transmits this audio data to the server in real time.

[1248] The emotion engine analyzes and recognizes that the user's tone of voice is sad.

[1249] The server uses conversational AI to analyze the voice data and, referring to the analysis results of the emotion engine, generates an appropriate response such as, "It's been a while. What's wrong? You seem down."

[1250] The server sends the generated response to the terminal, which then speaks the response to the user.

[1251] This allows users to have more emotionally enriching interactions with deceased loved ones within the metaverse space.

[1252] The processing flow will be explained below.

[1253] Step 1:

[1254] The user requests the collection of data on the deceased. The request is confirmed and preparations for collection are made.

[1255] Step 2:

[1256] The device interviews and photographs the deceased. Using a camera and microphone, it records the deceased's facial expressions, movements, voice, and speaking style for approximately seven hours. The collected data is then saved with a timestamp.

[1257] Step 3:

[1258] The device encodes the collected data and transmits it to the server using a secure protocol, along with a checksum to ensure data integrity and completeness.

[1259] Step 4:

[1260] The server analyzes the received data and uses deep learning models to extract features such as facial features, voice spectrum, speaking rhythm and tone.

[1261] Step 5:

[1262] The server performs 3D modeling based on the feature data, and generates a 3D virtual model using the extracted features to reproduce the natural facial expressions and movements of the deceased.

[1263] Step 6:

[1264] The server trains the conversational AI. Using the collected voice and language data, the AI ​​model is trained to reproduce the speaking style and tone of the deceased. Specifically, it analyzes and learns from the deceased's voice patterns and conversational flow.

[1265] Step 7:

[1266] The device collects information such as episodes and hobbies. For example, it asks, "What is your favorite food?" and records the answer.

[1267] Step 8:

[1268] The device sends episode data to the server, which organizes the data into categories such as episodes, hobbies, and favorite things.

[1269] Step 9:

[1270] The server stores the episode data in a database, where it is indexed and prepared for fast and efficient reference by the conversational AI.

[1271] Step 10:

[1272] A user logs into the metaverse space using a dedicated client application, and the login information is sent to the server via authentication means.

[1273] Step 11:

[1274] The server verifies the user's credentials and grants access. If authentication is successful, access is established within the metaverse space.

[1275] Step 12:

[1276] The server creates a 3D virtual model of the deceased person in the Metaverse space based on the user's location information, and adjusts the position of the virtual model based on the user's line of sight and location.

[1277] Step 13:

[1278] The user talks to the deceased, for example, saying, "Dad, it's been a long time."

[1279] Step 14:

[1280] The device transmits the user's voice in real time to a server, where the voice data is recorded and stored for analysis.

[1281] Step 15:

[1282] The server uses conversational AI to analyze the voice data and generate an appropriate response, such as "It's been a while, how have you been?"

[1283] Step 16:

[1284] The device captures the user's facial expressions and tone of voice in real time using a camera and microphone.

[1285] Step 17:

[1286] The device captures facial and voice data and sends it to the emotion engine, which analyzes this data and recognizes the user's emotions.

[1287] Step 18:

[1288] The server receives the emotion data sent from the emotion engine and reflects it in the conversational AI. For example, if the user looks sad, the conversational AI will generate a response such as "Cheer up, what happened recently?"

[1289] Step 19:

[1290] The server sends the generated response to the terminal, which then provides an audio output to the user.

[1291] This allows users to interact with deceased loved ones in a more emotionally rich and natural way within the metaverse space.

[1292] Example 2

[1293] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1294] Conventional systems have had difficulty using data from deceased people to recreate conversations with them in a virtual space. Furthermore, they lacked the ability to generate responses based on the user's emotions, making the conversations less emotionally rich and natural.

[1295] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for collecting personal data before death, means for analyzing the collected data to extract personal characteristics, means for generating a 3D virtual model using the characteristics, means for providing a virtual space in which the user can interact with the generated 3D virtual model, means for using interactive artificial intelligence to make the interaction in the virtual space natural, and means for recognizing the user's emotions and flexibly adjusting the response content. This allows the user to experience an emotionally rich and natural interaction with the deceased person in the virtual space.

[1296] "Personal data" is information about a person that is collected during their lifetime, including photographs, videos, audio files, and written material.

[1297] "Feature extraction means" refers to a process or technique for identifying and analyzing individual characteristics or features from collected data.

[1298] A "3D virtual model" is a three-dimensional digital representation of an individual's appearance, gestures, movements, etc., based on collected data and extracted features.

[1299] "Virtual space" refers to a virtual environment that users can access through a computer or dedicated client application, including the metaverse and VR space.

[1300] "Conversational artificial intelligence" refers to a computer program or system that can use natural language analysis and generation techniques to engage in natural conversations with humans.

[1301] "Means for recognizing emotions and flexibly adjusting response content" refers to a function that analyzes the user's facial expressions, tone of voice, and other emotional indicators in real time and dynamically changes the response content based on the results.

[1302] "Episode data" refers to information based on specific events, hobbies, or specific memories related to the deceased.

[1303] "Dedicated client application" refers to software specifically designed to access virtual spaces and interface with interactive AI and 3D virtual models.

[1304] MODE FOR CARRYING OUT THE INVENTION

[1305] This invention is a system that collects personal data from a person's life and analyzes it to generate a 3D virtual model. This system uses conversational artificial intelligence to allow users to experience natural interactions with the deceased in the metaverse space. It also combines an emotion engine that recognizes the user's emotions to realize more emotionally rich interactions.

[1306] Hardware and software used

[1307] Hardware:

[1308] Camera: Used to record the facial expressions and movements of the deceased.

[1309] Microphone: Used to record the voice and speaking style of the deceased.

[1310] Devices (e.g., computers, smartphones): used for data collection and transmission.

[1311] Server: Used for data processing, analysis, storage, AI training, and virtual space management.

[1312] software:

[1313] Dedicated client application: Software that allows users to access the virtual space.

[1314] Data collection software: collects data from the camera and microphone, encodes it, and sends it to a server.

[1315] Deep learning models: Algorithms for data analysis and feature extraction.

[1316] 3D modeling tools (e.g. Blender): Generate 3D virtual models based on collected feature data.

[1317] Conversational Artificial Intelligence (AI): Enables natural dialogue with users.

[1318] Emotion engine: Analyzes user emotions and reflects them in the dialogue AI.

[1319] Examples of concrete examples and prompts

[1320] For example, suppose a user logs into the metaverse space using a dedicated client application and says, "Dad, it's been a while." The specific processing at this time is as follows.

[1321] 1. The device sends this voice data to the server in real time.

[1322] 2. The server uses conversational artificial intelligence to analyze the voice data and generates an appropriate response by referring to the analysis results of the emotion engine.

[1323] 3. For example, the emotion engine analyzes and recognizes that the user's tone of voice is sad, and the server generates a response such as, "It's been a while, what's wrong? You seem depressed."

[1324] 4. The server sends the generated response to the terminal, which then speaks the response to the user.

[1325] In this way, users can have emotionally enriching interactions with deceased loved ones within the metaverse space.

[1326] Example prompt sentence:

[1327] "Dad, how have you been?"

[1328] "How are you feeling today?"

[1329] "I want to talk about recent events."

[1330] By inputting these prompt sentences, the conversational AI can understand the user's intentions and emotions and generate the most appropriate response.

[1331] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1332] Step 1: (Data collection request)

[1333] The user requests the collection of data on the deceased, specifying specific items to be collected, such as photos, videos, and audio files of the deceased.

[1334] The device records the request and begins preparations for collection, preparing the camera and microphone, and confirming the data format to be collected.

[1335] Input: User request.

[1336] Output: Collection readiness status.

[1337] Step 2: (Interview and filming)

[1338] The device uses a camera and microphone to interview and film the deceased, and the data collected includes facial expressions, movements, voice, and speaking style.

[1339] The device stores all data with a timestamp, splitting it into video clips every 10 minutes, for example.

[1340] Input: Real-time data from camera and microphone.

[1341] Output: Collected data with timestamp.

[1342] Step 3: (Data encoding and transmission)

[1343] The device encodes the collected data and transmits it to a server using a secure protocol (e.g., HTTPS).

[1344] The terminal generates a checksum to verify the integrity and completeness of the data and also sends this to the server.

[1345] Input: Collected data before encoding.

[1346] Output: The encoded data and a checksum.

[1347] Step 4: (Data Analysis)

[1348] The server then analyzes the received data using a deep learning model that extracts features such as facial features, voice spectrum, and speaking rhythm and tone.

[1349] The server specifically analyzes the deceased person's unique movements and expressions and stores them in a database.

[1350] Input: The encoded data.

[1351] Output: Extracted feature data.

[1352] Step 5: (3D modeling)

[1353] The server generates a 3D virtual model based on the extracted feature data, carefully adjusting it to reproduce the natural facial expressions and movements of the deceased.

[1354] The server uses a 3D modeling tool (e.g. Blender) to see in real time what needs to be fixed.

[1355] Input: Extracted feature data.

[1356] Output: 3D virtual model.

[1357] Step 6: (Training the conversational AI)

[1358] The server uses the collected speech and language data to train the conversational AI, specifically learning patterns to replicate the speech and narration of the deceased.

[1359] The server mimics voice patterns and adjusts them to create a natural conversational flow.

[1360] Input: Audio and language data.

[1361] Output: A trained conversational AI model.

[1362] Step 7: (Collect episode information)

[1363] Users provide information about the deceased, such as anecdotes, hobbies, and interests, in the form of an interview.

[1364] The device records this information and also saves it as text data.

[1365] Input: The answer given by the user.

[1366] Output: Episode data.

[1367] Step 8: (Submit episode data)

[1368] The episode data collected by the device is organized according to a format and sent to the server.

[1369] The device classifies the data (episodes, hobbies, interests) before sending it.

[1370] Input: Episode data.

[1371] Output: Organized episode data.

[1372] Step 9: (Storing episode data)

[1373] The server indexes and stores the received episode data in a database.

[1374] The server applies a search algorithm to allow quick access to the data.

[1375] Input: Organized episode data.

[1376] Output: Indexed episode data in a database.

[1377] Step 10: (Login and Authentication)

[1378] A user logs into the virtual space using a dedicated client application, and the login information is sent from the client to the server.

[1379] The server verifies the user's credentials and grants access.

[1380] Input: User login information.

[1381] Output: Authentication check and access granted.

[1382] Step 11: (Placing the virtual model)

[1383] The server creates a 3D virtual model of the deceased person in a virtual space based on the user's location information.

[1384] The server places the virtual model in an appropriate position, taking into account the user's line of sight and position.

[1385] Input: User's location.

[1386] Output: Positioning information of the virtual model in the virtual space.

[1387] Step 12: (User interaction)

[1388] The user talks to the deceased, for example, asking, "Dad, how are you?"

[1389] The device transmits this audio data to the server in real time.

[1390] The server uses conversational artificial intelligence to analyze the voice data and generates an appropriate response by referring to the analysis results of the emotion engine.

[1391] Input: User's question, voice data.

[1392] Output: A response from the conversational AI.

[1393] Step 13: (Capturing the user's facial expressions and voice)

[1394] The device captures the user's facial expressions and tone of voice in real time using a camera and microphone.

[1395] The device transmits the captured data to the emotion engine in real time.

[1396] Input: User's facial expression data and tone of voice.

[1397] Output: Emotion data sent to the emotion engine.

[1398] Step 14: (Analyze emotion data)

[1399] The server analyzes the user's emotions based on the data sent from the emotion engine. For example, it recognizes "joy" by detecting a smile.

[1400] Input: User emotion data from the emotion engine.

[1401] Output: Analyzed user emotion information.

[1402] Step 15: (Generating and reflecting dialogue responses)

[1403] The server generates a response from the conversational AI based on the emotional data. For example, if the server recognizes that the user is sad, it generates a response such as "Cheer up."

[1404] The server sends the generated response to the terminal, which then outputs it to the user.

[1405] Input: Parsed user emotion information.

[1406] Output: AI response based on emotion.

[1407] (Application example 2)

[1408] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1409] In modern society, people often seek emotional connections through conversations with deceased loved ones, but systems that can provide such experiences are not yet widespread. Furthermore, there are no systems incorporating conversational AI that can recognize and reflect emotions in real time, making it difficult to understand the user's emotions and provide natural conversations accordingly. Furthermore, for users to naturally converse with deceased loved ones in virtual spaces, it is necessary to manage and analyze episode data and emotional data, and an efficient system for this purpose is needed.

[1410] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1411] In this invention, the server includes means for collecting personal data before death, means for analyzing the collected data to extract personal characteristics, means for generating a 3D virtual model using the characteristics, means for providing a virtual space in which a user can interact with the generated 3D virtual model, means for using interactive AI to make the interaction in the virtual space natural, means for capturing and analyzing the user's voice and emotions in real time, and means for generating a response that reflects the analyzed emotional data, thereby enabling a user to have an emotional experience through interaction with the deceased.

[1412] "Personal data" refers to information provided by an individual during their lifetime, such as their voice, facial expressions, movements, and speaking style.

[1413] "Features" refers to identifying information such as facial features, voice spectrum, speaking rhythm and tone extracted from collected personal data.

[1414] "3D Virtual Model" means a three-dimensional model of an individual displayed in a virtual space, generated based on collected characteristics.

[1415] "Virtual space" refers to a virtual reality environment accessible to a user.

[1416] "Conversational AI" refers to an artificial intelligence system that uses natural language processing to engage in dialogue with users.

[1417] "Means for capturing and analyzing voice and emotions in real time" refers to the function of collecting the user's voice and facial expressions in real time using a voice input device or camera, and analyzing them using an emotion analysis engine.

[1418] "Emotion data" refers to information about emotions analyzed from the user's tone of voice, facial expressions, etc.

[1419] "Means for generating a response" refers to an appropriate response that corresponds to the user's emotions, which is generated by the conversational AI based on the analyzed emotional data.

[1420] This invention is a system that allows a user to have natural conversations with a deceased person in a virtual space based on personal data collected by the user while he or she was alive. An embodiment of this system will be described in detail below.

[1421] First, the user uses a device to collect data on the deceased. Data such as the deceased's facial expressions, movements, voice, and speaking style is recorded using a camera and microphone for approximately seven hours. This collected data is then sent to a server using a secure protocol. The server then uses a deep learning model to analyze the data and extract individual features. Based on the extracted features, a realistic and natural-looking 3D virtual model is then generated.

[1422] The generated 3D virtual model is displayed in a virtual space accessible to users. Users access this virtual space using a dedicated client application on their smartphone or head-mounted display. When the user speaks to the deceased, their voice is transmitted in real time to a server, where conversational AI analyzes the voice and generates an appropriate response.

[1423] The system also incorporates an emotion engine that captures and analyzes the user's voice and facial expressions in real time. The emotion engine analyzes emotions from the user's tone of voice and facial expressions and reflects them in the responses generated by the conversational AI. This allows the user to experience a more natural and emotional dialogue. For example, if a user says, "Dad, it's been a while," the emotion engine will analyze that the user's tone of voice sounds sad, and the conversational AI will generate a response such as, "It's been a while. What's wrong? You seem down."

[1424] The hardware used includes a smartphone, a head-mounted display, a camera, and a microphone, while the software used includes Unity (a 3D game engine), TensorFlow (a deep learning framework for sentiment analysis), Google Cloud Speech-to-Text API, and Azure Cognitive Services (conversational AI).

[1425] An example prompt is:

[1426] "User prompt"

[1427] We would like to request the development of an application that provides an emotional experience through interaction with the deceased, such as:

[1428] The user collects data about the deceased and conducts an interview using a camera and microphone. The collected data is sent to a server, and a 3D virtual model is generated using deep learning. The user logs in to the virtual space and begins a dialogue. The emotion engine analyzes the user's emotions in real time, and the conversational AI generates a response based on the user's emotions. Please help us develop an application that allows users to have an emotionally rich experience through natural dialogue.

[1429] In this way, a system is realized that allows users to naturally engage in emotional dialogue with the deceased in a virtual space.

[1430] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1431] Step 1:

[1432] The device collects data about the deceased.

[1433] Specifically, the user uses a device to record the deceased's facial expressions, movements, voice, speaking style, etc. for approximately seven hours using a camera and microphone. This data is time-stamped, and the audio data is saved in WAV format, and the video data in MP4 format. The input is the audio and video data of the deceased, and the output is the saved multimedia data.

[1434] Step 2:

[1435] The terminal transmits the collected data to the server.

[1436] The collected data is sent to the server using a secure protocol (e.g. HTTPS), along with a checksum to verify the integrity and completeness of the data. The input is the stored multimedia data, and the output is the data received by the server.

[1437] Step 3:

[1438] The server analyzes the data and extracts personal characteristics.

[1439] A deep learning model (for example, a model using TensorFlow) is used to analyze facial feature points, voice spectrum, speaking rhythm and tone, etc. The analysis results are saved as feature data. The input is the multimedia data received by the server, and the output is the extracted feature data.

[1440] Step 4:

[1441] The server generates a 3D virtual model based on the feature data.

[1442] Using the extracted features, 3D modeling software (e.g., Unity) is used to generate a 3D virtual model that reproduces the deceased's natural facial expressions and movements. The input is the feature data, and the output is the 3D virtual model.

[1443] Step 5:

[1444] Users access the virtual space using a dedicated client application.

[1445] Users log in to the Metaverse space using a smartphone or head-mounted display. The login information is sent to the server via authentication means. The input is the user's authentication information, and the output is the access permission granted upon successful authentication.

[1446] Step 6:

[1447] The server makes a 3D virtual model of the deceased appear in the virtual space based on the user's location information.

[1448] The 3D virtual model is placed in an appropriate position taking into account the user's line of sight and position information. The input is the user's position information, and the output is the position of the 3D virtual model in the metaverse space.

[1449] Step 7:

[1450] The terminal transmits the user's voice to the server in real time.

[1451] When a user speaks to the deceased, the audio data is captured in real time and sent to a server, where it is converted to text using the Google Cloud Speech-to-Text API. The input is the user's audio data, and the output is the converted audio data.

[1452] Step 8:

[1453] The server uses conversational AI to analyze the voice data and generate an appropriate response.

[1454] The server uses conversational AI to generate appropriate responses based on user input. It uses Azure Cognitive Services. The input is the converted voice data, and the output is the generated response text.

[1455] Step 9:

[1456] The terminal then vocalizes the generated response to the user.

[1457] The response text from the server is converted into audio format and played back to the user in real time. The input is the generated response text, and the output is the audio output.

[1458] Step 10:

[1459] The device captures the user's facial expressions and tone of voice in real time and sends them to the emotion engine.

[1460] It uses a camera and microphone to capture the user's facial expressions and tone of voice and sends them to the emotion engine, which then analyzes them using a deep learning model. The input is the user's facial and voice data, and the output is emotion data.

[1461] Step 11:

[1462] The server reflects the emotional data received from the emotion engine in the conversational AI.

[1463] The conversational AI adjusts the response it generates depending on the user's emotions. For example, if the user is recognized as sad, the response may be "Cheer up." The input is emotional data, and the output is the adjusted response text.

[1464] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1465] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1466] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1467] [Fourth embodiment]

[1468] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1469] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1470] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1471] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1472] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1473] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1474] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1475] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1476] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1477] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1478] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1479] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1480] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1481] This invention is a system that collects personal data from people while they are alive, analyzes the collected data to generate a 3D virtual model, and uses interactive AI in a metaverse space to allow users to experience natural conversations with the deceased.

[1482] System program generation and implementation method

[1483] Data collection

[1484] 1. A user requests data collection for a deceased person. For example, a user might say, "I want to create a virtual human of my father, so please collect his data."

[1485] 2. The device interviews and photographs the deceased. Using a camera and microphone, it records the father's facial expressions, movements, voice, and speaking style for approximately seven hours. The collected data is categorized into visual, audio, and linguistic information.

[1486] 3. The device encodes the collected data and sends it to the server using a secure protocol.

[1487] Data analysis and training of generative AI models

[1488] 1. The server analyzes the received data and uses deep learning models to extract features such as facial features, voice spectrum, speaking rhythm and tone.

[1489] 2. The server performs 3D modeling, generating a 3D virtual model that reproduces natural facial expressions and movements based on the extracted features.

[1490] 3. The server trains the conversational AI, learning from the collected voice and language data to recreate the speaking style and tone of the deceased.

[1491] Storing episode data

[1492] 1. The device collects information such as episodes, hobbies, etc. For example, it asks, "What is your favorite food?" and the father answers, "I love sushi."

[1493] 2. The device organizes this information according to a format and sends it to the server.

[1494] 3. The server stores the episode data in a database and indexes it for quick reference by the conversational AI.

[1495] Expanding into the Metaverse

[1496] 1. A user logs into the metaverse space using a dedicated client application.

[1497] 2. The server verifies the user's credentials and grants access.

[1498] 3. The server creates a 3D virtual model of the deceased person in the Metaverse space and adjusts the position of the virtual model based on the user's location information.

[1499] Specific examples

[1500] The user logs into the metaverse space and says, "Dad, it's been a while."

[1501] The device transmits this audio data to the server in real time.

[1502] The server uses conversational AI to analyze the voice data, generate an appropriate response such as, "It's been a really long time. How have you been lately?" and send it to the device.

[1503] The terminal returns the generated response to the user, allowing the user to experience a natural conversation with the deceased.

[1504] This allows users to enjoy natural conversations with their deceased loved ones within the metaverse space, providing a means to ease emotional pain.

[1505] The processing flow will be explained below.

[1506] Step 1:

[1507] The user requests the collection of data on the deceased. The request is recorded and the collection is prepared.

[1508] Step 2:

[1509] The device interviews and photographs the deceased. Using a camera and microphone, it records the deceased's facial expressions, movements, voice, and speaking style for approximately seven hours. Each piece of data is saved with a timestamp.

[1510] Step 3:

[1511] The device encodes the collected data and sends it to the server using a secure protocol, along with a checksum to verify the data's integrity and completeness.

[1512] Step 4:

[1513] The server analyzes the received data, first extracting features such as facial features, voice spectrum, and speaking rhythm and tone using deep learning algorithms.

[1514] Step 5:

[1515] The server performs 3D modeling based on the feature data, and generates a 3D virtual model using the extracted features to reproduce the natural facial expressions and movements of the deceased.

[1516] Step 6:

[1517] The server trains the conversational AI, using the collected voice and language data to train the AI ​​model to reproduce the speech style and tone of the deceased, specifically understanding the voice patterns and flow of the conversation.

[1518] Step 7:

[1519] The device separately collects information such as anecdotes and hobbies. For example, during an interview, it might ask, "What is your favorite food?" and record the answer.

[1520] Step 8:

[1521] The device sends episode data to the server, where the data is categorized into episodes, hobbies, favorite things, etc. and organized according to a format.

[1522] Step 9:

[1523] The server stores the episode data in a database, where it is indexed and made available for quick reference by the conversational AI.

[1524] Step 10:

[1525] A user logs into the metaverse space using a dedicated client application, and the login information is sent to the server via authentication means.

[1526] Step 11:

[1527] The server verifies the user's credentials and grants access. If authentication is successful, access is established within the metaverse space.

[1528] Step 12:

[1529] The server creates a 3D virtual model of the deceased person in the Metaverse space based on the user's location information, and places the model in an appropriate location, taking into account the user's line of sight and location.

[1530] Step 13:

[1531] The user talks to the deceased, for example, saying, "Dad, how are you?"

[1532] Step 14:

[1533] The device transmits the user's voice in real time to a server, where the voice data is recorded and then stored for analysis.

[1534] Step 15:

[1535] The server uses conversational AI to analyze the voice data and generate an appropriate response, such as "I've been doing well. How have you been lately?"

[1536] Step 16:

[1537] The server sends the generated response to the terminal, which then outputs the response to the user.

[1538] This allows users to experience natural interactions with the deceased within the metaverse space.

[1539] Example 1

[1540] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1541] There is a need for a system that allows users to communicate with deceased individuals and ease their emotional pain by recreating memories and conversations with them. However, existing technologies are insufficient in recreating the deceased, making it difficult to achieve natural conversations. To solve this problem, a system that accurately recreates the characteristics of the deceased and allows users to experience natural conversations is needed.

[1542] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1543] In this invention, the server includes: a means for a user to request data collection of the deceased; a means for recording the deceased's facial expressions, movements, voice, and speaking style using a camera and microphone; a means for encoding the recorded data and transmitting it to the server using a secure protocol; a means for the server to analyze the received data using a deep learning model and extract features; a means for generating a 3D virtual model based on the extracted features; a means for training an interactive AI using voice data and language data; a means for collecting information such as episodes and hobbies and storing it in a database; a means for a user to access the metaverse space using a dedicated client application; a means for the server to create a 3D virtual model of the deceased in the metaverse space and adjust it based on the user's location information; and a means for realizing natural conversation using the interactive AI, allowing users to enjoy natural conversations with the deceased in the metaverse space.

[1544] "User" refers to an individual who utilizes the system to collect data about or interact with a deceased person.

[1545] "Means for requesting data collection" refers to a function that allows a user to request data collection of a deceased person from the system.

[1546] "Camera and microphone" refers to visual and audio input devices used to record the facial expressions, movements, voice, and speaking style of the deceased.

[1547] "Means of recording" refers to the ability to use cameras and microphones to digitally record the facial expressions, movements, voice, and speaking style of the deceased.

[1548] "Encoding" refers to the process of compressing and converting recorded data into a format that can be transmitted efficiently.

[1549] "Secure protocol" refers to a communication protocol for ensuring the security of data communications, and includes SSL and TLS.

[1550] A "deep learning model" refers to an artificial intelligence technology that uses neural networks to analyze data and extract features.

[1551] "Feature extraction" refers to finding important patterns or attributes in data and using them for analysis.

[1552] A "3D virtual model" refers to a digital model that reproduces the features of the deceased and displays them in three-dimensional space.

[1553] "Conversational AI" refers to artificial intelligence technology that can mimic natural human conversation.

[1554] "Anecdotes and hobbies" refers to information that reveals the deceased's personal experiences and interests.

[1555] A "database" refers to an information management system that stores data in a structured format that makes it easy to search and reference.

[1556] "Client Application" refers to software that a user uses to access the metaverse space.

[1557] "Metaverse space" refers to a virtual reality environment constructed digitally, including a space that users can interact with.

[1558] "Location information" refers to data indicating the current coordinates or location of the user and virtual model.

[1559] This invention is a system that collects personal data from people while they are alive, analyzes the collected data to generate a 3D virtual model, and allows users to experience natural conversations with the deceased using interactive AI in a metaverse space.

[1560] Data collection

[1561] 1. A user requests data collection for a deceased person. For example, they might request, "I would like to create a virtual model of my father, so please collect his data." This request is made via a dedicated web portal or application.

[1562] 2. The device interviews and films the deceased. It uses a camera and microphone to record the father's facial expressions, movements, voice, and speaking style for approximately seven hours. The specific hardware used includes a high-resolution camera and a directional microphone. The collected data is categorized into visual, audio, and language information.

[1563] 3. The device encodes the collected data and sends it to the server using a secure protocol. A data encoder is used for this encoding, and the SSL / TLS protocol is used to ensure secure data transmission. Encoding is done using a format such as H.264 or FLAC.

[1564] Data analysis and training of generative AI models

[1565] 1. The server analyzes the received data and uses a deep learning model (e.g., TensorFlow or PyTorch) to extract features such as facial features, voice spectrum, speaking rhythm and tone.

[1566] 2. The server performs 3D modeling, using Blender, Maya, or other software to generate a 3D virtual model that reproduces natural facial expressions and movements based on the extracted features.

[1567] 3. The server trains the conversational AI, learning from the collected voice and language data and using language models such as GPT-3 and BERT to recreate the speaking style and tone of the deceased, enabling natural conversation.

[1568] Storing episode data

[1569] 1. The device collects information such as episodes, hobbies, etc. For example, it asks, "What is your favorite food?" and the father answers, "I like sushi."

[1570] 2. The terminal organizes the collected episode information according to a format and sends it to the server. This organization is done using a formatting tool.

[1571] 3. The server stores the episode data in a database and indexes it for quick reference by the conversational AI. Database systems used include MySQL and MongoDB.

[1572] Expanding into the Metaverse

[1573] 1. A user logs into the Metaverse space using a dedicated client application, including the client application for Oculus Rift.

[1574] 2. The server verifies the user's authentication information and grants access. The authentication system used is OAuth 2.0 or similar.

[1575] 3. The server creates a 3D virtual model of the deceased person in the Metaverse space, and adjusts the virtual model's position appropriately using Unity or Unreal Engine based on the user's location information.

[1576] Specific examples

[1577] 1. The user logs into the metaverse space and says, "Dad, it's been a while."

[1578] 2. The device sends this audio data to the server in real time.

[1579] 3. The server uses conversational AI to analyze the voice data, generate an appropriate response such as, "It's been a really long time. How have you been lately?" and send it to the device.

[1580] 4. The device returns the generated response to the user, allowing the user to experience a natural conversation with the deceased.

[1581] An example of a specific prompt is, "Dad, what did you do today?"

[1582] This allows users to enjoy natural conversations with their deceased loved ones within the metaverse space, providing a means to ease emotional pain.

[1583] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1584] System program processing flow

[1585] Data Collection Steps

[1586] Step 1:

[1587] A user requests data collection of a deceased person using a dedicated web portal or application. As input, the user enters details such as the name of the deceased, their relationship, and the information they wish to collect. As output, a request form is sent to the server.

[1588] Step 2:

[1589] The server receives the user's request and prepares the data collection. The user's request information is transmitted to the server as input. The collection plan is sent to the terminal as output.

[1590] Step 3:

[1591] The device uses a camera and microphone to record the facial expressions, movements, voice, and speaking style of the deceased. The camera and microphone capture the deceased's visual, audio, and linguistic information as input. The captured data is then stored on the device as output. Specifically, the camera captures high-resolution video and the microphone records audio.

[1592] Step 4:

[1593] The data recorded by the device is encoded and sent to the server using a secure protocol. Visual, audio, and language information collected as input is encoded. Encrypted data is sent to the server as output. Specifically, the data is encoded into H.264 or FLAC format and sent using the SSL / TLS protocol.

[1594] Steps for Data Analysis and Generative AI Model Training

[1595] Step 5:

[1596] The server analyzes the received data using a deep learning model. Encoded data securely transmitted to the server is provided as input. Feature-extracted data is generated as output. Specifically, TensorFlow and PyTorch are used to analyze features such as facial features, voice spectrum, speaking rhythm and tone.

[1597] Step 6:

[1598] The server generates a 3D virtual model based on the extracted features. The analyzed feature data is given as input. The 3D virtual model is generated as output. Specifically, the model is constructed using Blender or Maya, and skeletal animation is implemented.

[1599] Step 7:

[1600] The server trains the conversational AI. The collected speech and language data is used as input. The output is a trained conversational AI model that improves its ability to reproduce the speech style and tone of the deceased. Specifically, training is performed using GPT-3 and BERT to enhance the model's ability to mimic natural conversation.

[1601] Steps for storing episode data

[1602] Step 8:

[1603] The device collects information from the deceased, such as episodes and hobbies. The user or interviewer asks questions as input, and the deceased's responses are collected as voice data. The output is the collected episode information organized. Specifically, the system works by collecting a situation in which the interviewer asks, "What is your favorite food?" and the deceased answers, "I like sushi."

[1604] Step 9:

[1605] The episode information collected by the device is organized according to a format and sent to the server. Unorganized episode information is given as input, and formatted data is sent to the server as output. Specifically, the data is organized into JSON or XML format using a formatting tool and sent to the server.

[1606] Step 10:

[1607] The server stores the episode data in a database and indexes it for quick reference by the conversational AI. Formatted episode data is given as input. The output is episode information stored in a database and indexed. Specifically, data is organized and indexed using MySQL or MongoDB.

[1608] Steps for expanding into the metaverse space

[1609] Step 11:

[1610] A user logs into the metaverse space using a dedicated client application. The user's authentication information (username and password) is used as input. The output is the initiation of a login session. Specifically, the client application takes the user's credentials as input and sends them to the authentication system.

[1611] Step 12:

[1612] The server verifies the user's authentication information and grants access. User credentials are provided as input. An authentication result is generated as output. Specifically, an authentication system such as OAuth 2.0 is used, and a session begins if authentication is successful.

[1613] Step 13:

[1614] The server creates a 3D virtual model of the deceased person in the Metaverse space and adjusts it based on the user's location information. The user's current location information and 3D model information are provided as input. The output is a 3D model that has been adjusted to correspond to the user. Specifically, Unity or Unreal Engine is used to place the virtual model at a specified location in the Metaverse space.

[1615] Specific examples

[1616] 1. A user logs in to the metaverse space and says, "Dad, it's been a while." The user's voice is captured by a microphone as input. The voice data is sent to the server as output.

[1617] 2. The device sends this audio data in real time to the server. The input is the captured audio data. The output is the audio data that is transferred to the server.

[1618] 3. The server uses conversational AI to analyze the voice data, generate an appropriate response such as "It's been a really long time. How have you been lately?" and send it to the device. The user's voice data is given as input, and the response voice data is generated as output.

[1619] 4. The device returns the generated response to the user. The response audio data sent from the server is given as input. The response is played back to the user through the speaker as output.

[1620] An example of a specific prompt is, "Dad, what did you do today?"

[1621] (Application example 1)

[1622] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1623] Although existing technologies exist for systems that collect personal data from the deceased and recreate natural interactions with the deceased, they lack specific environments and memory-based interactions to improve the user's emotional satisfaction.This invention aims to provide a virtual memorial service that allows users to feel a deeper emotional connection by interacting with the deceased in a specific memorial store.

[1624] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1625] In this invention, the server includes a means for collecting personal data of the deceased while they are alive, a means for analyzing the collected data to extract personal characteristics, a means for generating a 3D virtual model using the characteristics, a means for providing a virtual space in which the user can interact with the generated 3D virtual model, a means for using interactive AI to make the interaction in the virtual space natural, a means for placing a specific memorial store in the virtual space in which the user can move, and a means for placing objects and music related to the deceased in the memorial store and realizing interaction based on memorable episodes. This allows the user to interact with the deceased in a specific environment, further strengthening the emotional connection.

[1626] "Personal data" refers to all information, including audio, video, and linguistic information, captured during a person's lifetime.

[1627] "Features" are information extracted from collected personal data, such as facial features, voice spectrum, speaking rhythm and tone.

[1628] A "3D virtual model" is a three-dimensional human model generated in a virtual space based on collected features.

[1629] A "virtual space" is a digitally recreated environment that a user can access.

[1630] "Conversational AI" is artificial intelligence that is trained to engage in natural conversations using collected voice and language data.

[1631] A "memorial store" is a specific area within the virtual space where objects, music, and stories related to the deceased are interactively arranged.

[1632] "Object" refers to any object or item placed in a virtual space.

[1633] "Music" refers to melodies and songs related to the deceased, and is sound information played within the virtual space.

[1634] "Memorable episodes" are information that provides conversations based on the deceased's hobbies, favorite things, and past events.

[1635] A "dedicated client application" is a specially designed software application used by a user to access a virtual space.

[1636] System program generation

[1637] To realize this invention, we start by collecting data on individuals during their lifetime. This process uses hardware such as a camera (e.g., OpenCV) and an audio recorder. The collected data is categorized into visual, audio, and language information.

[1638] The server receives the collected data and uses a deep learning model (e.g., TensorFlow) to extract features such as facial features, voice spectrum, speaking rhythm, and tone. Based on this, a 3D virtual model is generated. It also trains a conversational AI model based on the episode data and develops algorithms to achieve natural conversations. This model is then used during subsequent conversations.

[1639] Meanwhile, a specific memorial store will be set up in the virtual space, and will be stocked with objects and music related to the deceased. This setting is important for users to feel a deep emotional connection with the deceased. The memorial store will be rendered using Unity.

[1640] Processing procedure description

[1641] Data collection:

[1642] The server interviews and films the deceased. It uses a camera and microphone to record the deceased's facial expressions, movements, voice, and speaking style. For example, if a user requests, "I want to create a virtual human of my father, so please collect data," the device encodes the collected data and sends it to the server using a secure protocol.

[1643] Data analysis and generative AI model training:

[1644] The server analyzes the received data and uses deep learning models to extract features, which are then used to generate a 3D virtual model that reproduces natural facial expressions and movements. Furthermore, the conversational AI is trained to reproduce the speech style and tone of the deceased.

[1645] Episode data storage:

[1646] To ensure a more natural interaction when users visit the memorial store, the device collects information about the deceased's anecdotes and hobbies and stores it in a database, which the conversational AI can quickly reference.

[1647] Virtual Space Deployment:

[1648] A user logs into the virtual space using a dedicated client application, and after the server verifies the user's authentication information, a 3D virtual model of the deceased person appears in the virtual space. For example, a user can log into a "Memorial Cafe" and enjoy interacting with the 3D model of the deceased. The cafe is furnished with objects and music based on the preferences of the deceased.

[1649] Examples of concrete examples and prompts

[1650] As a concrete example, if a user logs in to Memorial Cafe and says, "Mom, you liked this cafe, didn't you?", the conversational AI will respond, "Yes, I really liked it. I love the atmosphere."

[1651] Example prompt sentence:

[1652] User: Mom, do you like this cafe?

[1653] Deceased Model: Yes, I love it. I love the atmosphere here.

[1654] This allows users to experience natural conversation with the deceased within the memorial store, further strengthening their emotional connection.

[1655] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1656] Step 1:

[1657] Collecting personal data before death.

[1658] Input: User requests data collection

[1659] How it works: The device uses a camera and microphone to record the facial expressions, movements, voice, and speaking style of the deceased. For example, based on a user's request to "create a virtual model of my father," the device collects data for approximately seven hours.

[1660] Output: A dataset containing visual, audio, and linguistic information

[1661] Step 2:

[1662] The collected data is sent to a server using a secure protocol.

[1663] Input: Dataset collected in step 1

[1664] Specific operation: The device encodes the data and sends it to the server using a secure protocol (e.g., SSL / TLS).

[1665] Output: The encoded data reaches the server

[1666] Step 3:

[1667] The server parses the data it receives.

[1668] Input: Encoded data

[1669] How it works: The server uses a deep learning model (e.g., TensorFlow) to decode the data and extract features such as facial features, voice spectrum, speaking rhythm and tone.

[1670] Output: A dataset of extracted facial feature points, voice spectrum, speaking rhythm and tone.

[1671] Step 4:

[1672] Generate a 3D virtual model.

[1673] Input: Feature data extracted in step 3

[1674] Specific operation: The server uses Unity to generate a 3D virtual model based on the extracted feature data, reproducing natural facial expressions and movements.

[1675] Output: Generated 3D virtual model

[1676] Step 5:

[1677] Train a conversational AI model.

[1678] Input: Collected speech and language data

[1679] How it works: The server uses collected voice and language data to train the conversational AI so that it can converse naturally, and uses deep learning algorithms to build a model that reproduces the speech style and tone of the deceased.

[1680] Output: A trained conversational AI model

[1681] Step 6:

[1682] Episode data is stored and indexed in a database.

[1683] Input: Information about episodes and hobbies collected from users

[1684] Specific operation: The device organizes information about episodes and hobbies, stores it in a database, and indexes it for quick reference by the conversational AI.

[1685] Output: Indexed episode data

[1686] Step 7:

[1687] A user logs into the virtual space using a dedicated client application.

[1688] Input: User credentials

[1689] How it works: A user logs into a virtual space using a dedicated client application. The server verifies the authentication information and grants the user access.

[1690] Output: The user can access the virtual space.

[1691] Step 8:

[1692] A memorial store will be placed in the virtual space, and a 3D virtual model of the deceased person will appear.

[1693] Input: User's location, 3D virtual model generated in step 4

[1694] Specific operation: The server places a memorial store in the virtual space and makes a 3D virtual model of the deceased person appear based on the user's location information.

[1695] Output: 3D virtual model of the deceased placed in a memorial store

[1696] Step 9:

[1697] The user engages in natural dialogue with conversational AI.

[1698] Input: User voice input, the conversational AI model trained in step 5, and episode data from step 6

[1699] How it works: When a user speaks to a model of the deceased person, the device sends the voice data in real time to the server. The server uses a conversational AI model to generate a response and sends it to the device. The device then returns the response to the user, realizing a natural dialogue.

[1700] Output: A natural dialogue experience between the user and conversational AI

[1701] Specific examples

[1702] For example, if a user logs in to Memorial Cafe and says, "Mom, do you like this cafe?" the conversational AI will respond, "Yes, I love it. I love the atmosphere here."

[1703] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1704] This invention combines a system that collects personal data from a person's lifetime, analyzes the collected data to generate a 3D virtual model, and enables users to experience natural conversations with the deceased using interactive AI within the metaverse space, with an emotion engine that recognizes the user's emotions.

[1705] System program generation and implementation method

[1706] Generate 3D virtual models from data collection

[1707] 1. A user requests the collection of data on a deceased person. The request is recorded and preparations for collection are made.

[1708] 2. The device interviews and photographs the deceased. Using a camera and microphone, it records the deceased's facial expressions, movements, voice, and speaking style for approximately seven hours. Each piece of data is saved with a timestamp.

[1709] 3. The device encodes the collected data and sends it to the server using a secure protocol, along with a checksum to verify the data's integrity and completeness.

[1710] 4. The server analyzes the received data, using deep learning models to extract features such as facial features, voice spectrum, and speaking rhythm and tone.

[1711] 5. The server performs 3D modeling based on the feature data. The extracted features are used to generate a 3D virtual model that reproduces the natural facial expressions and movements of the deceased.

[1712] 6. The server trains the conversational AI. Using the collected voice and language data, the AI ​​model is trained to reproduce the speaking style and tone of the deceased, specifically understanding the deceased's voice patterns and conversational flow.

[1713] 7. The device collects information about episodes, hobbies, etc. For example, during an interview, it might ask, "What is your favorite food?" and record the answer.

[1714] 8. The device sends the episode data to the server. The data is categorized into episodes, hobbies, favorite things, etc. and organized according to a format.

[1715] 9. The server stores the episode data in a database, where it is indexed and made available for quick reference by the conversational AI.

[1716] Expanding into the Metaverse

[1717] 1. A user logs into the metaverse space using a dedicated client application. The login information is sent to the server via authentication methods.

[1718] 2. The server verifies the user's credentials and grants access. If authentication is successful, access is established within the metaverse space.

[1719] 3. The server creates a 3D virtual model of the deceased person in the Metaverse space based on the user's location information, placing it in an appropriate position taking into account the user's line of sight and location.

[1720] 4. The user talks to the deceased, for example, "Dad, how are you?"

[1721] 5. The device transmits the user's voice in real time to the server, where it is recorded and stored for analysis.

[1722] 6. The server uses conversational AI to analyze the voice data and generate an appropriate response, such as "I've been doing well. How have you been lately?"

[1723] 7. The server sends the generated response to the terminal, which then outputs the response to the user.

[1724] Incorporating an emotion engine

[1725] 1. The device captures the user's facial expressions and tone of voice in real time using a camera and microphone.

[1726] 2. The device sends the captured data to the emotion engine, which then recognizes the user's emotion based on this data. For example, if a smile is detected, it will recognize that the user is feeling "joy."

[1727] 3. The server receives the emotion data sent from the emotion engine and reflects it in the conversational AI. For example, if the user is recognized as sad, the conversational AI generates a response such as "Cheer up."

[1728] Specific examples

[1729] For example, a user logs into the metaverse space and says, "Dad, it's been a while."

[1730] The device transmits this audio data to the server in real time.

[1731] The emotion engine analyzes and recognizes that the user's tone of voice is sad.

[1732] The server uses conversational AI to analyze the voice data and, referring to the analysis results of the emotion engine, generates an appropriate response such as, "It's been a while. What's wrong? You seem down."

[1733] The server sends the generated response to the terminal, which then speaks the response to the user.

[1734] This allows users to have more emotionally enriching interactions with deceased loved ones within the metaverse space.

[1735] The processing flow will be explained below.

[1736] Step 1:

[1737] The user requests the collection of data on the deceased. The request is confirmed and preparations for collection are made.

[1738] Step 2:

[1739] The device interviews and photographs the deceased. Using a camera and microphone, it records the deceased's facial expressions, movements, voice, and speaking style for approximately seven hours. The collected data is then saved with a timestamp.

[1740] Step 3:

[1741] The device encodes the collected data and transmits it to the server using a secure protocol, along with a checksum to ensure data integrity and completeness.

[1742] Step 4:

[1743] The server analyzes the received data and uses deep learning models to extract features such as facial features, voice spectrum, speaking rhythm and tone.

[1744] Step 5:

[1745] The server performs 3D modeling based on the feature data, and generates a 3D virtual model using the extracted features to reproduce the natural facial expressions and movements of the deceased.

[1746] Step 6:

[1747] The server trains the conversational AI. Using the collected voice and language data, the AI ​​model is trained to reproduce the speaking style and tone of the deceased. Specifically, it analyzes and learns from the deceased's voice patterns and conversational flow.

[1748] Step 7:

[1749] The device collects information such as episodes and hobbies. For example, it asks, "What is your favorite food?" and records the answer.

[1750] Step 8:

[1751] The device sends episode data to the server, which organizes the data into categories such as episodes, hobbies, and favorite things.

[1752] Step 9:

[1753] The server stores the episode data in a database, where it is indexed and prepared for fast and efficient reference by the conversational AI.

[1754] Step 10:

[1755] A user logs into the metaverse space using a dedicated client application, and the login information is sent to the server via authentication means.

[1756] Step 11:

[1757] The server verifies the user's credentials and grants access. If authentication is successful, access is established within the metaverse space.

[1758] Step 12:

[1759] The server creates a 3D virtual model of the deceased person in the Metaverse space based on the user's location information, and adjusts the position of the virtual model based on the user's line of sight and location.

[1760] Step 13:

[1761] The user talks to the deceased, for example, saying, "Dad, it's been a long time."

[1762] Step 14:

[1763] The device transmits the user's voice in real time to a server, where the voice data is recorded and stored for analysis.

[1764] Step 15:

[1765] The server uses conversational AI to analyze the voice data and generate an appropriate response, such as "It's been a while, how have you been?"

[1766] Step 16:

[1767] The device captures the user's facial expressions and tone of voice in real time using a camera and microphone.

[1768] Step 17:

[1769] The device captures facial and voice data and sends it to the emotion engine, which analyzes this data and recognizes the user's emotions.

[1770] Step 18:

[1771] The server receives the emotion data sent from the emotion engine and reflects it in the conversational AI. For example, if the user looks sad, the conversational AI will generate a response such as "Cheer up, what happened recently?"

[1772] Step 19:

[1773] The server sends the generated response to the terminal, which then provides an audio output to the user.

[1774] This allows users to interact with deceased loved ones in a more emotionally rich and natural way within the metaverse space.

[1775] Example 2

[1776] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1777] Conventional systems have had difficulty using data from deceased people to recreate conversations with them in a virtual space. Furthermore, they lacked the ability to generate responses based on the user's emotions, making the conversations less emotionally rich and natural.

[1778] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for collecting personal data before death, means for analyzing the collected data to extract personal characteristics, means for generating a 3D virtual model using the characteristics, means for providing a virtual space in which the user can interact with the generated 3D virtual model, means for using interactive artificial intelligence to make the interaction in the virtual space natural, and means for recognizing the user's emotions and flexibly adjusting the response content. This allows the user to experience an emotionally rich and natural interaction with the deceased person in the virtual space.

[1779] "Personal data" is information about a person that is collected during their lifetime, including photographs, videos, audio files, and written material.

[1780] "Feature extraction means" refers to a process or technique for identifying and analyzing individual characteristics or features from collected data.

[1781] A "3D virtual model" is a three-dimensional digital representation of an individual's appearance, gestures, movements, etc., based on collected data and extracted features.

[1782] "Virtual space" refers to a virtual environment that users can access through a computer or dedicated client application, including the metaverse and VR space.

[1783] "Conversational artificial intelligence" refers to a computer program or system that can use natural language analysis and generation techniques to engage in natural conversations with humans.

[1784] "Means for recognizing emotions and flexibly adjusting response content" refers to a function that analyzes the user's facial expressions, tone of voice, and other emotional indicators in real time and dynamically changes the response content based on the results.

[1785] "Episode data" refers to information based on specific events, hobbies, or specific memories related to the deceased.

[1786] "Dedicated client application" refers to software specifically designed to access virtual spaces and interface with interactive AI and 3D virtual models.

[1787] MODE FOR CARRYING OUT THE INVENTION

[1788] This invention is a system that collects personal data from a person's life and analyzes it to generate a 3D virtual model. This system uses conversational artificial intelligence to allow users to experience natural interactions with the deceased in the metaverse space. It also combines an emotion engine that recognizes the user's emotions to realize more emotionally rich interactions.

[1789] Hardware and software used

[1790] Hardware:

[1791] Camera: Used to record the facial expressions and movements of the deceased.

[1792] Microphone: Used to record the voice and speaking style of the deceased.

[1793] Devices (e.g., computers, smartphones): used for data collection and transmission.

[1794] Server: Used for data processing, analysis, storage, AI training, and virtual space management.

[1795] software:

[1796] Dedicated client application: Software that allows users to access the virtual space.

[1797] Data collection software: collects data from the camera and microphone, encodes it, and sends it to a server.

[1798] Deep learning models: Algorithms for data analysis and feature extraction.

[1799] 3D modeling tools (e.g. Blender): Generate 3D virtual models based on collected feature data.

[1800] Conversational Artificial Intelligence (AI): Enables natural dialogue with users.

[1801] Emotion engine: Analyzes user emotions and reflects them in the dialogue AI.

[1802] Examples of concrete examples and prompts

[1803] For example, suppose a user logs into the metaverse space using a dedicated client application and says, "Dad, it's been a while." The specific processing at this time is as follows.

[1804] 1. The device sends this voice data to the server in real time.

[1805] 2. The server uses conversational artificial intelligence to analyze the voice data and generates an appropriate response by referring to the analysis results of the emotion engine.

[1806] 3. For example, the emotion engine analyzes and recognizes that the user's tone of voice is sad, and the server generates a response such as, "It's been a while, what's wrong? You seem depressed."

[1807] 4. The server sends the generated response to the terminal, which then speaks the response to the user.

[1808] In this way, users can have emotionally enriching interactions with deceased loved ones within the metaverse space.

[1809] Example prompt sentence:

[1810] "Dad, how have you been?"

[1811] "How are you feeling today?"

[1812] "I want to talk about recent events."

[1813] By inputting these prompt sentences, the conversational AI can understand the user's intentions and emotions and generate the most appropriate response.

[1814] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1815] Step 1: (Data collection request)

[1816] The user requests the collection of data on the deceased, specifying specific items to be collected, such as photos, videos, and audio files of the deceased.

[1817] The device records the request and begins preparations for collection, preparing the camera and microphone, and confirming the data format to be collected.

[1818] Input: User request.

[1819] Output: Collection readiness status.

[1820] Step 2: (Interview and filming)

[1821] The device uses a camera and microphone to interview and film the deceased, and the data collected includes facial expressions, movements, voice, and speaking style.

[1822] The device stores all data with a timestamp, splitting it into video clips every 10 minutes, for example.

[1823] Input: Real-time data from camera and microphone.

[1824] Output: Collected data with timestamp.

[1825] Step 3: (Data encoding and transmission)

[1826] The device encodes the collected data and transmits it to a server using a secure protocol (e.g., HTTPS).

[1827] The terminal generates a checksum to verify the integrity and completeness of the data and also sends this to the server.

[1828] Input: Collected data before encoding.

[1829] Output: The encoded data and a checksum.

[1830] Step 4: (Data Analysis)

[1831] The server then analyzes the received data using a deep learning model that extracts features such as facial features, voice spectrum, and speaking rhythm and tone.

[1832] The server specifically analyzes the deceased person's unique movements and expressions and stores them in a database.

[1833] Input: The encoded data.

[1834] Output: Extracted feature data.

[1835] Step 5: (3D modeling)

[1836] The server generates a 3D virtual model based on the extracted feature data, carefully adjusting it to reproduce the natural facial expressions and movements of the deceased.

[1837] The server uses a 3D modeling tool (e.g. Blender) to see in real time what needs to be fixed.

[1838] Input: Extracted feature data.

[1839] Output: 3D virtual model.

[1840] Step 6: (Training the conversational AI)

[1841] The server uses the collected speech and language data to train the conversational AI, specifically learning patterns to replicate the speech and narration of the deceased.

[1842] The server mimics voice patterns and adjusts them to create a natural conversational flow.

[1843] Input: Audio and language data.

[1844] Output: A trained conversational AI model.

[1845] Step 7: (Collect episode information)

[1846] Users provide information about the deceased, such as anecdotes, hobbies, and interests, in the form of an interview.

[1847] The device records this information and also saves it as text data.

[1848] Input: The answer given by the user.

[1849] Output: Episode data.

[1850] Step 8: (Submit episode data)

[1851] The episode data collected by the device is organized according to a format and sent to the server.

[1852] The device classifies the data (episodes, hobbies, interests) before sending it.

[1853] Input: Episode data.

[1854] Output: Organized episode data.

[1855] Step 9: (Storing episode data)

[1856] The server indexes and stores the received episode data in a database.

[1857] The server applies a search algorithm to allow quick access to the data.

[1858] Input: Organized episode data.

[1859] Output: Indexed episode data in a database.

[1860] Step 10: (Login and Authentication)

[1861] A user logs into the virtual space using a dedicated client application, and the login information is sent from the client to the server.

[1862] The server verifies the user's credentials and grants access.

[1863] Input: User login information.

[1864] Output: Authentication check and access granted.

[1865] Step 11: (Placing the virtual model)

[1866] The server creates a 3D virtual model of the deceased person in a virtual space based on the user's location information.

[1867] The server places the virtual model in an appropriate position, taking into account the user's line of sight and position.

[1868] Input: User's location.

[1869] Output: Positioning information of the virtual model in the virtual space.

[1870] Step 12: (User interaction)

[1871] The user talks to the deceased, for example, asking, "Dad, how are you?"

[1872] The device transmits this audio data to the server in real time.

[1873] The server uses conversational artificial intelligence to analyze the voice data and generates an appropriate response by referring to the analysis results of the emotion engine.

[1874] Input: User's question, voice data.

[1875] Output: A response from the conversational AI.

[1876] Step 13: (Capturing the user's facial expressions and voice)

[1877] The device captures the user's facial expressions and tone of voice in real time using a camera and microphone.

[1878] The device transmits the captured data to the emotion engine in real time.

[1879] Input: User's facial expression data and tone of voice.

[1880] Output: Emotion data sent to the emotion engine.

[1881] Step 14: (Analyze emotion data)

[1882] The server analyzes the user's emotions based on the data sent from the emotion engine. For example, it recognizes "joy" by detecting a smile.

[1883] Input: User emotion data from the emotion engine.

[1884] Output: Analyzed user emotion information.

[1885] Step 15: (Generating and reflecting dialogue responses)

[1886] The server generates a response from the conversational AI based on the emotional data. For example, if the server recognizes that the user is sad, it generates a response such as "Cheer up."

[1887] The server sends the generated response to the terminal, which then outputs it to the user.

[1888] Input: Parsed user emotion information.

[1889] Output: AI response based on emotion.

[1890] (Application example 2)

[1891] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1892] In modern society, people often seek emotional connections through conversations with deceased loved ones, but systems that can provide such experiences are not yet widespread. Furthermore, there are no systems incorporating conversational AI that can recognize and reflect emotions in real time, making it difficult to understand the user's emotions and provide natural conversations accordingly. Furthermore, for users to naturally converse with deceased loved ones in virtual spaces, it is necessary to manage and analyze episode data and emotional data, and an efficient system for this purpose is needed.

[1893] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1894] In this invention, the server includes means for collecting personal data before death, means for analyzing the collected data to extract personal characteristics, means for generating a 3D virtual model using the characteristics, means for providing a virtual space in which a user can interact with the generated 3D virtual model, means for using interactive AI to make the interaction in the virtual space natural, means for capturing and analyzing the user's voice and emotions in real time, and means for generating a response that reflects the analyzed emotional data, thereby enabling a user to have an emotional experience through interaction with the deceased.

[1895] "Personal data" refers to information provided by an individual during their lifetime, such as their voice, facial expressions, movements, and speaking style.

[1896] "Features" refers to identifying information such as facial features, voice spectrum, speaking rhythm and tone extracted from collected personal data.

[1897] "3D Virtual Model" means a three-dimensional model of an individual displayed in a virtual space, generated based on collected characteristics.

[1898] "Virtual space" refers to a virtual reality environment accessible to a user.

[1899] "Conversational AI" refers to an artificial intelligence system that uses natural language processing to engage in dialogue with users.

[1900] "Means for capturing and analyzing voice and emotions in real time" refers to the function of collecting the user's voice and facial expressions in real time using a voice input device or camera, and analyzing them using an emotion analysis engine.

[1901] "Emotion data" refers to information about emotions analyzed from the user's tone of voice, facial expressions, etc.

[1902] "Means for generating a response" refers to an appropriate response that corresponds to the user's emotions, which is generated by the conversational AI based on the analyzed emotional data.

[1903] This invention is a system that allows a user to have natural conversations with a deceased person in a virtual space based on personal data collected by the user while he or she was alive. An embodiment of this system will be described in detail below.

[1904] First, the user uses a device to collect data on the deceased. Data such as the deceased's facial expressions, movements, voice, and speaking style is recorded using a camera and microphone for approximately seven hours. This collected data is then sent to a server using a secure protocol. The server then uses a deep learning model to analyze the data and extract individual features. Based on the extracted features, a realistic and natural-looking 3D virtual model is then generated.

[1905] The generated 3D virtual model is displayed in a virtual space accessible to users. Users access this virtual space using a dedicated client application on their smartphone or head-mounted display. When the user speaks to the deceased, their voice is transmitted in real time to a server, where conversational AI analyzes the voice and generates an appropriate response.

[1906] The system also incorporates an emotion engine that captures and analyzes the user's voice and facial expressions in real time. The emotion engine analyzes emotions from the user's tone of voice and facial expressions and reflects them in the responses generated by the conversational AI. This allows the user to experience a more natural and emotional dialogue. For example, if a user says, "Dad, it's been a while," the emotion engine will analyze that the user's tone of voice sounds sad, and the conversational AI will generate a response such as, "It's been a while. What's wrong? You seem down."

[1907] The hardware used includes a smartphone, a head-mounted display, a camera, and a microphone, while the software used includes Unity (a 3D game engine), TensorFlow (a deep learning framework for sentiment analysis), Google Cloud Speech-to-Text API, and Azure Cognitive Services (conversational AI).

[1908] An example prompt is:

[1909] "User prompt"

[1910] We would like to request the development of an application that provides an emotional experience through interaction with the deceased, such as:

[1911] The user collects data about the deceased and conducts an interview using a camera and microphone. The collected data is sent to a server, and a 3D virtual model is generated using deep learning. The user logs in to the virtual space and begins a dialogue. The emotion engine analyzes the user's emotions in real time, and the conversational AI generates a response based on the user's emotions. Please help us develop an application that allows users to have an emotionally rich experience through natural dialogue.

[1912] In this way, a system is realized that allows users to naturally engage in emotional dialogue with the deceased in a virtual space.

[1913] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1914] Step 1:

[1915] The device collects data about the deceased.

[1916] Specifically, the user uses a device to record the deceased's facial expressions, movements, voice, and speaking style using a camera and microphone for approximately seven hours. This data is time-stamped, and the audio data is saved in WAV format, and the video data in MP4 format. The input is the audio and video data of the deceased, and the output is the saved multimedia data.

[1917] Step 2:

[1918] The terminal transmits the collected data to the server.

[1919] The collected data is sent to the server using a secure protocol (e.g. HTTPS), along with a checksum to verify the integrity and completeness of the data. The input is the stored multimedia data, and the output is the data received by the server.

[1920] Step 3:

[1921] The server analyzes the data and extracts personal characteristics.

[1922] A deep learning model (for example, a model using TensorFlow) is used to analyze facial feature points, voice spectrum, speaking rhythm and tone, etc. The analysis results are saved as feature data. The input is the multimedia data received by the server, and the output is the extracted feature data.

[1923] Step 4:

[1924] The server generates a 3D virtual model based on the feature data.

[1925] Using the extracted features, 3D modeling software (e.g., Unity) is used to generate a 3D virtual model that reproduces the deceased's natural facial expressions and movements. The input is the feature data, and the output is the 3D virtual model.

[1926] Step 5:

[1927] Users access the virtual space using a dedicated client application.

[1928] Users log in to the Metaverse space using a smartphone or head-mounted display. The login information is sent to the server via authentication means. The input is the user's authentication information, and the output is the access permission granted upon successful authentication.

[1929] Step 6:

[1930] The server makes a 3D virtual model of the deceased appear in the virtual space based on the user's location information.

[1931] The 3D virtual model is placed in an appropriate position taking into account the user's line of sight and position information. The input is the user's position information, and the output is the position of the 3D virtual model in the metaverse space.

[1932] Step 7:

[1933] The terminal transmits the user's voice to the server in real time.

[1934] When a user speaks to the deceased, the audio data is captured in real time and sent to a server, where it is converted to text using the Google Cloud Speech-to-Text API. The input is the user's audio data, and the output is the converted audio data.

[1935] Step 8:

[1936] The server uses conversational AI to analyze the voice data and generate an appropriate response.

[1937] The server uses conversational AI to generate appropriate responses based on user input. It uses Azure Cognitive Services. The input is the converted voice data, and the output is the generated response text.

[1938] Step 9:

[1939] The terminal then vocalizes the generated response to the user.

[1940] The response text from the server is converted into audio format and played back to the user in real time. The input is the generated response text, and the output is the audio output.

[1941] Step 10:

[1942] The device captures the user's facial expressions and tone of voice in real time and sends them to the emotion engine.

[1943] It uses a camera and microphone to capture the user's facial expressions and tone of voice and sends them to the emotion engine, which then analyzes them using a deep learning model. The input is the user's facial and voice data, and the output is emotion data.

[1944] Step 11:

[1945] The server reflects the emotional data received from the emotion engine in the conversational AI.

[1946] The conversational AI adjusts the response it generates depending on the user's emotions. For example, if the user is recognized as sad, the response may be "Cheer up." The input is emotional data, and the output is the adjusted response text.

[1947] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1948] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1949] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1950] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1951] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1952] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1953] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1954] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1955] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1956] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1957] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1958] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1959] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1960] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1961] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1962] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1963] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1964] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1965] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1966] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1967] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1968] The following is further disclosed regarding the above embodiment.

[1969] (Claim 1)

[1970] means of collecting personal data before death;

[1971] means for analyzing the collected data to extract personal characteristics;

[1972] means for generating a 3D virtual model using said features;

[1973] A means for providing a metaverse space in which users can interact with the generated 3D virtual model;

[1974] A means for using conversational AI to make interactions within the metaverse space natural;

[1975] A system including:

[1976] (Claim 2)

[1977] The system according to claim 1, further comprising means for storing collected episode data in a database, and for the interactive AI to refer to episodes from the database and generate responses.

[1978] (Claim 3)

[1979] 10. The system of claim 1, further comprising means for a user to access the metaverse space using a dedicated client application.

[1980] "Example 1"

[1981] (Claim 1)

[1982] a means by which a user may request the collection of data about a deceased person;

[1983] A means of recording the deceased's facial expressions, movements, voice, and manner of speaking using a camera and microphone;

[1984] means for encoding the recorded data and transmitting it to a server using a secure protocol;

[1985] A means for analyzing the data received by the server using a deep learning model and extracting features;

[1986] means for generating a 3D virtual model based on the extracted features;

[1987] a means for training a conversational AI using speech and language data;

[1988] A means of collecting information such as episodes and hobbies and storing it in a database,

[1989] A means for users to access the metaverse space using a dedicated client application;

[1990] A means for the server to create a 3D virtual model of the deceased person in the metaverse space and adjust it based on the user's location information;

[1991] A means of realizing natural dialogue using conversational AI,

[1992] A system including:

[1993] (Claim 2)

[1994] 10. The system of claim 1, further comprising means for organizing the collected episode data, transmitting the data to a server via a secure protocol, and storing the data in a database.

[1995] (Claim 3)

[1996] The system of claim 1, further comprising a means for a user to engage in natural dialogue with a deceased person using conversational AI within the metaverse space.

[1997] "Application Example 1"

[1998] (Claim 1)

[1999] means of collecting personal data before death;

[2000] means for analyzing the collected data to extract personal characteristics;

[2001] means for generating a 3D virtual model using said features;

[2002] a means for providing a virtual space in which a user can interact with the generated 3D virtual model;

[2003] A means for using an interactive AI to make the dialogue in the virtual space natural;

[2004] A means for arranging a specific commemorative store in a virtual space that a user can move around in;

[2005] A means for placing objects and music related to the deceased in the memorial store and realizing dialogue based on memorable episodes;

[2006] A system including:

[2007] (Claim 2)

[2008] The system according to claim 1, further comprising means for storing collected episode data in a database, and for the interactive AI to refer to episodes from the database and generate responses.

[2009] (Claim 3)

[2010] 10. The system of claim 1, further comprising means for a user to access the virtual space using a dedicated client application.

[2011] "Example 2: Combining Emotion Engines"

[2012] (Claim 1)

[2013] means of collecting personal data before death;

[2014] means for analyzing the collected data to extract personal characteristics;

[2015] means for generating a 3D virtual model using said features;

[2016] a means for providing a virtual space in which a user can interact with the generated 3D virtual model;

[2017] means for using interactive artificial intelligence to make the interaction in the virtual space natural;

[2018] A means for recognizing a user's emotions and flexibly adjusting the response content;

[2019] A system including:

[2020] (Claim 2)

[2021] 2. The system according to claim 1, further comprising means for storing collected episode data in a database, and for the interactive artificial intelligence to refer to episodes from the database and generate a response.

[2022] (Claim 3)

[2023] 10. The system of claim 1, further comprising means for a user to access the virtual space using a dedicated client application.

[2024] "Application example 2 when combining emotion engines"

[2025] (Claim 1)

[2026] means of collecting personal data before death;

[2027] means for analyzing the collected data to extract personal characteristics;

[2028] means for generating a 3D virtual model using said features;

[2029] A means for providing a virtual space in which a user can interact with the generated 3D virtual model;

[2030] A means for using an interactive AI to make the dialogue in the virtual space natural;

[2031] A means of capturing and analyzing the user's voice and emotions in real time;

[2032] means for generating a response that reflects the analyzed emotion data;

[2033] A system including:

[2034] (Claim 2)

[2035] The system according to claim 1, further comprising means for storing collected episode data in a database, and for the interactive AI to refer to episodes from the database and generate responses.

[2036] (Claim 3)

[2037] 10. The system of claim 1, further comprising means for a user to access the virtual space using a dedicated client application. [Explanation of symbols]

[2038] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. means of collecting personal data before death; means for analyzing the collected data to extract personal characteristics; means for generating a 3D virtual model using said features; A means for providing a metaverse space in which users can interact with the generated 3D virtual model; A means for using conversational AI to make interactions within the metaverse space natural; A system including:

2. The system according to claim 1, further comprising means for storing collected episode data in a database, and for the interactive AI to generate a response by referring to an episode from the database.

3. The system of claim 1 , further comprising means for a user to access the metaverse space using a dedicated client application.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A