system

The system generates a three-dimensional model and reproduces the voice of a deceased person for a realistic virtual experience, addressing the lack of emotional satisfaction in existing technologies by enabling natural and personalized conversations.

JP2026068317APending Publication Date: 2026-04-22SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-10
Publication Date
2026-04-22

AI Technical Summary

Technical Problem

Existing technologies struggle to provide a sufficiently realistic and emotionally satisfying virtual experience for interacting with deceased loved ones, particularly in reproducing conversations based on their characteristic appearance, voice, and individual memories.

Method used

A system that analyzes voice and video of a deceased person to generate a three-dimensional model and reproduce their voice, enabling dialogue in a virtual space using a database to reference relevant information and apply machine learning for natural interactions.

Benefits of technology

Provides a realistic and emotionally satisfying virtual experience by recreating the deceased's appearance and voice, allowing users to have personalized and natural conversations, thereby strengthening emotional connections.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026068317000001_ABST
    Figure 2026068317000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] A means for analyzing voice and video information of a person acquired by the user and generating a three-dimensional model of the said person, A means for reproducing the voice of the person based on the aforementioned audio information, A means to enable dialogue with the person in a virtual space using a generated three-dimensional model and reproduced voice, A system that includes this.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] For people who wish to interact with or reunite with their deceased loved ones, it is difficult for existing technologies to provide a sufficiently realistic and emotionally satisfying virtual experience. In particular, in reproducing conversations based on the characteristic appearance, voice, individual memories, and episodes of the deceased, there are many deficiencies in current systems. An object of the present invention is to provide means for providing an interactive experience with a deceased loved one in a more natural and individualized virtual space.

Means for Solving the Problems

[0005] This invention solves the problem by providing a means for analyzing the voice and video of a deceased person acquired by a user and generating a three-dimensional model of the deceased. It also includes means for reproducing the voice of the deceased person based on the voice information, and means for enabling dialogue with the deceased person in a virtual space using the generated three-dimensional model and the reproduced voice. Furthermore, by combining means for generating conversations by referencing relevant information from a database based on user input in the virtual space, natural interactions based on individual memories and episodes are realized. Moreover, machine learning technology is applied to the core of these means to enable the reproduction of more realistic appearances and voices, thereby providing users with an emotionally satisfying experience.

[0006] "Audio information" refers to audio data acquired by the user, which records characteristics such as the way a specific person speaks and their voice quality.

[0007] "Video information" refers to video data acquired by the user, which visually records the appearance and actions of a specific person.

[0008] A "three-dimensional model" is a digital model that reproduces the appearance of a specific person in three dimensions, reflecting that person's characteristics.

[0009] "Reproduced voice" refers to voice synthesized based on acquired audio information, reproducing the vocal characteristics of a specific person.

[0010] A "virtual space" is a digital space constructed using computer technology, within which users can have visual and auditory experiences.

[0011] "Means that enable dialogue" refers to systems and methods for users and three-dimensional models to communicate with each other in a virtual space.

[0012] A "database" is a system for efficiently and systematically organizing and storing information, enabling specific searches and references.

[0013] "Machine learning technology" is a technique that uses algorithms and methods to analyze large amounts of data and automatically learn patterns and features to make predictions and decisions. [Brief explanation of the drawing]

[0014] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] This is a sequence diagram showing the processing flow of the data processing system in Example 2, which incorporates an emotion engine. [Figure 14] It is a sequence diagram showing the processing flow of a data processing system in Application Example 2 when a sentiment engine is combined.

Embodiments for Carrying Out the Invention

[0015] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.

[0016] First, the terms used in the following description will be explained.

[0017] In the following embodiments, a numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.

[0018] In the following embodiments, a numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0019] In the following embodiments, a numbered storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, etc.

[0020] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0021] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0022] [First Embodiment]

[0023] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0024] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0025] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0026] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0027] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0028] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0029] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0030] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0031] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0032] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0033] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0034] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0035] This invention is a system that generates a three-dimensional model using the voice and video information of a deceased person, enabling interaction with the deceased person in a virtual space. This system operates as follows:

[0036] First, the user uses their device to collect audio and video information about the deceased. The collected information is securely uploaded from the device to a server. The server analyzes this information and extracts the deceased's physical and vocal characteristics. The extracted characteristics are used to generate a three-dimensional model and reconstruct the voice.

[0037] The server uses a facial recognition algorithm to identify the deceased's distinctive appearance from video information and generates a three-dimensional model based on this. For audio information, a speech synthesis engine is used to analyze it and generate audio data to reproduce the deceased's speaking style and voice quality.

[0038] The generated 3D model and audio data are converted into a format usable in a virtual space. This virtual space is accessible to users via their devices and serves as a place to interact with the deceased person's avatar. When a user logs into the virtual space, the server displays the deceased person's 3D avatar and provides a conversation using reproduced audio.

[0039] For example, if a user asks the deceased to "tell me about a past trip," the device sends this request to the server. The server searches for relevant information in its database and generates a response based on shared memories with the deceased. For instance, it might provide a response that reflects the deceased's personality, such as, "That trip to the beach was so much fun. I remember you collecting seashells on the sand." The generated response is then reproduced in the virtual space using the deceased's voice.

[0040] This invention thus provides a realistic and personalized virtual experience, strengthening the emotional connection with the deceased. Users can find emotional healing through reuniting with the deceased and recreating memories in the virtual space.

[0041] The following describes the processing flow.

[0042] Step 1:

[0043] The user uses the device to record the deceased's voice and video. The device uses a microphone and camera to acquire high-quality audio and video data, which is then stored in temporary data storage.

[0044] Step 2:

[0045] The device uploads the collected audio and video data to the server. The upload uses an encrypted, secure communication protocol to protect the confidentiality of the data.

[0046] Step 3:

[0047] The server analyzes the received audio data and extracts speech features such as the deceased person's speaking style, pitch, and intonation. This process uses a speech recognition algorithm.

[0048] Step 4:

[0049] The server analyzes the video data to identify the deceased's appearance and distinctive facial patterns. Using facial recognition technology, it extracts feature points of the face in three-dimensional space.

[0050] Step 5:

[0051] The server uses the extracted speech features to recreate the deceased person's voice using a speech synthesis engine. The synthesized voice is then converted into a format that can be played back as natural-sounding speech.

[0052] Step 6:

[0053] The server uses feature points extracted from the video footage and 3D modeling software to generate a 3D avatar of the deceased. This avatar is a digital model that preserves the visual characteristics of the deceased, including their appearance.

[0054] Step 7:

[0055] The user accesses the virtual space using a terminal and begins a conversation with the deceased based on a 3D avatar and voice data provided by the server.

[0056] Step 8:

[0057] The server receives input from the user in the virtual space, searches the database, and generates natural conversations based on relevant episodes and information about the deceased.

[0058] Step 9:

[0059] The generated conversation content is provided to the user in a virtual space using reconstructed audio. The user can continue the conversation with the deceased and have an emotional experience.

[0060] (Example 1)

[0061] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0062] In modern society, there is a growing need to maintain emotional connections with deceased loved ones by recreating conversations and memories in virtual spaces. However, conventional technologies struggle to reproduce the realistic voice and appearance of the deceased, resulting in an unrealistic experience.

[0063] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0064] In this invention, the server includes means for analyzing audio and image information acquired by the user and extracting the characteristics of a person from the audio and image information; means for generating a three-dimensional representation of the person based on the extracted characteristics; and means for generating acoustic information based on the characteristics and reproducing the voice of the person. This makes it possible to have a realistic and personalized conversation with a deceased person in a virtual space.

[0065] "Audio and visual information" refers to information that includes a person's voice data and visual data, and is fundamental information for understanding biological characteristics.

[0066] "Characteristics" refer to the distinctive features of a person extracted from audio and image information, and are the basic data used to generate three-dimensional displays and acoustic information.

[0067] "Three-dimensional representation" refers to a three-dimensional visual representation of a person generated based on their characteristics, and is used as an avatar in a virtual space.

[0068] "Acoustic information" refers to voice data of a person, reproduced based on their characteristics, and is the voice used in dialogue in a virtual space.

[0069] A "virtual space" is a digital environment created using computer technology, where users access and interact with it through an interface.

[0070] "Conversation information" refers to the content of dialogue generated in a virtual space, specifically the response content generated from a set of information based on the user's input.

[0071] An "information set" refers to a collection of information, including databases and knowledge bases, used to generate dialogue in a virtual space.

[0072] This invention is a system for recreating conversations with deceased persons in a virtual space, and its embodiments are described below.

[0073] First, the user uses their device to collect audio and image information of the deceased. This process involves collecting high-quality data using smart devices or dedicated recording equipment. This data is then encrypted and uploaded to a server.

[0074] The server analyzes the received audio and image information. For audio data, speech synthesis software is used to extract the characteristics of the deceased. This uses common speech recognition services and synthesis engines. For image information, a facial recognition algorithm is used to understand the person's appearance and extract features to generate a 3D representation. At this stage, for example, OpenCV or similar image processing libraries are used.

[0075] Based on the extracted characteristics, the server generates a three-dimensional representation. This process is achieved using a three-dimensional modeling tool such as Blender. Furthermore, based on these characteristics, acoustic information is generated to recreate the deceased's voice. This allows the user to interact with the deceased both visually and aurally in a virtual space.

[0076] The virtual space is a digital environment accessible to users. Users log in to this space via a terminal and enjoy conversations using a 3D display and audio information of the deceased provided by the server. In this process, conversations are generated in real time using a generative AI model, and prompt sentences are processed in response to user input, resulting in natural and personalized conversations.

[0077] For example, if a user uses the prompt "Tell me about a past trip," the server will use this request to search for memories and related information about the deceased and generate a response such as, "That trip to the beach was really fun. I remember you collecting seashells on the sand." This response is reproduced in a voice that mimics the deceased's voice.

[0078] The distinguishing feature of this invention is that, by integrating these technologies into a system, it can provide users with emotional healing by offering a realistic and personalized virtual experience with the deceased and strengthening their emotional connection.

[0079] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0080] Step 1:

[0081] The user uses a device to collect audio and image information. The input used is video of a person speaking or a portrait image. The device temporarily stores the collected data and prepares it for transmission to a server. Specifically, it records high-resolution video using a smartphone or digital camera and records audio using a microphone.

[0082] Step 2:

[0083] The device uploads the collected audio and image information to the server. The input is the data collected in step 1, and the output is the data sent to the server. Here, data is encrypted using a secure communication protocol such as HTTPS to protect privacy. The device notifies the user when the transmission is complete.

[0084] Step 3:

[0085] The server analyzes the received audio and image data. The input is data received from the terminal, and the output is extracted characteristic data of the person. A speech recognition engine is used to analyze the audio data, extracting the pitch and tone of the voice. Image data is processed by a facial recognition algorithm to extract data on distinctive appearances. This process provides basic information to identify the characteristics of the deceased.

[0086] Step 4:

[0087] The server generates a three-dimensional representation using the extracted feature data. The input is the feature data obtained in step 3, and the output is a three-dimensional representation that can be displayed in a virtual space. Specifically, 3D modeling software is used to create an avatar based on the extracted data. In addition, acoustic information is generated, and speech synthesis is performed to reproduce the deceased person's voice.

[0088] Step 5:

[0089] The server integrates the generated 3D display and audio information into the virtual space. The input is the data generated in step 4, and the output is avatar and audio data usable within the virtual space. These are set up by virtual environment creation software and made accessible to the user. Specifically, a digital environment is created using tools such as Unity or Unreal Engine.

[0090] Step 6:

[0091] Users can access a virtual space using a device and converse with an avatar of a deceased person. As input, users send questions and comments to the virtual space using prompt text. Based on this prompt, the server searches a knowledge base for relevant information and generates a response. The output is a voice response presented as a natural conversation to the user. Specifically, using an AI model, a response such as "Tell me about your old trip" might be generated, for example, "That trip to the beach was really fun."

[0092] (Application Example 1)

[0093] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0094] Conventional virtual reality dialogue systems have struggled to foster deep emotional connections with deceased loved ones because they lack sufficient integration with the real world. Furthermore, recreating memorable experiences in places associated with the deceased within the virtual space has been difficult. Therefore, there is a need to provide a more immersive and emotionally impactful virtual experience.

[0095] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0096] In this invention, the server includes means for analyzing voice and video information of a person acquired by the user and generating a three-dimensional model; means for reproducing the person's voice based on the voice information; means for enabling dialogue with the person in a virtual reality space using the generated three-dimensional model and voice; and means for recognizing real-world scenery and providing content linked to the person in the virtual reality space. This makes it possible for the user to enjoy a deeper reminiscing experience by linking the emotional reunion with the deceased with real-world scenery.

[0097] A "user" is a person who uses the system to experience a conversation with the deceased.

[0098] "Audio and video information" refers to recorded voice and video data related to the deceased.

[0099] A "three-dimensional model" is a digital representation created to recreate the appearance of a deceased person in three dimensions.

[0100] A "sound reproduction device" is a device that includes technology for artificially generating the voice of a deceased person based on collected audio information.

[0101] A "virtual reality space" is an interactive artificial environment recreated using digital technology.

[0102] A "device for recognizing real-world landscapes" is a device that uses cameras and sensors to digitize the surrounding physical environment and acquire information.

[0103] An "information storage database" is a digital storage device used to organize and store information about the aforementioned person or related matters.

[0104] This invention is a system for facilitating dialogue with a deceased person in a virtual reality space. This system includes a terminal owned by the user, a server for data processing, and a device capable of displaying the virtual reality environment.

[0105] Users acquire audio and video information about the deceased through their devices and securely upload it to a server in the cloud. The server uses this data to generate a three-dimensional model of the deceased. Facial recognition algorithms are used to generate the 3D model, reproducing the deceased's distinctive appearance. For audio information, speech synthesis technology is used to reproduce the deceased's voice quality and speaking style. Specifically, Amazon Polly and Google® Cloud Text-to-Speech are used for speech synthesis, and AR technology is used for 3D model generation.

[0106] The server recognizes real-world landscape data and provides a responsive experience with the deceased when the user views that landscape through a virtual reality device. When the user visits a specific location, they can share experiences and memories associated with the deceased and that location in the virtual reality space. An example of a prompt message is, "The user has identified a location. Please share your memories associated with that location."

[0107] This system allows users to experience emotional reunions with deceased loved ones in a more realistic way, and to recreate memories with deep emotion. This service, utilizing virtual reality space, goes beyond mere data viewing, offering new possibilities that broaden the scope of experiences.

[0108] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0109] Step 1:

[0110] The user obtains audio and video information of the deceased through the device. The input consists of recorded audio and video data. This data is temporarily stored on the device and used for subsequent processing. Specific actions include the user recording audio or taking photos and videos using a mobile device or computer.

[0111] Step 2:

[0112] The device uploads the acquired audio and video information to a server in the cloud. The input is the audio and video data on the device, and the output is the data securely transferred to the server. This process uses data encryption and data transfer protocols (e.g., HTTPS) to ensure that the information is transmitted securely.

[0113] Step 3:

[0114] The server analyzes uploaded audio and video information to generate a three-dimensional model of the deceased. The input is audio and video data stored on the server, and the output is a three-dimensional model of the deceased. This process utilizes machine learning techniques and facial recognition algorithms to extract the deceased's physical characteristics and reproduce them as a three-dimensional model. The actual operation involves extracting feature points from video data and constructing a three-dimensional digital model based on these points.

[0115] Step 4:

[0116] The server analyzes audio information and performs speech synthesis to recreate the deceased person's voice. The input is audio data, and the output is the recreated voice of the deceased. Amazon Polly and Google Cloud Text-to-Speech are used for speech synthesis technology, which generates digital speech that recreates the deceased person's voice quality and speaking style. The specific operations are analysis of voice characteristics and generation of recreated speech data.

[0117] Step 5:

[0118] The user accesses a virtual space using a virtual reality device and initiates an interaction with a generated three-dimensional model using voice. The input is the user's visual and auditory information, and the output is the experience of interacting with the deceased's avatar. When the user moves to a specific location, the system recognizes the location and generates relevant prompts, providing content that is relevant to the deceased. Specific actions include obtaining the user's location information and initiating a conversation about memories of the deceased based on that information.

[0119] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0120] This invention combines an emotion engine with a system that generates a three-dimensional model by analyzing audio and video information of a deceased person acquired by the user, enabling interaction with the deceased person in a virtual space. This makes it possible to recognize the user's emotions in real time and provide interaction accordingly.

[0121] First, the user uses a device to record audio and video information of the deceased. This information is sent from the device to a server. The server analyzes the characteristics of the deceased's voice from the transmitted audio information and extracts the characteristics of the deceased's appearance from the video information. This generates a three-dimensional model of the deceased and a reconstructed voice.

[0122] Next, the system uses an emotion engine to recognize the user's emotions. When the user accesses the virtual space, the terminal captures the user's facial expressions and voice, and uses the emotion engine to perform real-time emotion analysis. Based on this analysis, the server adjusts the dialogue. For example, if the user has a sad expression, the server generates a response that includes comfort and empathy, and the deceased person's avatar in the virtual space responds appropriately.

[0123] As a concrete example, consider a situation where a user is overwhelmed with emotion and about to cry when reunited with a deceased loved one in a virtual space. At this point, the emotion engine detects the user's tears and sends "sadness" emotion data to the server. Based on this information, the server generates a warm message such as, "Don't cry, we can always meet here," and the avatar of the deceased person in the virtual space speaks that message to the user.

[0124] This invention provides users with an experience that allows them to connect deeply and emotionally with deceased loved ones in a virtual space, enabling flexible conversations that respond to changes in their emotions. As a result, users can experience deeper healing and satisfaction.

[0125] The following describes the processing flow.

[0126] Step 1:

[0127] The user uses the device to record audio and video of the deceased. The recording and video are done in high resolution, and the saved data is stored in the device's temporary storage.

[0128] Step 2:

[0129] The device uploads the collected audio and video data to the server. This process is carried out through an encrypted, secure communication channel to ensure the confidentiality of the data.

[0130] Step 3:

[0131] The server receives the transmitted audio data and applies a speech recognition algorithm to analyze the characteristics of the deceased's voice. From the results of this analysis, the deceased's speaking style and voice quality are digitized.

[0132] Step 4:

[0133] The server processes the video data and uses facial recognition technology to extract the deceased's physical characteristics. This process allows the server to understand the three-dimensional structure in 3D space, which is then used to construct a 3D model.

[0134] Step 5:

[0135] The server generates a three-dimensional model and a reproduced voice of the deceased. These serve as materials for users to experience visually and aurally within the virtual space.

[0136] Step 6:

[0137] When a user logs into the virtual space, the device uses the user's camera and microphone to continuously capture the user's facial expressions and voice, and sends that data to the emotion engine.

[0138] Step 7:

[0139] The emotion engine analyzes received data and evaluates the user's emotional state in real time. The analysis results identify the user's current emotion, such as sadness, happiness, or surprise.

[0140] Step 8:

[0141] The server dynamically generates dialogue content within the virtual space based on emotional data transmitted from the emotion engine. It selects and adjusts appropriate responses according to the user's emotions.

[0142] Step 9:

[0143] Within the virtual space, the conversation generated by the server is reproduced as audio and delivered to the user through the deceased person's avatar. This allows the user to naturally engage in emotional dialogue with the deceased.

[0144] (Example 2)

[0145] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0146] In modern society, there is a growing demand for new forms of interaction with deceased loved ones. However, it is difficult to retain memories of the deceased, and there is a lack of methods to recreate emotional connections. Conventional systems have limitations in the quality of recreating the deceased and the interaction, and in particular, they have the challenge of not being able to have flexible conversations that respond to the user's emotions.

[0147] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0148] In this invention, the server includes means for analyzing voice and video information of a person acquired by the user and generating a three-dimensional model of the person; means for reproducing the person's voice based on the voice information; means for enabling dialogue with the person in a virtual space using the generated three-dimensional model and the reproduced voice; and means for recognizing the user's emotions in real time and generating a corresponding dialogue. This makes it possible to provide flexible and appropriate dialogue that responds to the user's emotions while maintaining a deep emotional connection with the deceased.

[0149] "User" refers to an individual who operates the system and inputs audio and video information.

[0150] "Person" refers to an individual from whom audio and video information is acquired, and who is the subject of a three-dimensional model or a reconstructed voice.

[0151] A "three-dimensional model" refers to a three-dimensional digital representation generated based on a person's physical characteristics.

[0152] "Audio information" refers to digital audio data that includes the characteristics of a person's voice.

[0153] "Reproduced voice" refers to artificial voice data that reproduces a person's voice, generated based on the original audio information.

[0154] A "virtual space" refers to a digital environment constructed using digital technology, where three-dimensional models and reproduced sounds are placed, and users can interact with them.

[0155] "Methods for recognizing emotions in real time" refers to technologies that analyze data such as the user's facial expressions and voice to identify their emotions at that moment.

[0156] "Means of generating dialogue" refers to the process of generating appropriate responses in response to the user's emotions and input.

[0157] A "generative AI model" refers to an algorithm used to learn from data and accomplish a specific task.

[0158] A "prompt sentence" refers to an instruction sentence input into a generative AI model, which serves as a criterion for determining the content of the output dialogue.

[0159] This invention provides a virtual space that enables users to have emotional interactions with deceased loved ones. The user first records audio and video information of the deceased using a device. Suitable hardware includes common smartphones, tablets, and camera devices. This information is transmitted from the device to a server. The server receives this data and analyzes the characteristics of the deceased's voice using an audio feature extraction algorithm. It also extracts the deceased's physical characteristics using deep learning-based image processing software. This generates a three-dimensional model and synthesized voice of the deceased.

[0160] When a user accesses the virtual space, the terminal captures the user's facial expressions and voice, and analyzes the user's emotions in real time through emotion recognition software. This process requires a camera and microphone. Based on this analysis, the server uses a generative AI model to generate appropriate dialogue. An example of a prompt might be, "Generate flexible dialogue that responds to the user's emotions and convey the content in the deceased's voice."

[0161] As a concrete example, suppose a user reunites with a deceased loved one in a virtual space and displays a sad expression. At this point, emotion recognition software identifies the user's expression as "sadness" and sends that data to a server. The server then generates a warm message such as, "Don't cry, I'll always be waiting here for you." This message is conveyed to the user in the virtual space through the deceased person's avatar.

[0162] This system allows users to experience deep emotional connection with their deceased loved ones and enjoy intimate and realistic interactions that respond to their changing emotions.

[0163] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0164] Step 1:

[0165] Users record audio and video information of the deceased using their own devices. Input includes audio and video files obtained from smartphones and cameras. This information is saved in digital format and prepared for transmission to a database.

[0166] Step 2:

[0167] The terminal transmits the recorded audio and video information to the server. In this step, the information is transmitted over the internet as data packets. The input data consists of the audio and video files obtained in the previous step. The output is the encrypted information that arrives securely at the server.

[0168] Step 3:

[0169] The server analyzes the received audio information. It processes the input audio data through an audio feature extraction algorithm, outputting features such as voice intonation, tone, and speaking speed as numerical data. Specifically, an audio signal processing library is used.

[0170] Step 4:

[0171] The server analyzes the received video information. It receives video data as input and uses image processing techniques to extract features such as facial shape and posture. The output is data for a three-dimensional model based on these features. Specifically, it performs image analysis using deep learning.

[0172] Step 5:

[0173] The server generates a 3D model of the deceased and a reconstructed voice. This step uses the voice and video feature data obtained in the previous step. Based on the input data, a generating AI model is used to output a digital avatar and synthesized voice.

[0174] Step 6:

[0175] When a user accesses the virtual space, the terminal captures the user's facial expressions and voice in real time. Input consists of audio and video data obtained from the camera and microphone. Output is a data stream of the user's facial expressions and voice, which is sent to emotion recognition software.

[0176] Step 7:

[0177] The server analyzes the user's emotions in real time. It uses facial and audio data received from the terminal as input. An emotion recognition algorithm analyzes this data and outputs the user's emotional state as numerical data.

[0178] Step 8:

[0179] The server uses a generative AI model to generate dialogue that responds to the user's emotions. The prompt is set to "Generate a warm response based on the user's emotions," and the input is the emotion data obtained in the previous step. The output is the dialogue that the deceased person's avatar should say.

[0180] Step 9:

[0181] The server sends the generated dialogue to the deceased person's avatar in the virtual space, allowing it to interact with the user. The input is the generated dialogue, and the output is the audio and visual dialogue experience provided to the user. The virtual space performs the necessary visual rendering and speech synthesis to reproduce this in real time.

[0182] (Application Example 2)

[0183] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0184] Conventional virtual reality dialogue systems have struggled to provide flexible dialogue that responds to the user's emotional state. There is a growing need to generate appropriate responses based on the user's emotions to foster deeper emotional connections and improve the user experience.

[0185] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0186] In this invention, the server includes means for analyzing voice and video information of a person acquired by the user and generating a three-dimensional model of the person; means for evaluating the user's emotional state and adjusting the dialogue content based on that evaluation; and means for generating a conversation that corresponds to the user's emotions. This enables natural and personalized dialogue that is adapted to emotions.

[0187] A "user" is an individual who operates the system and interacts within the virtual space.

[0188] "Audio information" refers to digital data related to a person's speech, and is fundamental information for a system to reproduce.

[0189] "Visual information" refers to visual data about a person's appearance, which is used to generate a three-dimensional model.

[0190] A "three-dimensional model" is a digital representation of a person in three dimensions, created based on acquired video information.

[0191] "Reproduced speech" refers to the speech expressions of a person that are generated based on audio information.

[0192] A "virtual space" is a digital environment built on a computer, where users interact with each other.

[0193] "Emotional analysis tools" are technologies that evaluate the user's current emotional state and allow the system to determine a response accordingly.

[0194] A "database" is a collection of information that a system uses to store and refer to as needed.

[0195] "Machine learning technology" is a technique that uses algorithms to extract patterns from data and then uses those patterns to make predictions or generate new data.

[0196] The system for implementing this invention is based on a user terminal and a server. First, the user acquires audio and video information using a smart device, such as smart glasses or a head-mounted display. The terminal captures this information and transmits it to the server. The server generates a three-dimensional model from the video information and reconstructs the audio based on the audio information.

[0197] The server incorporates the "Microsoft® Azure® Cognitive Services" sentiment analysis API and the "OpenAI® speech synthesis API." The user's facial expressions and voice data are analyzed in real time using these APIs. Based on the analysis results, the server generates emotionally adaptive dialogue, enabling natural conversations within the virtual space.

[0198] As a concrete example, consider a virtual store scenario. A user reunites with a deceased person, represented as an avatar, in a virtual space, and they tour their favorite places together. At this time, the system detects when the user is smiling, and the deceased person's avatar provides a message such as, "I remember the smile on your face when you chose this." An example of a prompt used here would be, "Generate a warm message when the user is smiling. For example, provide a conversation that recreates a happy moment at a memorable store."

[0199] In this way, by adjusting the interaction in real time according to the user's emotions, the system can provide a deeper emotional connection and create a more satisfying experience for the user.

[0200] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0201] Step 1:

[0202] The terminal captures audio and video information from the user and sends it to the server as digital data. The input is the user's audio and video, and the output is the transmission of this data to the server. This process prepares the basic data necessary for the virtual space.

[0203] Step 2:

[0204] The server generates a three-dimensional model from the received video information. This process utilizes computer graphics technology to create a three-dimensional model of a person. The input is the user's video data, and the output is a three-dimensional model. This generates a virtual avatar based on any given person.

[0205] Step 3:

[0206] The server analyzes voice information and performs speech synthesis based on this analysis. Machine learning techniques are used to capture voice characteristics and reproduce realistic speech. The input is the user's voice data, and the output is the synthesized speech. A generative AI model is used, enabling natural-sounding speech.

[0207] Step 4:

[0208] The server senses the user's facial expressions in real time and evaluates their emotional state using an emotion analysis API. The input is the user's facial expression data, and the output is the result of the emotion analysis. Specifically, the API quantifies the characteristics of each facial expression and determines the emotional state.

[0209] Step 5:

[0210] The server generates prompt sentences based on the sentiment analysis results and constructs an appropriate dialogue. In this step, the generating AI model determines the dialogue content based on the prompt sentences and presents it to the user. The input is data on the emotional state, and the output is a dialogue adapted to that emotion.

[0211] Step 6:

[0212] Within the virtual space, a conversation generated by the user's device is played back. The user can use a smart device to enjoy a conversation with the deceased through a generated 3D model and reproduced voice. The input is the 3D model and voice generated in the previous step, and the output is the conversation experienced by the user.

[0213] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0214] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0215] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0216] [Second Embodiment]

[0217] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0218] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0219] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0220] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0221] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0222] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0223] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0224] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0225] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0226] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0227] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0228] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0229] This invention is a system that generates a three-dimensional model using the voice and video information of a deceased person, enabling interaction with the deceased person in a virtual space. This system operates as follows:

[0230] First, the user uses their device to collect audio and video information about the deceased. The collected information is securely uploaded from the device to a server. The server analyzes this information and extracts the deceased's physical and vocal characteristics. The extracted characteristics are used to generate a three-dimensional model and reconstruct the voice.

[0231] The server uses a facial recognition algorithm to identify the deceased's distinctive appearance from video information and generates a three-dimensional model based on this. For audio information, a speech synthesis engine is used to analyze it and generate audio data to reproduce the deceased's speaking style and voice quality.

[0232] The generated 3D model and audio data are converted into a format usable in a virtual space. This virtual space is accessible to users via their devices and serves as a place to interact with the deceased person's avatar. When a user logs into the virtual space, the server displays the deceased person's 3D avatar and provides a conversation using reproduced audio.

[0233] For example, if a user asks the deceased to "tell me about a past trip," the device sends this request to the server. The server searches for relevant information in its database and generates a response based on shared memories with the deceased. For instance, it might provide a response that reflects the deceased's personality, such as, "That trip to the beach was so much fun. I remember you collecting seashells on the sand." The generated response is then reproduced in the virtual space using the deceased's voice.

[0234] This invention thus provides a realistic and personalized virtual experience, strengthening the emotional connection with the deceased. Users can find emotional healing through reuniting with the deceased and recreating memories in the virtual space.

[0235] The following describes the processing flow.

[0236] Step 1:

[0237] The user uses the device to record the deceased's voice and video. The device uses a microphone and camera to acquire high-quality audio and video data, which is then stored in temporary data storage.

[0238] Step 2:

[0239] The device uploads the collected audio and video data to the server. The upload uses an encrypted, secure communication protocol to protect the confidentiality of the data.

[0240] Step 3:

[0241] The server analyzes the received audio data and extracts speech features such as the deceased person's speaking style, pitch, and intonation. This process uses a speech recognition algorithm.

[0242] Step 4:

[0243] The server analyzes the video data to identify the deceased's appearance and distinctive facial patterns. Using facial recognition technology, it extracts feature points of the face in three-dimensional space.

[0244] Step 5:

[0245] The server uses the extracted speech features to recreate the deceased person's voice using a speech synthesis engine. The synthesized voice is then converted into a format that can be played back as natural-sounding speech.

[0246] Step 6:

[0247] The server uses feature points extracted from the video footage and 3D modeling software to generate a 3D avatar of the deceased. This avatar is a digital model that preserves the visual characteristics of the deceased, including their appearance.

[0248] Step 7:

[0249] The user accesses the virtual space using a terminal and begins a conversation with the deceased based on a 3D avatar and voice data provided by the server.

[0250] Step 8:

[0251] The server receives input from the user in the virtual space, searches the database, and generates natural conversations based on relevant episodes and information about the deceased.

[0252] Step 9:

[0253] The generated conversation content is provided to the user in a virtual space using reconstructed audio. The user can continue the conversation with the deceased and have an emotional experience.

[0254] (Example 1)

[0255] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0256] In modern society, there is a growing need to maintain emotional connections with deceased loved ones by recreating conversations and memories in virtual spaces. However, conventional technologies struggle to reproduce the realistic voice and appearance of the deceased, resulting in an unrealistic experience.

[0257] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0258] In this invention, the server includes means for analyzing audio and image information acquired by the user and extracting the characteristics of a person from the audio and image information; means for generating a three-dimensional representation of the person based on the extracted characteristics; and means for generating acoustic information based on the characteristics and reproducing the voice of the person. This makes it possible to have a realistic and personalized conversation with a deceased person in a virtual space.

[0259] "Audio and visual information" refers to information that includes a person's voice data and visual data, and is fundamental information for understanding biological characteristics.

[0260] "Characteristics" refer to the distinctive features of a person extracted from audio and image information, and are the basic data used to generate three-dimensional displays and acoustic information.

[0261] "Three-dimensional representation" refers to a three-dimensional visual representation of a person generated based on their characteristics, and is used as an avatar in a virtual space.

[0262] "Acoustic information" refers to voice data of a person, reproduced based on their characteristics, and is the voice used in dialogue in a virtual space.

[0263] A "virtual space" is a digital environment created using computer technology, where users access and interact with it through an interface.

[0264] "Conversation information" refers to the content of dialogue generated in a virtual space, specifically the response content generated from a set of information based on the user's input.

[0265] An "information set" refers to a collection of information, including databases and knowledge bases, used to generate dialogue in a virtual space.

[0266] This invention is a system for recreating conversations with deceased persons in a virtual space, and its embodiments are described below.

[0267] First, the user uses their device to collect audio and image information of the deceased. This process involves collecting high-quality data using smart devices or dedicated recording equipment. This data is then encrypted and uploaded to a server.

[0268] The server analyzes the received audio and image information. For audio data, speech synthesis software is used to extract the characteristics of the deceased. This uses common speech recognition services and synthesis engines. For image information, a facial recognition algorithm is used to understand the person's appearance and extract features to generate a 3D representation. At this stage, for example, OpenCV or similar image processing libraries are used.

[0269] Based on the extracted characteristics, the server generates a three-dimensional representation. This process is achieved using a three-dimensional modeling tool such as Blender. Furthermore, based on these characteristics, acoustic information is generated to recreate the deceased's voice. This allows the user to interact with the deceased both visually and aurally in a virtual space.

[0270] The virtual space is a digital environment accessible to users. Users log in to this space via a terminal and enjoy conversations using a 3D display and audio information of the deceased provided by the server. In this process, conversations are generated in real time using a generative AI model, and prompt sentences are processed in response to user input, resulting in natural and personalized conversations.

[0271] For example, if a user uses the prompt "Tell me about a past trip," the server will use this request to search for memories and related information about the deceased and generate a response such as, "That trip to the beach was really fun. I remember you collecting seashells on the sand." This response is reproduced in a voice that mimics the deceased's voice.

[0272] The distinguishing feature of this invention is that, by integrating these technologies into a system, it can provide users with emotional healing by offering a realistic and personalized virtual experience with the deceased and strengthening their emotional connection.

[0273] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0274] Step 1:

[0275] The user uses a device to collect audio and image information. The input used is video of a person speaking or a portrait image. The device temporarily stores the collected data and prepares it for transmission to a server. Specifically, it records high-resolution video using a smartphone or digital camera and records audio using a microphone.

[0276] Step 2:

[0277] The device uploads the collected audio and image information to the server. The input is the data collected in step 1, and the output is the data sent to the server. Here, data is encrypted using a secure communication protocol such as HTTPS to protect privacy. The device notifies the user when the transmission is complete.

[0278] Step 3:

[0279] The server analyzes the received audio and image data. The input is data received from the terminal, and the output is extracted characteristic data of the person. A speech recognition engine is used to analyze the audio data, extracting the pitch and tone of the voice. Image data is processed by a facial recognition algorithm to extract data on distinctive appearances. This process provides basic information to identify the characteristics of the deceased.

[0280] Step 4:

[0281] The server generates a three-dimensional representation using the extracted feature data. The input is the feature data obtained in step 3, and the output is a three-dimensional representation that can be displayed in a virtual space. Specifically, 3D modeling software is used to create an avatar based on the extracted data. In addition, acoustic information is generated, and speech synthesis is performed to reproduce the deceased person's voice.

[0282] Step 5:

[0283] The server integrates the generated 3D display and audio information into the virtual space. The input is the data generated in step 4, and the output is avatar and audio data usable within the virtual space. These are set up by virtual environment creation software and made accessible to the user. Specifically, a digital environment is created using tools such as Unity or Unreal Engine.

[0284] Step 6:

[0285] The user can access the virtual space using the terminal and have conversations with the avatar of the deceased. As input, the user sends questions and comments to the virtual space using prompt sentences. Based on the prompt, the server searches for relevant information from the knowledge base and generates a response. The output is the response voice presented as a natural conversation to the user. Specifically, using an AI model, a response such as "That trip to the seaside was really enjoyable." is generated for the prompt "Tell me about past trips."

[0286] (Application Example 1)

[0287] Next, Application Example 1 will be described. In the following description, the data processing device 12 is referred to as the "server", and the smart glasses 214 are referred to as the "terminal".

[0288] In the conventional dialogue system in the virtual space, when the user experiences a conversation with the deceased, the linkage with the real world is not sufficient, so it is difficult to deepen the emotional connection. Also, it was difficult to reproduce the experience at a memorable place with the deceased within the virtual space. Therefore, there is a need to provide a more immersive and touching virtual experience.

[0289] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0290] In this invention, the server includes means for analyzing the voice and video information of the person acquired by the user and generating a three-dimensional model, means for reproducing the voice of the person based on the voice information, means for enabling a conversation with the person in the virtual reality space using the generated three-dimensional model and voice, and means for recognizing the scenery of the real world and providing content linked with the person in the virtual reality space. Thereby, the user can enjoy a deeper memory experience while linking the emotional reunion scene with the deceased to the scenery of the real world.

[0291] A "user" is a person who uses the system to experience a conversation with the deceased.

[0292] "Audio and video information" refers to recorded voice and video data related to the deceased.

[0293] A "three-dimensional model" is a digital representation created to recreate the appearance of a deceased person in three dimensions.

[0294] A "sound reproduction device" is a device that includes technology for artificially generating the voice of a deceased person based on collected audio information.

[0295] A "virtual reality space" is an interactive artificial environment recreated using digital technology.

[0296] A "device for recognizing real-world landscapes" is a device that uses cameras and sensors to digitize the surrounding physical environment and acquire information.

[0297] An "information storage database" is a digital storage device used to organize and store information about the aforementioned person or related matters.

[0298] This invention is a system for facilitating dialogue with a deceased person in a virtual reality space. This system includes a terminal owned by the user, a server for data processing, and a device capable of displaying the virtual reality environment.

[0299] Users acquire audio and video information about the deceased through their devices and securely upload it to a server in the cloud. The server uses this data to generate a 3D model of the deceased. Facial recognition algorithms are used to generate the 3D model, reproducing the deceased's distinctive appearance. For audio information, speech synthesis technology is used to reproduce the deceased's voice quality and speaking style. Specifically, Amazon Polly and Google Cloud Text-to-Speech are used for speech synthesis, and AR technology is used for 3D model generation.

[0300] The server recognizes real - world landscape data and provides an interactive experience with the deceased when the user views that landscape through a virtual - reality device. When the user visits a specific location, they can share experiences and memories related to that location with the deceased in the virtual - reality space. An example of a prompt sentence is, "The user has identified a location. Please share the memories related to that location."

[0301] With this system, users can more realistically experience an emotional reunion with the deceased and reproduce memories with deep emotions. This service that utilizes the virtual - reality space provides new possibilities to expand the scope of experiences beyond simply viewing data.

[0302] The flow of the specific process in Application Example 1 will be described using FIG. 12.

[0303] Step 1:

[0304] The user obtains the voice and video information of the deceased through the terminal. The input is the recorded voice data and video data. These data are temporarily stored in the terminal and used for subsequent processing. Specific operations include the user using a mobile terminal or computer to record voice, take photos, or shoot videos.

[0305] Step 2:

[0306] The terminal uploads the obtained voice and video information to the server on the cloud. The input is the voice and video data in the terminal, and the output is the data safely transferred to the server. In this process, data encryption and data transfer protocols (e.g., HTTPS) are used to ensure that the information is transmitted securely.

[0307] Step 3:

[0308] The server analyzes uploaded audio and video information to generate a three-dimensional model of the deceased. The input is audio and video data stored on the server, and the output is a three-dimensional model of the deceased. This process utilizes machine learning techniques and facial recognition algorithms to extract the deceased's physical characteristics and reproduce them as a three-dimensional model. The actual operation involves extracting feature points from video data and constructing a three-dimensional digital model based on these points.

[0309] Step 4:

[0310] The server analyzes audio information and performs speech synthesis to recreate the deceased person's voice. The input is audio data, and the output is the recreated voice of the deceased. Amazon Polly and Google Cloud Text-to-Speech are used for speech synthesis technology, which generates digital speech that recreates the deceased person's voice quality and speaking style. The specific operations are analysis of voice characteristics and generation of recreated speech data.

[0311] Step 5:

[0312] The user accesses a virtual space using a virtual reality device and initiates an interaction with a generated three-dimensional model using voice. The input is the user's visual and auditory information, and the output is the experience of interacting with the deceased's avatar. When the user moves to a specific location, the system recognizes the location and generates relevant prompts, providing content that is relevant to the deceased. Specific actions include obtaining the user's location information and initiating a conversation about memories of the deceased based on that information.

[0313] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0314] This invention combines an emotion engine with a system that generates a three-dimensional model by analyzing audio and video information of a deceased person acquired by the user, enabling interaction with the deceased person in a virtual space. This makes it possible to recognize the user's emotions in real time and provide interaction accordingly.

[0315] First, the user uses a device to record audio and video information of the deceased. This information is sent from the device to a server. The server analyzes the characteristics of the deceased's voice from the transmitted audio information and extracts the characteristics of the deceased's appearance from the video information. This generates a three-dimensional model of the deceased and a reconstructed voice.

[0316] Next, the system uses an emotion engine to recognize the user's emotions. When the user accesses the virtual space, the terminal captures the user's facial expressions and voice, and uses the emotion engine to perform real-time emotion analysis. Based on this analysis, the server adjusts the dialogue. For example, if the user has a sad expression, the server generates a response that includes comfort and empathy, and the deceased person's avatar in the virtual space responds appropriately.

[0317] As a concrete example, consider a situation where a user is overwhelmed with emotion and about to cry when reunited with a deceased loved one in a virtual space. At this point, the emotion engine detects the user's tears and sends "sadness" emotion data to the server. Based on this information, the server generates a warm message such as, "Don't cry, we can always meet here," and the avatar of the deceased person in the virtual space speaks that message to the user.

[0318] This invention provides users with an experience that allows them to connect deeply and emotionally with deceased loved ones in a virtual space, enabling flexible conversations that respond to changes in their emotions. As a result, users can experience deeper healing and satisfaction.

[0319] The following describes the processing flow.

[0320] Step 1:

[0321] The user uses the device to record audio and video of the deceased. The recording and video are done in high resolution, and the saved data is stored in the device's temporary storage.

[0322] Step 2:

[0323] The device uploads the collected audio and video data to the server. This process is carried out through an encrypted, secure communication channel to ensure the confidentiality of the data.

[0324] Step 3:

[0325] The server receives the transmitted audio data and applies a speech recognition algorithm to analyze the characteristics of the deceased's voice. From the results of this analysis, the deceased's speaking style and voice quality are digitized.

[0326] Step 4:

[0327] The server processes the video data and uses facial recognition technology to extract the deceased's physical characteristics. This process allows the server to understand the three-dimensional structure in 3D space, which is then used to construct a 3D model.

[0328] Step 5:

[0329] The server generates a three-dimensional model and a reproduced voice of the deceased. These serve as materials for users to experience visually and aurally within the virtual space.

[0330] Step 6:

[0331] When a user logs into the virtual space, the device uses the user's camera and microphone to continuously capture the user's facial expressions and voice, and sends that data to the emotion engine.

[0332] Step 7:

[0333] The emotion engine analyzes received data and evaluates the user's emotional state in real time. The analysis results identify the user's current emotion, such as sadness, happiness, or surprise.

[0334] Step 8:

[0335] The server dynamically generates dialogue content within the virtual space based on emotional data transmitted from the emotion engine. It selects and adjusts appropriate responses according to the user's emotions.

[0336] Step 9:

[0337] Within the virtual space, the conversation generated by the server is reproduced as audio and delivered to the user through the deceased person's avatar. This allows the user to naturally engage in emotional dialogue with the deceased.

[0338] (Example 2)

[0339] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0340] In modern society, there is a growing demand for new forms of interaction with deceased loved ones. However, it is difficult to retain memories of the deceased, and there is a lack of methods to recreate emotional connections. Conventional systems have limitations in the quality of recreating the deceased and the interaction, and in particular, they have the challenge of not being able to have flexible conversations that respond to the user's emotions.

[0341] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0342] In this invention, the server includes means for analyzing voice and video information of a person acquired by the user and generating a three-dimensional model of the person; means for reproducing the person's voice based on the voice information; means for enabling dialogue with the person in a virtual space using the generated three-dimensional model and the reproduced voice; and means for recognizing the user's emotions in real time and generating a corresponding dialogue. This makes it possible to provide flexible and appropriate dialogue that responds to the user's emotions while maintaining a deep emotional connection with the deceased.

[0343] "User" refers to an individual who operates the system and inputs audio and video information.

[0344] "Person" refers to an individual from whom audio and video information is acquired, and who is the subject of a three-dimensional model or a reconstructed voice.

[0345] A "three-dimensional model" refers to a three-dimensional digital representation generated based on a person's physical characteristics.

[0346] "Audio information" refers to digital audio data that includes the characteristics of a person's voice.

[0347] "Reproduced voice" refers to artificial voice data that reproduces a person's voice, generated based on the original audio information.

[0348] A "virtual space" refers to a digital environment constructed using digital technology, where three-dimensional models and reproduced sounds are placed, and users can interact with them.

[0349] "Methods for recognizing emotions in real time" refers to technologies that analyze data such as the user's facial expressions and voice to identify their emotions at that moment.

[0350] "Means of generating dialogue" refers to the process of generating appropriate responses in response to the user's emotions and input.

[0351] A "generative AI model" refers to an algorithm used to learn from data and accomplish a specific task.

[0352] A "prompt sentence" refers to an instruction sentence input into a generative AI model, which serves as a criterion for determining the content of the output dialogue.

[0353] This invention provides a virtual space that enables users to have emotional interactions with deceased loved ones. The user first records audio and video information of the deceased using a device. Suitable hardware includes common smartphones, tablets, and camera devices. This information is transmitted from the device to a server. The server receives this data and analyzes the characteristics of the deceased's voice using an audio feature extraction algorithm. It also extracts the deceased's physical characteristics using deep learning-based image processing software. This generates a three-dimensional model and synthesized voice of the deceased.

[0354] When a user accesses the virtual space, the terminal captures the user's facial expressions and voice, and analyzes the user's emotions in real time through emotion recognition software. This process requires a camera and microphone. Based on this analysis, the server uses a generative AI model to generate appropriate dialogue. An example of a prompt might be, "Generate flexible dialogue that responds to the user's emotions and convey the content in the deceased's voice."

[0355] As a concrete example, suppose a user reunites with a deceased loved one in a virtual space and displays a sad expression. At this point, emotion recognition software identifies the user's expression as "sadness" and sends that data to a server. The server then generates a warm message such as, "Don't cry, I'll always be waiting here for you." This message is conveyed to the user in the virtual space through the deceased person's avatar.

[0356] This system allows users to experience deep emotional connection with their deceased loved ones and enjoy intimate and realistic interactions that respond to their changing emotions.

[0357] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0358] Step 1:

[0359] Users record audio and video information of the deceased using their own devices. Input includes audio and video files obtained from smartphones and cameras. This information is saved in digital format and prepared for transmission to a database.

[0360] Step 2:

[0361] The terminal transmits the recorded audio and video information to the server. In this step, the information is transmitted over the internet as data packets. The input data consists of the audio and video files obtained in the previous step. The output is the encrypted information that arrives securely at the server.

[0362] Step 3:

[0363] The server analyzes the received audio information. It processes the input audio data through an audio feature extraction algorithm, outputting features such as voice intonation, tone, and speaking speed as numerical data. Specifically, an audio signal processing library is used.

[0364] Step 4:

[0365] The server analyzes the received video information. It receives video data as input and uses image processing techniques to extract features such as facial shape and posture. The output is data for a three-dimensional model based on these features. Specifically, it performs image analysis using deep learning.

[0366] Step 5:

[0367] The server generates a 3D model of the deceased and a reconstructed voice. This step uses the voice and video feature data obtained in the previous step. Based on the input data, a generating AI model is used to output a digital avatar and synthesized voice.

[0368] Step 6:

[0369] When a user accesses the virtual space, the terminal captures the user's facial expressions and voice in real time. Input consists of audio and video data obtained from the camera and microphone. Output is a data stream of the user's facial expressions and voice, which is sent to emotion recognition software.

[0370] Step 7:

[0371] The server analyzes the user's emotions in real time. It uses facial and audio data received from the terminal as input. An emotion recognition algorithm analyzes this data and outputs the user's emotional state as numerical data.

[0372] Step 8:

[0373] The server uses a generative AI model to generate dialogue that responds to the user's emotions. The prompt is set to "Generate a warm response based on the user's emotions," and the input is the emotion data obtained in the previous step. The output is the dialogue that the deceased person's avatar should say.

[0374] Step 9:

[0375] The server sends the generated dialogue to the deceased person's avatar in the virtual space, allowing it to interact with the user. The input is the generated dialogue, and the output is the audio and visual dialogue experience provided to the user. The virtual space performs the necessary visual rendering and speech synthesis to reproduce this in real time.

[0376] (Application Example 2)

[0377] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0378] Conventional virtual reality dialogue systems have struggled to provide flexible dialogue that responds to the user's emotional state. There is a growing need to generate appropriate responses based on the user's emotions to foster deeper emotional connections and improve the user experience.

[0379] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0380] In this invention, the server includes means for analyzing voice and video information of a person acquired by the user and generating a three-dimensional model of the person; means for evaluating the user's emotional state and adjusting the dialogue content based on that evaluation; and means for generating a conversation that corresponds to the user's emotions. This enables natural and personalized dialogue that is adapted to emotions.

[0381] A "user" is an individual who operates the system and interacts within the virtual space.

[0382] "Audio information" refers to digital data related to a person's speech, and is fundamental information for a system to reproduce.

[0383] "Visual information" refers to visual data about a person's appearance, which is used to generate a three-dimensional model.

[0384] A "three-dimensional model" is a digital representation of a person in three dimensions, created based on acquired video information.

[0385] "Reproduced speech" refers to the speech expressions of a person that are generated based on audio information.

[0386] A "virtual space" is a digital environment built on a computer, where users interact with each other.

[0387] "Emotional analysis tools" are technologies that evaluate the user's current emotional state and allow the system to determine a response accordingly.

[0388] A "database" is a collection of information that a system uses to store and refer to as needed.

[0389] "Machine learning technology" is a technique that uses algorithms to extract patterns from data and then uses those patterns to make predictions or generate new data.

[0390] The system for implementing this invention is based on a user terminal and a server. First, the user acquires audio and video information using a smart device, such as smart glasses or a head-mounted display. The terminal captures this information and transmits it to the server. The server generates a three-dimensional model from the video information and reconstructs the audio based on the audio information.

[0391] The server incorporates the "Microsoft Azure Cognitive Services" sentiment analysis API and the "OpenAI speech synthesis API." The user's facial expressions and voice data are analyzed in real time using these APIs. Based on the analysis results, the server generates emotionally adaptive dialogue, enabling natural conversations within the virtual space.

[0392] As a concrete example, consider a virtual store scenario. A user reunites with a deceased person, represented as an avatar, in a virtual space, and they tour their favorite places together. At this time, the system detects when the user is smiling, and the deceased person's avatar provides a message such as, "I remember the smile on your face when you chose this." An example of a prompt used here would be, "Generate a warm message when the user is smiling. For example, provide a conversation that recreates a happy moment at a memorable store."

[0393] In this way, by adjusting the interaction in real time according to the user's emotions, the system can provide a deeper emotional connection and create a more satisfying experience for the user.

[0394] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0395] Step 1:

[0396] The terminal captures audio and video information from the user and sends it to the server as digital data. The input is the user's audio and video, and the output is the transmission of this data to the server. This process prepares the basic data necessary for the virtual space.

[0397] Step 2:

[0398] The server generates a three-dimensional model from the received video information. This process utilizes computer graphics technology to create a three-dimensional model of a person. The input is the user's video data, and the output is a three-dimensional model. This generates a virtual avatar based on any given person.

[0399] Step 3:

[0400] The server analyzes voice information and performs speech synthesis based on this analysis. Machine learning techniques are used to capture voice characteristics and reproduce realistic speech. The input is the user's voice data, and the output is the synthesized speech. A generative AI model is used, enabling natural-sounding speech.

[0401] Step 4:

[0402] The server senses the user's facial expressions in real time and evaluates their emotional state using an emotion analysis API. The input is the user's facial expression data, and the output is the result of the emotion analysis. Specifically, the API quantifies the characteristics of each facial expression and determines the emotional state.

[0403] Step 5:

[0404] The server generates prompt sentences based on the sentiment analysis results and constructs an appropriate dialogue. In this step, the generating AI model determines the dialogue content based on the prompt sentences and presents it to the user. The input is data on the emotional state, and the output is a dialogue adapted to that emotion.

[0405] Step 6:

[0406] Within the virtual space, a conversation generated by the user's device is played back. The user can use a smart device to enjoy a conversation with the deceased through a generated 3D model and reproduced voice. The input is the 3D model and voice generated in the previous step, and the output is the conversation experienced by the user.

[0407] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0408] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0409] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0410] [Third Embodiment]

[0411] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0412] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0413] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0414] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0415] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0416] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0417] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0418] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0419] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0420] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0421] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0422] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0423] This invention is a system that generates a three-dimensional model using the voice and video information of a deceased person, enabling interaction with the deceased person in a virtual space. This system operates as follows:

[0424] First, the user uses their device to collect audio and video information about the deceased. The collected information is securely uploaded from the device to a server. The server analyzes this information and extracts the deceased's physical and vocal characteristics. The extracted characteristics are used to generate a three-dimensional model and reconstruct the voice.

[0425] The server uses a facial recognition algorithm to identify the deceased's distinctive appearance from video information and generates a three-dimensional model based on this. For audio information, a speech synthesis engine is used to analyze it and generate audio data to reproduce the deceased's speaking style and voice quality.

[0426] The generated 3D model and audio data are converted into a format usable in a virtual space. This virtual space is accessible to users via their devices and serves as a place to interact with the deceased person's avatar. When a user logs into the virtual space, the server displays the deceased person's 3D avatar and provides a conversation using reproduced audio.

[0427] For example, if a user asks the deceased to "tell me about a past trip," the device sends this request to the server. The server searches for relevant information in its database and generates a response based on shared memories with the deceased. For instance, it might provide a response that reflects the deceased's personality, such as, "That trip to the beach was so much fun. I remember you collecting seashells on the sand." The generated response is then reproduced in the virtual space using the deceased's voice.

[0428] This invention thus provides a realistic and personalized virtual experience, strengthening the emotional connection with the deceased. Users can find emotional healing through reuniting with the deceased and recreating memories in the virtual space.

[0429] The following describes the processing flow.

[0430] Step 1:

[0431] The user uses the device to record the deceased's voice and video. The device uses a microphone and camera to acquire high-quality audio and video data, which is then stored in temporary data storage.

[0432] Step 2:

[0433] The device uploads the collected audio and video data to the server. The upload uses an encrypted, secure communication protocol to protect the confidentiality of the data.

[0434] Step 3:

[0435] The server analyzes the received audio data and extracts speech features such as the deceased person's speaking style, pitch, and intonation. This process uses a speech recognition algorithm.

[0436] Step 4:

[0437] The server analyzes the video data to identify the deceased's appearance and distinctive facial patterns. Using facial recognition technology, it extracts feature points of the face in three-dimensional space.

[0438] Step 5:

[0439] The server uses the extracted speech features to recreate the deceased person's voice using a speech synthesis engine. The synthesized voice is then converted into a format that can be played back as natural-sounding speech.

[0440] Step 6:

[0441] The server uses feature points extracted from the video footage and 3D modeling software to generate a 3D avatar of the deceased. This avatar is a digital model that preserves the visual characteristics of the deceased, including their appearance.

[0442] Step 7:

[0443] The user accesses the virtual space using a terminal and begins a conversation with the deceased based on a 3D avatar and voice data provided by the server.

[0444] Step 8:

[0445] The server receives input from the user in the virtual space, searches the database, and generates natural conversations based on relevant episodes and information about the deceased.

[0446] Step 9:

[0447] The generated conversation content is provided to the user in a virtual space using reconstructed audio. The user can continue the conversation with the deceased and have an emotional experience.

[0448] (Example 1)

[0449] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0450] In modern society, there is a growing need to maintain emotional connections with deceased loved ones by recreating conversations and memories in virtual spaces. However, conventional technologies struggle to reproduce the realistic voice and appearance of the deceased, resulting in an unrealistic experience.

[0451] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0452] In this invention, the server includes means for analyzing audio and image information acquired by the user and extracting the characteristics of a person from the audio and image information; means for generating a three-dimensional representation of the person based on the extracted characteristics; and means for generating acoustic information based on the characteristics and reproducing the voice of the person. This makes it possible to have a realistic and personalized conversation with a deceased person in a virtual space.

[0453] "Audio and visual information" refers to information that includes a person's voice data and visual data, and is fundamental information for understanding biological characteristics.

[0454] "Characteristics" refer to the distinctive features of a person extracted from audio and image information, and are the basic data used to generate three-dimensional displays and acoustic information.

[0455] "Three-dimensional representation" refers to a three-dimensional visual representation of a person generated based on their characteristics, and is used as an avatar in a virtual space.

[0456] "Acoustic information" refers to voice data of a person, reproduced based on their characteristics, and is the voice used in dialogue in a virtual space.

[0457] A "virtual space" is a digital environment created using computer technology, where users access and interact with it through an interface.

[0458] "Conversation information" refers to the content of dialogue generated in a virtual space, specifically the response content generated from a set of information based on the user's input.

[0459] An "information set" refers to a collection of information, including databases and knowledge bases, used to generate dialogue in a virtual space.

[0460] This invention is a system for recreating conversations with deceased persons in a virtual space, and its embodiments are described below.

[0461] First, the user uses their device to collect audio and image information of the deceased. This process involves collecting high-quality data using smart devices or dedicated recording equipment. This data is then encrypted and uploaded to a server.

[0462] The server analyzes the received audio and image information. For audio data, speech synthesis software is used to extract the characteristics of the deceased. This uses common speech recognition services and synthesis engines. For image information, a facial recognition algorithm is used to understand the person's appearance and extract features to generate a 3D representation. At this stage, for example, OpenCV or similar image processing libraries are used.

[0463] Based on the extracted characteristics, the server generates a three-dimensional representation. This process is achieved using a three-dimensional modeling tool such as Blender. Furthermore, based on these characteristics, acoustic information is generated to recreate the deceased's voice. This allows the user to interact with the deceased both visually and aurally in a virtual space.

[0464] The virtual space is a digital environment accessible to users. Users log in to this space via a terminal and enjoy conversations using a 3D display and audio information of the deceased provided by the server. In this process, conversations are generated in real time using a generative AI model, and prompt sentences are processed in response to user input, resulting in natural and personalized conversations.

[0465] For example, if a user uses the prompt "Tell me about a past trip," the server will use this request to search for memories and related information about the deceased and generate a response such as, "That trip to the beach was really fun. I remember you collecting seashells on the sand." This response is reproduced in a voice that mimics the deceased's voice.

[0466] The distinguishing feature of this invention is that, by integrating these technologies into a system, it can provide users with emotional healing by offering a realistic and personalized virtual experience with the deceased and strengthening their emotional connection.

[0467] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0468] Step 1:

[0469] The user uses a device to collect audio and image information. The input used is video of a person speaking or a portrait image. The device temporarily stores the collected data and prepares it for transmission to a server. Specifically, it records high-resolution video using a smartphone or digital camera and records audio using a microphone.

[0470] Step 2:

[0471] The device uploads the collected audio and image information to the server. The input is the data collected in step 1, and the output is the data sent to the server. Here, data is encrypted using a secure communication protocol such as HTTPS to protect privacy. The device notifies the user when the transmission is complete.

[0472] Step 3:

[0473] The server analyzes the received audio and image data. The input is data received from the terminal, and the output is extracted characteristic data of the person. A speech recognition engine is used to analyze the audio data, extracting the pitch and tone of the voice. Image data is processed by a facial recognition algorithm to extract data on distinctive appearances. This process provides basic information to identify the characteristics of the deceased.

[0474] Step 4:

[0475] The server generates a three-dimensional representation using the extracted feature data. The input is the feature data obtained in step 3, and the output is a three-dimensional representation that can be displayed in a virtual space. Specifically, 3D modeling software is used to create an avatar based on the extracted data. In addition, acoustic information is generated, and speech synthesis is performed to reproduce the deceased person's voice.

[0476] Step 5:

[0477] The server integrates the generated 3D display and audio information into the virtual space. The input is the data generated in step 4, and the output is avatar and audio data usable within the virtual space. These are set up by virtual environment creation software and made accessible to the user. Specifically, a digital environment is created using tools such as Unity or Unreal Engine.

[0478] Step 6:

[0479] Users can access a virtual space using a device and converse with an avatar of a deceased person. As input, users send questions and comments to the virtual space using prompt text. Based on this prompt, the server searches a knowledge base for relevant information and generates a response. The output is a voice response presented as a natural conversation to the user. Specifically, using an AI model, a response such as "Tell me about your old trip" might be generated, for example, "That trip to the beach was really fun."

[0480] (Application Example 1)

[0481] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0482] Conventional virtual reality dialogue systems have struggled to foster deep emotional connections with deceased loved ones because they lack sufficient integration with the real world. Furthermore, recreating memorable experiences in places associated with the deceased within the virtual space has been difficult. Therefore, there is a need to provide a more immersive and emotionally impactful virtual experience.

[0483] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0484] In this invention, the server includes means for analyzing voice and video information of a person acquired by the user and generating a three-dimensional model; means for reproducing the person's voice based on the voice information; means for enabling dialogue with the person in a virtual reality space using the generated three-dimensional model and voice; and means for recognizing real-world scenery and providing content linked to the person in the virtual reality space. This makes it possible for the user to enjoy a deeper reminiscing experience by linking the emotional reunion with the deceased with real-world scenery.

[0485] A "user" is a person who uses the system to experience a conversation with the deceased.

[0486] "Audio and video information" refers to recorded voice and video data related to the deceased.

[0487] A "three-dimensional model" is a digital representation created to recreate the appearance of a deceased person in three dimensions.

[0488] A "sound reproduction device" is a device that includes technology for artificially generating the voice of a deceased person based on collected audio information.

[0489] A "virtual reality space" is an interactive artificial environment recreated using digital technology.

[0490] A "device for recognizing real-world landscapes" is a device that uses cameras and sensors to digitize the surrounding physical environment and acquire information.

[0491] An "information storage database" is a digital storage device used to organize and store information about the aforementioned person or related matters.

[0492] This invention is a system for facilitating dialogue with a deceased person in a virtual reality space. This system includes a terminal owned by the user, a server for data processing, and a device capable of displaying the virtual reality environment.

[0493] Users acquire audio and video information about the deceased through their devices and securely upload it to a server in the cloud. The server uses this data to generate a 3D model of the deceased. Facial recognition algorithms are used to generate the 3D model, reproducing the deceased's distinctive appearance. For audio information, speech synthesis technology is used to reproduce the deceased's voice quality and speaking style. Specifically, Amazon Polly and Google Cloud Text-to-Speech are used for speech synthesis, and AR technology is used for 3D model generation.

[0494] The server recognizes real-world landscape data and provides a responsive experience with the deceased when the user views that landscape through a virtual reality device. When the user visits a specific location, they can share experiences and memories associated with the deceased and that location in the virtual reality space. An example of a prompt message is, "The user has identified a location. Please share your memories associated with that location."

[0495] This system allows users to experience emotional reunions with deceased loved ones in a more realistic way, and to recreate memories with deep emotion. This service, utilizing virtual reality space, goes beyond mere data viewing, offering new possibilities that broaden the scope of experiences.

[0496] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0497] Step 1:

[0498] The user obtains audio and video information of the deceased through the device. The input consists of recorded audio and video data. This data is temporarily stored on the device and used for subsequent processing. Specific actions include the user recording audio or taking photos and videos using a mobile device or computer.

[0499] Step 2:

[0500] The device uploads the acquired audio and video information to a server in the cloud. The input is the audio and video data on the device, and the output is the data securely transferred to the server. This process uses data encryption and data transfer protocols (e.g., HTTPS) to ensure that the information is transmitted securely.

[0501] Step 3:

[0502] The server analyzes uploaded audio and video information to generate a three-dimensional model of the deceased. The input is audio and video data stored on the server, and the output is a three-dimensional model of the deceased. This process utilizes machine learning techniques and facial recognition algorithms to extract the deceased's physical characteristics and reproduce them as a three-dimensional model. The actual operation involves extracting feature points from video data and constructing a three-dimensional digital model based on these points.

[0503] Step 4:

[0504] The server analyzes audio information and performs speech synthesis to recreate the deceased person's voice. The input is audio data, and the output is the recreated voice of the deceased. Amazon Polly and Google Cloud Text-to-Speech are used for speech synthesis technology, which generates digital speech that recreates the deceased person's voice quality and speaking style. The specific operations are analysis of voice characteristics and generation of recreated speech data.

[0505] Step 5:

[0506] The user accesses a virtual space using a virtual reality device and initiates an interaction with a generated three-dimensional model using voice. The input is the user's visual and auditory information, and the output is the experience of interacting with the deceased's avatar. When the user moves to a specific location, the system recognizes the location and generates relevant prompts, providing content that is relevant to the deceased. Specific actions include obtaining the user's location information and initiating a conversation about memories of the deceased based on that information.

[0507] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0508] This invention combines an emotion engine with a system that generates a three-dimensional model by analyzing audio and video information of a deceased person acquired by the user, enabling interaction with the deceased person in a virtual space. This makes it possible to recognize the user's emotions in real time and provide interaction accordingly.

[0509] First, the user uses a device to record audio and video information of the deceased. This information is sent from the device to a server. The server analyzes the characteristics of the deceased's voice from the transmitted audio information and extracts the characteristics of the deceased's appearance from the video information. This generates a three-dimensional model of the deceased and a reconstructed voice.

[0510] Next, the system uses an emotion engine to recognize the user's emotions. When the user accesses the virtual space, the terminal captures the user's facial expressions and voice, and uses the emotion engine to perform real-time emotion analysis. Based on this analysis, the server adjusts the dialogue. For example, if the user has a sad expression, the server generates a response that includes comfort and empathy, and the deceased person's avatar in the virtual space responds appropriately.

[0511] As a concrete example, consider a situation where a user is overwhelmed with emotion and about to cry when reunited with a deceased loved one in a virtual space. At this point, the emotion engine detects the user's tears and sends "sadness" emotion data to the server. Based on this information, the server generates a warm message such as, "Don't cry, we can always meet here," and the avatar of the deceased person in the virtual space speaks that message to the user.

[0512] This invention provides users with an experience that allows them to connect deeply and emotionally with deceased loved ones in a virtual space, enabling flexible conversations that respond to changes in their emotions. As a result, users can experience deeper healing and satisfaction.

[0513] The following describes the processing flow.

[0514] Step 1:

[0515] The user uses the device to record audio and video of the deceased. The recording and video are done in high resolution, and the saved data is stored in the device's temporary storage.

[0516] Step 2:

[0517] The device uploads the collected audio and video data to the server. This process is carried out through an encrypted, secure communication channel to ensure the confidentiality of the data.

[0518] Step 3:

[0519] The server receives the transmitted audio data and applies a speech recognition algorithm to analyze the characteristics of the deceased's voice. From the results of this analysis, the deceased's speaking style and voice quality are digitized.

[0520] Step 4:

[0521] The server processes the video data and uses facial recognition technology to extract the deceased's physical characteristics. This process allows the server to understand the three-dimensional structure in 3D space, which is then used to construct a 3D model.

[0522] Step 5:

[0523] The server generates a three-dimensional model and a reproduced voice of the deceased. These serve as materials for users to experience visually and aurally within the virtual space.

[0524] Step 6:

[0525] When a user logs into the virtual space, the device uses the user's camera and microphone to continuously capture the user's facial expressions and voice, and sends that data to the emotion engine.

[0526] Step 7:

[0527] The emotion engine analyzes received data and evaluates the user's emotional state in real time. The analysis results identify the user's current emotion, such as sadness, happiness, or surprise.

[0528] Step 8:

[0529] The server dynamically generates dialogue content within the virtual space based on emotional data transmitted from the emotion engine. It selects and adjusts appropriate responses according to the user's emotions.

[0530] Step 9:

[0531] Within the virtual space, the conversation generated by the server is reproduced as audio and delivered to the user through the deceased person's avatar. This allows the user to naturally engage in emotional dialogue with the deceased.

[0532] (Example 2)

[0533] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0534] In modern society, there is a growing demand for new forms of interaction with deceased loved ones. However, it is difficult to retain memories of the deceased, and there is a lack of methods to recreate emotional connections. Conventional systems have limitations in the quality of recreating the deceased and the interaction, and in particular, they have the challenge of not being able to have flexible conversations that respond to the user's emotions.

[0535] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0536] In this invention, the server includes means for analyzing voice and video information of a person acquired by the user and generating a three-dimensional model of the person; means for reproducing the person's voice based on the voice information; means for enabling dialogue with the person in a virtual space using the generated three-dimensional model and the reproduced voice; and means for recognizing the user's emotions in real time and generating a corresponding dialogue. This makes it possible to provide flexible and appropriate dialogue that responds to the user's emotions while maintaining a deep emotional connection with the deceased.

[0537] "User" refers to an individual who operates the system and inputs audio and video information.

[0538] "Person" refers to an individual from whom audio and video information is acquired, and who is the subject of a three-dimensional model or a reconstructed voice.

[0539] A "three-dimensional model" refers to a three-dimensional digital representation generated based on a person's physical characteristics.

[0540] "Audio information" refers to digital audio data that includes the characteristics of a person's voice.

[0541] "Reproduced voice" refers to artificial voice data that reproduces a person's voice, generated based on the original audio information.

[0542] A "virtual space" refers to a digital environment constructed using digital technology, where three-dimensional models and reproduced sounds are placed, and users can interact with them.

[0543] "Methods for recognizing emotions in real time" refers to technologies that analyze data such as the user's facial expressions and voice to identify their emotions at that moment.

[0544] "Means of generating dialogue" refers to the process of generating appropriate responses in response to the user's emotions and input.

[0545] A "generative AI model" refers to an algorithm used to learn from data and accomplish a specific task.

[0546] A "prompt sentence" refers to an instruction sentence input into a generative AI model, which serves as a criterion for determining the content of the output dialogue.

[0547] This invention provides a virtual space that enables users to have emotional interactions with deceased loved ones. The user first records audio and video information of the deceased using a device. Suitable hardware includes common smartphones, tablets, and camera devices. This information is transmitted from the device to a server. The server receives this data and analyzes the characteristics of the deceased's voice using an audio feature extraction algorithm. It also extracts the deceased's physical characteristics using deep learning-based image processing software. This generates a three-dimensional model and synthesized voice of the deceased.

[0548] When a user accesses the virtual space, the terminal captures the user's facial expressions and voice, and analyzes the user's emotions in real time through emotion recognition software. This process requires a camera and microphone. Based on this analysis, the server uses a generative AI model to generate appropriate dialogue. An example of a prompt might be, "Generate flexible dialogue that responds to the user's emotions and convey the content in the deceased's voice."

[0549] As a concrete example, suppose a user reunites with a deceased loved one in a virtual space and displays a sad expression. At this point, emotion recognition software identifies the user's expression as "sadness" and sends that data to a server. The server then generates a warm message such as, "Don't cry, I'll always be waiting here for you." This message is conveyed to the user in the virtual space through the deceased person's avatar.

[0550] This system allows users to experience deep emotional connection with their deceased loved ones and enjoy intimate and realistic interactions that respond to their changing emotions.

[0551] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0552] Step 1:

[0553] Users record audio and video information of the deceased using their own devices. Input includes audio and video files obtained from smartphones and cameras. This information is saved in digital format and prepared for transmission to a database.

[0554] Step 2:

[0555] The terminal transmits the recorded audio and video information to the server. In this step, the information is transmitted over the internet as data packets. The input data consists of the audio and video files obtained in the previous step. The output is the encrypted information that arrives securely at the server.

[0556] Step 3:

[0557] The server analyzes the received audio information. It processes the input audio data through an audio feature extraction algorithm, outputting features such as voice intonation, tone, and speaking speed as numerical data. Specifically, an audio signal processing library is used.

[0558] Step 4:

[0559] The server analyzes the received video information. It receives video data as input and uses image processing techniques to extract features such as facial shape and posture. The output is data for a three-dimensional model based on these features. Specifically, it performs image analysis using deep learning.

[0560] Step 5:

[0561] The server generates a 3D model of the deceased and a reconstructed voice. This step uses the voice and video feature data obtained in the previous step. Based on the input data, a generating AI model is used to output a digital avatar and synthesized voice.

[0562] Step 6:

[0563] When a user accesses the virtual space, the terminal captures the user's facial expressions and voice in real time. Input consists of audio and video data obtained from the camera and microphone. Output is a data stream of the user's facial expressions and voice, which is sent to emotion recognition software.

[0564] Step 7:

[0565] The server analyzes the user's emotions in real time. It uses facial and audio data received from the terminal as input. An emotion recognition algorithm analyzes this data and outputs the user's emotional state as numerical data.

[0566] Step 8:

[0567] The server uses a generative AI model to generate dialogue that responds to the user's emotions. The prompt is set to "Generate a warm response based on the user's emotions," and the input is the emotion data obtained in the previous step. The output is the dialogue that the deceased person's avatar should say.

[0568] Step 9:

[0569] The server sends the generated dialogue to the deceased person's avatar in the virtual space, allowing it to interact with the user. The input is the generated dialogue, and the output is the audio and visual dialogue experience provided to the user. The virtual space performs the necessary visual rendering and speech synthesis to reproduce this in real time.

[0570] (Application Example 2)

[0571] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0572] Conventional virtual reality dialogue systems have struggled to provide flexible dialogue that responds to the user's emotional state. There is a growing need to generate appropriate responses based on the user's emotions to foster deeper emotional connections and improve the user experience.

[0573] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0574] In this invention, the server includes means for analyzing voice and video information of a person acquired by the user and generating a three-dimensional model of the person; means for evaluating the user's emotional state and adjusting the dialogue content based on that evaluation; and means for generating a conversation that corresponds to the user's emotions. This enables natural and personalized dialogue that is adapted to emotions.

[0575] A "user" is an individual who operates the system and interacts within the virtual space.

[0576] "Audio information" refers to digital data related to a person's speech, and is fundamental information for a system to reproduce.

[0577] "Visual information" refers to visual data about a person's appearance, which is used to generate a three-dimensional model.

[0578] A "three-dimensional model" is a digital representation of a person in three dimensions, created based on acquired video information.

[0579] "Reproduced speech" refers to the speech expressions of a person that are generated based on audio information.

[0580] A "virtual space" is a digital environment built on a computer, where users interact with each other.

[0581] "Emotional analysis tools" are technologies that evaluate the user's current emotional state and allow the system to determine a response accordingly.

[0582] A "database" is a collection of information that a system uses to store and refer to as needed.

[0583] "Machine learning technology" is a technique that uses algorithms to extract patterns from data and then uses those patterns to make predictions or generate new data.

[0584] The system for implementing this invention is based on a user terminal and a server. First, the user acquires audio and video information using a smart device, such as smart glasses or a head-mounted display. The terminal captures this information and transmits it to the server. The server generates a three-dimensional model from the video information and reconstructs the audio based on the audio information.

[0585] The server incorporates the "Microsoft Azure Cognitive Services" sentiment analysis API and the "OpenAI speech synthesis API." The user's facial expressions and voice data are analyzed in real time using these APIs. Based on the analysis results, the server generates emotionally adaptive dialogue, enabling natural conversations within the virtual space.

[0586] As a concrete example, consider a virtual store scenario. A user reunites with a deceased person, represented as an avatar, in a virtual space, and they tour their favorite places together. At this time, the system detects when the user is smiling, and the deceased person's avatar provides a message such as, "I remember the smile on your face when you chose this." An example of a prompt used here would be, "Generate a warm message when the user is smiling. For example, provide a conversation that recreates a happy moment at a memorable store."

[0587] In this way, by adjusting the interaction in real time according to the user's emotions, the system can provide a deeper emotional connection and create a more satisfying experience for the user.

[0588] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0589] Step 1:

[0590] The terminal captures audio and video information from the user and sends it to the server as digital data. The input is the user's audio and video, and the output is the transmission of this data to the server. This process prepares the basic data necessary for the virtual space.

[0591] Step 2:

[0592] The server generates a three-dimensional model from the received video information. This process utilizes computer graphics technology to create a three-dimensional model of a person. The input is the user's video data, and the output is a three-dimensional model. This generates a virtual avatar based on any given person.

[0593] Step 3:

[0594] The server analyzes voice information and performs speech synthesis based on this analysis. Machine learning techniques are used to capture voice characteristics and reproduce realistic speech. The input is the user's voice data, and the output is the synthesized speech. A generative AI model is used, enabling natural-sounding speech.

[0595] Step 4:

[0596] The server senses the user's facial expressions in real time and evaluates their emotional state using an emotion analysis API. The input is the user's facial expression data, and the output is the result of the emotion analysis. Specifically, the API quantifies the characteristics of each facial expression and determines the emotional state.

[0597] Step 5:

[0598] The server generates prompt sentences based on the sentiment analysis results and constructs an appropriate dialogue. In this step, the generating AI model determines the dialogue content based on the prompt sentences and presents it to the user. The input is data on the emotional state, and the output is a dialogue adapted to that emotion.

[0599] Step 6:

[0600] Within the virtual space, a conversation generated by the user's device is played back. The user can use a smart device to enjoy a conversation with the deceased through a generated 3D model and reproduced voice. The input is the 3D model and voice generated in the previous step, and the output is the conversation experienced by the user.

[0601] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0602] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0603] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[0604] [Fourth Embodiment]

[0605] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0606] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0607] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0608] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0609] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0610] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0611] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0612] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0613] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0614] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0615] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0616] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0617] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0618] This invention is a system that generates a three-dimensional model using the voice and video information of a deceased person, enabling interaction with the deceased person in a virtual space. This system operates as follows:

[0619] First, the user uses their device to collect audio and video information about the deceased. The collected information is securely uploaded from the device to a server. The server analyzes this information and extracts the deceased's physical and vocal characteristics. The extracted characteristics are used to generate a three-dimensional model and reconstruct the voice.

[0620] The server uses a facial recognition algorithm to identify the deceased's distinctive appearance from video information and generates a three-dimensional model based on this. For audio information, a speech synthesis engine is used to analyze it and generate audio data to reproduce the deceased's speaking style and voice quality.

[0621] The generated 3D model and audio data are converted into a format usable in a virtual space. This virtual space is accessible to users via their devices and serves as a place to interact with the deceased person's avatar. When a user logs into the virtual space, the server displays the deceased person's 3D avatar and provides a conversation using reproduced audio.

[0622] For example, if a user asks the deceased to "tell me about a past trip," the device sends this request to the server. The server searches for relevant information in its database and generates a response based on shared memories with the deceased. For instance, it might provide a response that reflects the deceased's personality, such as, "That trip to the beach was so much fun. I remember you collecting seashells on the sand." The generated response is then reproduced in the virtual space using the deceased's voice.

[0623] This invention thus provides a realistic and personalized virtual experience, strengthening the emotional connection with the deceased. Users can find emotional healing through reuniting with the deceased and recreating memories in the virtual space.

[0624] The following describes the processing flow.

[0625] Step 1:

[0626] The user uses the device to record the deceased's voice and video. The device uses a microphone and camera to acquire high-quality audio and video data, which is then stored in temporary data storage.

[0627] Step 2:

[0628] The device uploads the collected audio and video data to the server. The upload uses an encrypted, secure communication protocol to protect the confidentiality of the data.

[0629] Step 3:

[0630] The server analyzes the received audio data and extracts speech features such as the deceased person's speaking style, pitch, and intonation. This process uses a speech recognition algorithm.

[0631] Step 4:

[0632] The server analyzes the video data to identify the deceased's appearance and distinctive facial patterns. Using facial recognition technology, it extracts feature points of the face in three-dimensional space.

[0633] Step 5:

[0634] The server uses the extracted speech features to recreate the deceased person's voice using a speech synthesis engine. The synthesized voice is then converted into a format that can be played back as natural-sounding speech.

[0635] Step 6:

[0636] The server uses feature points extracted from the video footage and 3D modeling software to generate a 3D avatar of the deceased. This avatar is a digital model that preserves the visual characteristics of the deceased, including their appearance.

[0637] Step 7:

[0638] The user accesses the virtual space using a terminal and begins a conversation with the deceased based on a 3D avatar and voice data provided by the server.

[0639] Step 8:

[0640] The server receives input from the user in the virtual space, searches the database, and generates natural conversations based on relevant episodes and information about the deceased.

[0641] Step 9:

[0642] The generated conversation content is provided to the user in a virtual space using reconstructed audio. The user can continue the conversation with the deceased and have an emotional experience.

[0643] (Example 1)

[0644] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0645] In modern society, there is a growing need to maintain emotional connections with deceased loved ones by recreating conversations and memories in virtual spaces. However, conventional technologies struggle to reproduce the realistic voice and appearance of the deceased, resulting in an unrealistic experience.

[0646] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0647] In this invention, the server includes means for analyzing audio and image information acquired by the user and extracting the characteristics of a person from the audio and image information; means for generating a three-dimensional representation of the person based on the extracted characteristics; and means for generating acoustic information based on the characteristics and reproducing the voice of the person. This makes it possible to have a realistic and personalized conversation with a deceased person in a virtual space.

[0648] "Audio and visual information" refers to information that includes a person's voice data and visual data, and is fundamental information for understanding biological characteristics.

[0649] "Characteristics" refer to the distinctive features of a person extracted from audio and image information, and are the basic data used to generate three-dimensional displays and acoustic information.

[0650] "Three-dimensional representation" refers to a three-dimensional visual representation of a person generated based on their characteristics, and is used as an avatar in a virtual space.

[0651] "Acoustic information" refers to voice data of a person, reproduced based on their characteristics, and is the voice used in dialogue in a virtual space.

[0652] A "virtual space" is a digital environment created using computer technology, where users access and interact with it through an interface.

[0653] "Conversation information" refers to the content of dialogue generated in a virtual space, specifically the response content generated from a set of information based on the user's input.

[0654] An "information set" refers to a collection of information, including databases and knowledge bases, used to generate dialogue in a virtual space.

[0655] This invention is a system for recreating conversations with deceased persons in a virtual space, and its embodiments are described below.

[0656] First, the user uses their device to collect audio and image information of the deceased. This process involves collecting high-quality data using smart devices or dedicated recording equipment. This data is then encrypted and uploaded to a server.

[0657] The server analyzes the received audio and image information. For audio data, speech synthesis software is used to extract the characteristics of the deceased. This uses common speech recognition services and synthesis engines. For image information, a facial recognition algorithm is used to understand the person's appearance and extract features to generate a 3D representation. At this stage, for example, OpenCV or similar image processing libraries are used.

[0658] Based on the extracted characteristics, the server generates a three-dimensional representation. This process is achieved using a three-dimensional modeling tool such as Blender. Furthermore, based on these characteristics, acoustic information is generated to recreate the deceased's voice. This allows the user to interact with the deceased both visually and aurally in a virtual space.

[0659] The virtual space is a digital environment accessible to users. Users log in to this space via a terminal and enjoy conversations using a 3D display and audio information of the deceased provided by the server. In this process, conversations are generated in real time using a generative AI model, and prompt sentences are processed in response to user input, resulting in natural and personalized conversations.

[0660] For example, if a user uses the prompt "Tell me about a past trip," the server will use this request to search for memories and related information about the deceased and generate a response such as, "That trip to the beach was really fun. I remember you collecting seashells on the sand." This response is reproduced in a voice that mimics the deceased's voice.

[0661] The distinguishing feature of this invention is that, by integrating these technologies into a system, it can provide users with emotional healing by offering a realistic and personalized virtual experience with the deceased and strengthening their emotional connection.

[0662] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0663] Step 1:

[0664] The user uses a device to collect audio and image information. The input used is video of a person speaking or a portrait image. The device temporarily stores the collected data and prepares it for transmission to a server. Specifically, it records high-resolution video using a smartphone or digital camera and records audio using a microphone.

[0665] Step 2:

[0666] The device uploads the collected audio and image information to the server. The input is the data collected in step 1, and the output is the data sent to the server. Here, data is encrypted using a secure communication protocol such as HTTPS to protect privacy. The device notifies the user when the transmission is complete.

[0667] Step 3:

[0668] The server analyzes the received audio and image data. The input is data received from the terminal, and the output is extracted characteristic data of the person. A speech recognition engine is used to analyze the audio data, extracting the pitch and tone of the voice. Image data is processed by a facial recognition algorithm to extract data on distinctive appearances. This process provides basic information to identify the characteristics of the deceased.

[0669] Step 4:

[0670] The server generates a three-dimensional representation using the extracted feature data. The input is the feature data obtained in step 3, and the output is a three-dimensional representation that can be displayed in a virtual space. Specifically, 3D modeling software is used to create an avatar based on the extracted data. In addition, acoustic information is generated, and speech synthesis is performed to reproduce the deceased person's voice.

[0671] Step 5:

[0672] The server integrates the generated 3D display and audio information into the virtual space. The input is the data generated in step 4, and the output is avatar and audio data usable within the virtual space. These are set up by virtual environment creation software and made accessible to the user. Specifically, a digital environment is created using tools such as Unity or Unreal Engine.

[0673] Step 6:

[0674] Users can access a virtual space using a device and converse with an avatar of a deceased person. As input, users send questions and comments to the virtual space using prompt text. Based on this prompt, the server searches a knowledge base for relevant information and generates a response. The output is a voice response presented as a natural conversation to the user. Specifically, using an AI model, a response such as "Tell me about your old trip" might be generated, for example, "That trip to the beach was really fun."

[0675] (Application Example 1)

[0676] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0677] Conventional virtual reality dialogue systems have struggled to foster deep emotional connections with deceased loved ones because they lack sufficient integration with the real world. Furthermore, recreating memorable experiences in places associated with the deceased within the virtual space has been difficult. Therefore, there is a need to provide a more immersive and emotionally impactful virtual experience.

[0678] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0679] In this invention, the server includes means for analyzing voice and video information of a person acquired by the user and generating a three-dimensional model; means for reproducing the person's voice based on the voice information; means for enabling dialogue with the person in a virtual reality space using the generated three-dimensional model and voice; and means for recognizing real-world scenery and providing content linked to the person in the virtual reality space. This makes it possible for the user to enjoy a deeper reminiscing experience by linking the emotional reunion with the deceased with real-world scenery.

[0680] A "user" is a person who uses the system to experience a conversation with the deceased.

[0681] "Audio and video information" refers to recorded voice and video data related to the deceased.

[0682] A "three-dimensional model" is a digital representation created to recreate the appearance of a deceased person in three dimensions.

[0683] A "sound reproduction device" is a device that includes technology for artificially generating the voice of a deceased person based on collected audio information.

[0684] A "virtual reality space" is an interactive artificial environment recreated using digital technology.

[0685] A "device for recognizing real-world landscapes" is a device that uses cameras and sensors to digitize the surrounding physical environment and acquire information.

[0686] An "information storage database" is a digital storage device used to organize and store information about the aforementioned person or related matters.

[0687] This invention is a system for facilitating dialogue with a deceased person in a virtual reality space. This system includes a terminal owned by the user, a server for data processing, and a device capable of displaying the virtual reality environment.

[0688] Users acquire audio and video information about the deceased through their devices and securely upload it to a server in the cloud. The server uses this data to generate a 3D model of the deceased. Facial recognition algorithms are used to generate the 3D model, reproducing the deceased's distinctive appearance. For audio information, speech synthesis technology is used to reproduce the deceased's voice quality and speaking style. Specifically, Amazon Polly and Google Cloud Text-to-Speech are used for speech synthesis, and AR technology is used for 3D model generation.

[0689] The server recognizes real-world landscape data and provides a responsive experience with the deceased when the user views that landscape through a virtual reality device. When the user visits a specific location, they can share experiences and memories associated with the deceased and that location in the virtual reality space. An example of a prompt message is, "The user has identified a location. Please share your memories associated with that location."

[0690] This system allows users to experience emotional reunions with deceased loved ones in a more realistic way, and to recreate memories with deep emotion. This service, utilizing virtual reality space, goes beyond mere data viewing, offering new possibilities that broaden the scope of experiences.

[0691] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0692] Step 1:

[0693] The user obtains audio and video information of the deceased through the device. The input consists of recorded audio and video data. This data is temporarily stored on the device and used for subsequent processing. Specific actions include the user recording audio or taking photos and videos using a mobile device or computer.

[0694] Step 2:

[0695] The device uploads the acquired audio and video information to a server in the cloud. The input is the audio and video data on the device, and the output is the data securely transferred to the server. This process uses data encryption and data transfer protocols (e.g., HTTPS) to ensure that the information is transmitted securely.

[0696] Step 3:

[0697] The server analyzes uploaded audio and video information to generate a three-dimensional model of the deceased. The input is audio and video data stored on the server, and the output is a three-dimensional model of the deceased. This process utilizes machine learning techniques and facial recognition algorithms to extract the deceased's physical characteristics and reproduce them as a three-dimensional model. The actual operation involves extracting feature points from video data and constructing a three-dimensional digital model based on these points.

[0698] Step 4:

[0699] The server analyzes audio information and performs speech synthesis to recreate the deceased person's voice. The input is audio data, and the output is the recreated voice of the deceased. Amazon Polly and Google Cloud Text-to-Speech are used for speech synthesis technology, which generates digital speech that recreates the deceased person's voice quality and speaking style. The specific operations are analysis of voice characteristics and generation of recreated speech data.

[0700] Step 5:

[0701] The user accesses a virtual space using a virtual reality device and initiates an interaction with a generated three-dimensional model using voice. The input is the user's visual and auditory information, and the output is the experience of interacting with the deceased's avatar. When the user moves to a specific location, the system recognizes the location and generates relevant prompts, providing content that is relevant to the deceased. Specific actions include obtaining the user's location information and initiating a conversation about memories of the deceased based on that information.

[0702] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0703] This invention combines an emotion engine with a system that generates a three-dimensional model by analyzing audio and video information of a deceased person acquired by the user, enabling interaction with the deceased person in a virtual space. This makes it possible to recognize the user's emotions in real time and provide interaction accordingly.

[0704] First, the user uses a device to record audio and video information of the deceased. This information is sent from the device to a server. The server analyzes the characteristics of the deceased's voice from the transmitted audio information and extracts the characteristics of the deceased's appearance from the video information. This generates a three-dimensional model of the deceased and a reconstructed voice.

[0705] Next, the system uses an emotion engine to recognize the user's emotions. When the user accesses the virtual space, the terminal captures the user's facial expressions and voice, and uses the emotion engine to perform real-time emotion analysis. Based on this analysis, the server adjusts the dialogue. For example, if the user has a sad expression, the server generates a response that includes comfort and empathy, and the deceased person's avatar in the virtual space responds appropriately.

[0706] As a concrete example, consider a situation where a user is overwhelmed with emotion and about to cry when reunited with a deceased loved one in a virtual space. At this point, the emotion engine detects the user's tears and sends "sadness" emotion data to the server. Based on this information, the server generates a warm message such as, "Don't cry, we can always meet here," and the avatar of the deceased person in the virtual space speaks that message to the user.

[0707] This invention provides users with an experience that allows them to connect deeply and emotionally with deceased loved ones in a virtual space, enabling flexible conversations that respond to changes in their emotions. As a result, users can experience deeper healing and satisfaction.

[0708] The following describes the processing flow.

[0709] Step 1:

[0710] The user uses the device to record audio and video of the deceased. The recording and video are done in high resolution, and the saved data is stored in the device's temporary storage.

[0711] Step 2:

[0712] The device uploads the collected audio and video data to the server. This process is carried out through an encrypted, secure communication channel to ensure the confidentiality of the data.

[0713] Step 3:

[0714] The server receives the transmitted audio data and applies a speech recognition algorithm to analyze the characteristics of the deceased's voice. From the results of this analysis, the deceased's speaking style and voice quality are digitized.

[0715] Step 4:

[0716] The server processes the video data and uses facial recognition technology to extract the deceased's physical characteristics. This process allows the server to understand the three-dimensional structure in 3D space, which is then used to construct a 3D model.

[0717] Step 5:

[0718] The server generates a three-dimensional model and a reproduced voice of the deceased. These serve as materials for users to experience visually and aurally within the virtual space.

[0719] Step 6:

[0720] When a user logs into the virtual space, the device uses the user's camera and microphone to continuously capture the user's facial expressions and voice, and sends that data to the emotion engine.

[0721] Step 7:

[0722] The emotion engine analyzes received data and evaluates the user's emotional state in real time. The analysis results identify the user's current emotion, such as sadness, happiness, or surprise.

[0723] Step 8:

[0724] The server dynamically generates dialogue content within the virtual space based on emotional data transmitted from the emotion engine. It selects and adjusts appropriate responses according to the user's emotions.

[0725] Step 9:

[0726] Within the virtual space, the conversation generated by the server is reproduced as audio and delivered to the user through the deceased person's avatar. This allows the user to naturally engage in emotional dialogue with the deceased.

[0727] (Example 2)

[0728] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0729] In modern society, there is a growing demand for new forms of interaction with deceased loved ones. However, it is difficult to retain memories of the deceased, and there is a lack of methods to recreate emotional connections. Conventional systems have limitations in the quality of recreating the deceased and the interaction, and in particular, they have the challenge of not being able to have flexible conversations that respond to the user's emotions.

[0730] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0731] In this invention, the server includes means for analyzing voice and video information of a person acquired by the user and generating a three-dimensional model of the person; means for reproducing the person's voice based on the voice information; means for enabling dialogue with the person in a virtual space using the generated three-dimensional model and the reproduced voice; and means for recognizing the user's emotions in real time and generating a corresponding dialogue. This makes it possible to provide flexible and appropriate dialogue that responds to the user's emotions while maintaining a deep emotional connection with the deceased.

[0732] "User" refers to an individual who operates the system and inputs audio and video information.

[0733] "Person" refers to an individual from whom audio and video information is acquired, and who is the subject of a three-dimensional model or a reconstructed voice.

[0734] A "three-dimensional model" refers to a three-dimensional digital representation generated based on a person's physical characteristics.

[0735] "Audio information" refers to digital audio data that includes the characteristics of a person's voice.

[0736] "Reproduced voice" refers to artificial voice data that reproduces a person's voice, generated based on the original audio information.

[0737] A "virtual space" refers to a digital environment constructed using digital technology, where three-dimensional models and reproduced sounds are placed, and users can interact with them.

[0738] "Methods for recognizing emotions in real time" refers to technologies that analyze data such as the user's facial expressions and voice to identify their emotions at that moment.

[0739] "Means of generating dialogue" refers to the process of generating appropriate responses in response to the user's emotions and input.

[0740] A "generative AI model" refers to an algorithm used to learn from data and accomplish a specific task.

[0741] A "prompt sentence" refers to an instruction sentence input into a generative AI model, which serves as a criterion for determining the content of the output dialogue.

[0742] This invention provides a virtual space that enables users to have emotional interactions with deceased loved ones. The user first records audio and video information of the deceased using a device. Suitable hardware includes common smartphones, tablets, and camera devices. This information is transmitted from the device to a server. The server receives this data and analyzes the characteristics of the deceased's voice using an audio feature extraction algorithm. It also extracts the deceased's physical characteristics using deep learning-based image processing software. This generates a three-dimensional model and synthesized voice of the deceased.

[0743] When a user accesses the virtual space, the terminal captures the user's facial expressions and voice, and analyzes the user's emotions in real time through emotion recognition software. This process requires a camera and microphone. Based on this analysis, the server uses a generative AI model to generate appropriate dialogue. An example of a prompt might be, "Generate flexible dialogue that responds to the user's emotions and convey the content in the deceased's voice."

[0744] As a concrete example, suppose a user reunites with a deceased loved one in a virtual space and displays a sad expression. At this point, emotion recognition software identifies the user's expression as "sadness" and sends that data to a server. The server then generates a warm message such as, "Don't cry, I'll always be waiting here for you." This message is conveyed to the user in the virtual space through the deceased person's avatar.

[0745] This system allows users to experience deep emotional connection with their deceased loved ones and enjoy intimate and realistic interactions that respond to their changing emotions.

[0746] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0747] Step 1:

[0748] Users record audio and video information of the deceased using their own devices. Input includes audio and video files obtained from smartphones and cameras. This information is saved in digital format and prepared for transmission to a database.

[0749] Step 2:

[0750] The terminal transmits the recorded audio and video information to the server. In this step, the information is transmitted over the internet as data packets. The input data consists of the audio and video files obtained in the previous step. The output is the encrypted information that arrives securely at the server.

[0751] Step 3:

[0752] The server analyzes the received audio information. It processes the input audio data through an audio feature extraction algorithm, outputting features such as voice intonation, tone, and speaking speed as numerical data. Specifically, an audio signal processing library is used.

[0753] Step 4:

[0754] The server analyzes the received video information. It receives video data as input and uses image processing techniques to extract features such as facial shape and posture. The output is data for a three-dimensional model based on these features. Specifically, it performs image analysis using deep learning.

[0755] Step 5:

[0756] The server generates a 3D model of the deceased and a reconstructed voice. This step uses the voice and video feature data obtained in the previous step. Based on the input data, a generating AI model is used to output a digital avatar and synthesized voice.

[0757] Step 6:

[0758] When a user accesses the virtual space, the terminal captures the user's facial expressions and voice in real time. Input consists of audio and video data obtained from the camera and microphone. Output is a data stream of the user's facial expressions and voice, which is sent to emotion recognition software.

[0759] Step 7:

[0760] The server analyzes the user's emotions in real time. It uses facial and audio data received from the terminal as input. An emotion recognition algorithm analyzes this data and outputs the user's emotional state as numerical data.

[0761] Step 8:

[0762] The server uses a generative AI model to generate dialogue that responds to the user's emotions. The prompt is set to "Generate a warm response based on the user's emotions," and the input is the emotion data obtained in the previous step. The output is the dialogue that the deceased person's avatar should say.

[0763] Step 9:

[0764] The server sends the generated dialogue to the deceased person's avatar in the virtual space, allowing it to interact with the user. The input is the generated dialogue, and the output is the audio and visual dialogue experience provided to the user. The virtual space performs the necessary visual rendering and speech synthesis to reproduce this in real time.

[0765] (Application Example 2)

[0766] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0767] Conventional virtual reality dialogue systems have struggled to provide flexible dialogue that responds to the user's emotional state. There is a growing need to generate appropriate responses based on the user's emotions to foster deeper emotional connections and improve the user experience.

[0768] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0769] In this invention, the server includes means for analyzing voice and video information of a person acquired by the user and generating a three-dimensional model of the person; means for evaluating the user's emotional state and adjusting the dialogue content based on that evaluation; and means for generating a conversation that corresponds to the user's emotions. This enables natural and personalized dialogue that is adapted to emotions.

[0770] A "user" is an individual who operates the system and interacts within the virtual space.

[0771] "Audio information" refers to digital data related to a person's speech, and is fundamental information for a system to reproduce.

[0772] "Visual information" refers to visual data about a person's appearance, which is used to generate a three-dimensional model.

[0773] A "three-dimensional model" is a digital representation of a person in three dimensions, created based on acquired video information.

[0774] "Reproduced speech" refers to the speech expressions of a person that are generated based on audio information.

[0775] A "virtual space" is a digital environment built on a computer, where users interact with each other.

[0776] "Emotional analysis tools" are technologies that evaluate the user's current emotional state and allow the system to determine a response accordingly.

[0777] A "database" is a collection of information that a system uses to store and refer to as needed.

[0778] "Machine learning technology" is a technique that uses algorithms to extract patterns from data and then uses those patterns to make predictions or generate new data.

[0779] The system for implementing this invention is based on a user terminal and a server. First, the user acquires audio and video information using a smart device, such as smart glasses or a head-mounted display. The terminal captures this information and transmits it to the server. The server generates a three-dimensional model from the video information and reconstructs the audio based on the audio information.

[0780] The server incorporates the "Microsoft Azure Cognitive Services" sentiment analysis API and the "OpenAI speech synthesis API." The user's facial expressions and voice data are analyzed in real time using these APIs. Based on the analysis results, the server generates emotionally adaptive dialogue, enabling natural conversations within the virtual space.

[0781] As a concrete example, consider a virtual store scenario. A user reunites with a deceased person, represented as an avatar, in a virtual space, and they tour their favorite places together. At this time, the system detects when the user is smiling, and the deceased person's avatar provides a message such as, "I remember the smile on your face when you chose this." An example of a prompt used here would be, "Generate a warm message when the user is smiling. For example, provide a conversation that recreates a happy moment at a memorable store."

[0782] In this way, by adjusting the interaction in real time according to the user's emotions, the system can provide a deeper emotional connection and create a more satisfying experience for the user.

[0783] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0784] Step 1:

[0785] The terminal captures audio and video information from the user and sends it to the server as digital data. The input is the user's audio and video, and the output is the transmission of this data to the server. This process prepares the basic data necessary for the virtual space.

[0786] Step 2:

[0787] The server generates a three-dimensional model from the received video information. This process utilizes computer graphics technology to create a three-dimensional model of a person. The input is the user's video data, and the output is a three-dimensional model. This generates a virtual avatar based on any given person.

[0788] Step 3:

[0789] The server analyzes voice information and performs speech synthesis based on this analysis. Machine learning techniques are used to capture voice characteristics and reproduce realistic speech. The input is the user's voice data, and the output is the synthesized speech. A generative AI model is used, enabling natural-sounding speech.

[0790] Step 4:

[0791] The server senses the user's facial expressions in real time and evaluates their emotional state using an emotion analysis API. The input is the user's facial expression data, and the output is the result of the emotion analysis. Specifically, the API quantifies the characteristics of each facial expression and determines the emotional state.

[0792] Step 5:

[0793] The server generates prompt sentences based on the sentiment analysis results and constructs an appropriate dialogue. In this step, the generating AI model determines the dialogue content based on the prompt sentences and presents it to the user. The input is data on the emotional state, and the output is a dialogue adapted to that emotion.

[0794] Step 6:

[0795] Within the virtual space, a conversation generated by the user's device is played back. The user can use a smart device to enjoy a conversation with the deceased through a generated 3D model and reproduced voice. The input is the 3D model and voice generated in the previous step, and the output is the conversation experienced by the user.

[0796] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0797] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0798] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[0799] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0800] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0801] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0802] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0803] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0804] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0805] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0806] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[0807] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[0808] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0809] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0810] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0811] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0812] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0813] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0814] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0815] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0816] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.

[0817] The following is further disclosed regarding the embodiments described above.

[0818] (Claim 1)

[0819] A means for analyzing voice and video information of a person acquired by the user and generating a three-dimensional model of the said person,

[0820] A means for reproducing the voice of the person based on the aforementioned audio information,

[0821] A means to enable dialogue with the person in a virtual space using a generated three-dimensional model and reproduced voice,

[0822] A system that includes this.

[0823] (Claim 2)

[0824] The system according to claim 1, which generates a conversation by referencing relevant information from a database based on input from a user in the virtual space.

[0825] (Claim 3)

[0826] The system according to claim 1, wherein in generating the three-dimensional model and reproducing the voice, machine learning technology is used to extract the features of the person, and a realistic appearance and voice are reproduced based on these features.

[0827] "Example 1"

[0828] (Claim 1)

[0829] A means for analyzing audio and image information acquired by the user and for extracting the characteristics of a person from the audio and image information,

[0830] A means for generating a three-dimensional representation of a person based on the extracted characteristics,

[0831] Based on the aforementioned characteristics, means for generating acoustic information and reproducing the voice of the person,

[0832] A means for enabling conversation with the person in a virtual space using the generated three-dimensional display and acoustic information,

[0833] A system that includes this.

[0834] (Claim 2)

[0835] The system according to claim 1, which, in response to input from a user in the virtual space, references relevant information from an information set and generates conversational information.

[0836] (Claim 3)

[0837] The system according to claim 1, wherein in generating the three-dimensional display and reproducing the sound, learning technology is used to extract the characteristics of the person, and based on these, a realistic appearance and sound are reproduced.

[0838] "Application Example 1"

[0839] (Claim 1)

[0840] A device that analyzes voice and video information of a person acquired by the user and generates a three-dimensional model of the person,

[0841] A device for reproducing the voice of the person based on the aforementioned audio information,

[0842] A device that enables interaction with the person in a virtual reality space using a generated three-dimensional model and reproduced voice,

[0843] A device that recognizes real-world scenery and provides content that interacts with people in the virtual reality space,

[0844] A system that includes this.

[0845] (Claim 2)

[0846] The system according to claim 1, which generates a conversation by referencing relevant information from an information storage database based on input from a user in the virtual reality space.

[0847] (Claim 3)

[0848] The system according to claim 1, wherein in generating the three-dimensional model and reproducing the voice, machine learning technology is used to extract the features of the person, and a realistic appearance and voice are reproduced based on these features.

[0849] "Example 2 of combining an emotion engine"

[0850] (Claim 1)

[0851] A means for analyzing voice and video information of a person acquired by the user and generating a three-dimensional model of the said person,

[0852] A means for reproducing the voice of the person based on the aforementioned audio information,

[0853] A means to enable dialogue with the person in a virtual space using a generated three-dimensional model and reproduced voice,

[0854] A means of recognizing the user's emotions in real time and generating corresponding dialogue,

[0855] A system that includes this.

[0856] (Claim 2)

[0857] The system according to claim 1, which, based on input from a user in the virtual space, refers to relevant information from a database and generates a conversation using a generation AI model.

[0858] (Claim 3)

[0859] The system according to claim 1, wherein in generating the three-dimensional model and reproducing the voice, machine learning technology is used to extract the features of the person, and based on these, a realistic appearance and voice are reproduced, and dialogue content is generated based on the user's emotions.

[0860] "Application example 2 when combining with an emotional engine"

[0861] (Claim 1)

[0862] A means for analyzing voice and video information of a person acquired by the user and generating a three-dimensional model of the said person,

[0863] A means for reproducing the voice of the person based on the aforementioned audio information,

[0864] A means to enable dialogue with the person in a virtual space using a generated three-dimensional model and reproduced voice,

[0865] A means of sentiment analysis for evaluating the user's emotional state and adjusting the content of the dialogue based on that evaluation,

[0866] A system that includes this.

[0867] (Claim 2)

[0868] The system according to claim 1, which, in the virtual space, references relevant information from a database based on user input and real-time emotional state, and generates a conversation that corresponds to the user's emotions.

[0869] (Claim 3)

[0870] The system according to claim 1, wherein in the generation of the three-dimensional model, the reproduction of voice, and the emotion analysis, machine learning technology is used to extract the characteristics of the person and the emotional state of the user, and based on these, a realistic appearance, voice, and emotion-adaptive dialogue are reproduced. [Explanation of Symbols]

[0871] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. A means for analyzing voice and video information of a person acquired by the user and generating a three-dimensional model of the said person, A means for reproducing the voice of the person based on the aforementioned audio information, A means for enabling dialogue with the person in a virtual space using a generated three-dimensional model and reproduced voice, A system that includes this.

2. The system according to claim 1, which generates a conversation by referring to relevant information from a database based on input from a user in the virtual space.

3. The system according to claim 1, wherein in generating the three-dimensional model and reproducing the voice, machine learning technology is used to extract the features of the person, and a realistic appearance and voice are reproduced based on these features.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A