System

The system addresses the challenge of preserving deceased memories by generating a personality model and simulating conversations, allowing users to maintain emotional connections through realistic voice and video interactions.

JP2026026904APending Publication Date: 2026-02-18SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024129325
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-08-05
Publication Date
2026-02-18

AI Technical Summary

Technical Problem

Existing technologies lack the ability to realistically preserve and communicate the memories of deceased individuals, particularly through their voice and conversation patterns, making it difficult for grieving individuals to maintain an emotional connection.

Method used

A system that collects communication history, extracts and analyzes text data using natural language processing to generate a personality model, reproduces voice, and generates video frames to simulate conversations with the deceased, ensuring ethical consent is obtained.

Benefits of technology

Enables users to experience realistic and emotional connections with the deceased by recreating their voice and conversation patterns, providing a more meaningful way to cherish memories.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026026904000001_ABST
    Figure 2026026904000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system comprising: means for collecting a communication history of a deceased; means for extracting text data from the communication history; means for analyzing the text data with a natural language processing technique to identify linguistic features of the deceased; means for generating a personality model of the deceased based on the identified linguistic features; means for reproducing speech based on the personality model; and means for generating a video frame using the reproduced speech.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] Losing a loved one brings deep grief. In particular, the reality that one will never be able to communicate with the deceased can be a heavy psychological burden. Therefore, there is a need for a means to more realistically preserve memories of the deceased and communicate as if the deceased were alive. The present invention aims to solve this problem by reproducing the deceased's voice, phrasing, and conversation patterns, thereby simulating communication with the deceased. [Means for solving the problem]

[0005] The present invention provides a system including: means for collecting a communication history of a deceased person; means for extracting text data from the communication history; means for analyzing the text data using natural language processing technology to identify the linguistic characteristics of the deceased person; means for generating a personality model of the deceased person based on the identified linguistic characteristics; means for reproducing voice based on the personality model; and means for generating video frames using the reproduced voice. This allows a user to virtually experience communication with the deceased person through video frames that make the deceased feel as if they were alive. Furthermore, the system further includes means for confirming the deceased person's consent to use the service while they were alive, and means for providing the service after confirming the consent, thereby avoiding ethical and legal issues. Furthermore, a more realistic experience can be provided by incorporating the deceased person's intonation and frequently used phrases into the reproduced voice and synchronizing the video frames with footage of the deceased person while they were alive.

[0006] "Communication history" refers to records of text, audio, and other information left by the deceased during their lifetime, including social media posts, call records, message history, search history, and so on.

[0007] "Text data" refers to data that has been extracted from the communication history and converted into the format of a character string or sentence.

[0008] "Natural language processing technology" is a technology that enables computers to understand, interpret, and generate human language, and involves analyzing text and extracting linguistic features.

[0009] "Linguistic features" refer to characteristics such as a speaker's individual phrasing, intonation, and frequently used phrases, and are elements that form speaking styles and patterns.

[0010] A "personality model" is an algorithmic model constructed to replicate the speech patterns and linguistic characteristics of a particular individual, mimicking that individual's speaking style and tone.

[0011] "Voice reproduction" means generating a voice that mimics the voice and linguistic characteristics of a specific individual based on collected and analyzed data.

[0012] A "video frame" is a sequence of images or video that is used to generate and visually display a video in conjunction with reproduced audio.

[0013] "Service Use Agreement" means obtaining the legal and ethical permissions required from a User or their agent to use a particular Service. [Brief explanation of the drawings]

[0014] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0015] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0016] First, the terms used in the following description will be explained.

[0017] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0018] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0019] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0020] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0021] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0022] [First embodiment]

[0023] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0024] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0025] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0026] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0027] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0028] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0029] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0030] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0031] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0032] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0033] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0034] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0035] This system collects and analyzes the communication history of a deceased person, and generates and recreates a personality model based on the linguistic features of the person. This system allows users to virtually experience communication with the deceased person through video frames that make the deceased appear as if they were alive.

[0036] System configuration

[0037] The system of the present invention comprises the following components:

[0038] 1. Data collection method (server)

[0039] 2. Data extraction method (server)

[0040] 3. Data analysis method (server)

[0041] 4. Personality model generation means (server)

[0042] 5. Consent confirmation means (user)

[0043] 6. Audio reproduction means (terminal)

[0044] 7. Video frame generation means (terminal)

[0045] Program processing overview

[0046] Data collection method (server)

[0047] The user provides the server with the social networking service's API key and other login information, which the server uses to collect the deceased person's communication history, including social networking posts, call records, and message history.

[0048] Example: The server uses the Twitter API to retrieve the tweet history of a specified user ID.

[0049] Data extraction method (server)

[0050] The server converts the collected data into an appropriate format and extracts it as text data.

[0051] Example: A server converts JSON-formatted tweet data into a text format.

[0052] Data analysis method (server)

[0053] The server uses natural language processing technology to analyze the text data, extract keywords and analyze writing style, and identify linguistic features.

[0054] Example: The server uses an NLP library to extract frequently used phrases and their context.

[0055] Personality model generation means (server)

[0056] The server generates a personality model based on the identified linguistic features, including the deceased's speech patterns and intonation.

[0057] Example: The server models typical response patterns during a conversation based on the analyzed features.

[0058] Consent confirmation method (user)

[0059] The user must confirm whether the deceased person has agreed to use the service before they die, and if necessary, upload a consent form to the server.

[0060] Example: A user uploads a scanned consent form through a web portal.

[0061] Audio reproduction means (terminal)

[0062] The device downloads a personality model from the server and reproduces the voice using a voice changer.

[0063] Example: The device uses voice changer software to generate the voice of a deceased person.

[0064] Video frame generation means (terminal)

[0065] The terminal generates video frames based on the reproduced audio and presents them to the user.

[0066] Example: A device uses synthetic video techniques to create video that accompanies the reproduced audio.

[0067] Specific examples

[0068] For example, if a deceased person frequently used the greeting "Good morning," the server can identify that phrase and generate a video message saying "Good morning" through a voice changer on the device. Looking at the device screen, the user can feel as if the deceased person was greeting them in person.

[0069] In this way, the system of the present invention can generate and provide users with realistic interactive video frames that incorporate the memories and characteristics of the deceased, thereby helping to ease the grief of losing a loved one.

[0070] The processing flow will be explained below.

[0071] Step 1:

[0072] The user provides the server with the API key and login information of the SNS. Specifically, the user enters and submits the API key of Twitter or other SNS through a web form.

[0073] Step 2:

[0074] The server uses the provided API key to collect the post data of the deceased person from the SNS. Specifically, the server calls the Twitter API to obtain the tweet history of the specified user ID.

[0075] Step 3:

[0076] The server converts the collected data into an appropriate format and extracts it as text data. Specifically, the server converts the JSON-formatted tweet data into text format and extracts each sentence.

[0077] Step 4:

[0078] The server analyzes the extracted text data using natural language processing techniques. Specifically, the server uses an NLP library to extract frequently used phrases and keywords from the text.

[0079] Step 5:

[0080] The server then uses the text data analyzed in the previous step to identify the deceased's linguistic characteristics, such as writing style, tone, specific phrases, and intonation patterns.

[0081] Step 6:

[0082] Based on the identified features, the server generates a personality model that includes the deceased's speech patterns and intonation. Specifically, the analysis results are used to create a model that incorporates the deceased's typical speech patterns and reactions.

[0083] Step 7:

[0084] The user uploads a document to the server to confirm that they have agreed to use the service. Specifically, the user uploads a scanned copy of the consent form via a web portal.

[0085] Step 8:

[0086] The device downloads the generated personality model from the server and reproduces the voice using voice changer technology. Specifically, the device obtains the personality model via the Internet and reproduces it using voice synthesis software.

[0087] Step 9:

[0088] The device generates video frames based on the reproduced audio. Specifically, it combines the synthesized audio with footage of the deceased (previous video clips or photos) to create natural-sounding video frames.

[0089] Step 10:

[0090] The terminal provides the generated video frames to the user, specifically, displays the video frames on the terminal screen so that the user can view them.

[0091] Example 1

[0092] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0093] In modern society, there is a lack of ways to continuously remember the deceased, and it is particularly difficult to remember them through conversation or audio. Furthermore, there is no technology that can reproduce the specific linguistic characteristics and mannerisms of the deceased and imitate their personality, making it impossible to continue interacting with them. Therefore, there is a need for a way for users to experience memories of the deceased more realistically and maintain an emotional connection.

[0094] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0095] In this invention, the server includes means for collecting communication data of a user, means for extracting text data from the communication data, means for analyzing the text data using natural language processing technology and identifying linguistic characteristics of the user, means for generating a character model of the user based on the identified linguistic characteristics, means for reproducing audio based on the character model, and means for generating video frames using the reproduced audio. This allows the user to experience realistic conversations and audio with the deceased, enabling them to maintain an emotional connection with the deceased.

[0096] "User" refers to an individual or group of people who use the System.

[0097] "Communication data" refers to historical information about a user's digital communications, such as social media posts, call records, and message history.

[0098] "Text data" refers to text information extracted from communication data.

[0099] "Natural language processing technology" refers to computer technology for analyzing text data, extracting keywords, analyzing writing style, etc.

[0100] "Linguistic features" refer to the phrases, expressions, style, and other characteristics that users frequently use in a particular language.

[0101] "Person model" refers to a digital model that is generated based on a user's linguistic characteristics and mimics the user's speech patterns and intonation.

[0102] "Means for reproducing voice" refers to technologies and tools for generating voice that reproduces the characteristics of a user using a person model.

[0103] "Means for generating video frames" refers to techniques and tools for generating video frames in accordance with the reproduced audio.

[0104] "Service Use Agreement" refers to an agreement indicating that a user gives permission to use the system before death.

[0105] "Pronunciation characteristics" refers to the characteristics of a user's voice, including intonation, accent, frequently used expressions, etc.

[0106] "Frequent expressions" refer to phrases and expressions that users use frequently.

[0107] "Images" refers to your visual content, such as still images and videos.

[0108] "Synchronization" refers to the technique of playing audio and video in sync.

[0109] The present invention is a system for collecting, analyzing, and reproducing the communication history of a deceased person. Detailed procedures for implementing the present invention will be described below.

[0110] System configuration

[0111] The system of the present invention comprises the following components:

[0112] 1. Data collection means (server): Collects user communication data.

[0113] 2. Data extraction means (server): Extracts text data from the collected communication data.

[0114] 3. Data analysis means (server): Using natural language processing technology, the text data is analyzed and the user's linguistic characteristics are identified.

[0115] 4. Personality model generation means (server): Generates a personality model of the user based on the identified linguistic features.

[0116] 5. Consent confirmation means (user): Confirm whether the user agreed to use the service before death.

[0117] 6. Voice reproduction means (terminal): Reproduces the voice using the personality model downloaded from the server.

[0118] 7. Video frame generation means (terminal): Generates video frames based on the reproduced audio.

[0119] Hardware and software used

[0120] Servers: High-performance servers and cloud computing platforms (e.g., AWS, Google Cloud) are used.

[0121] Natural Language Processing library: SpaCy, NLTK, or other NLP library.

[0122] Generative AI models: Advanced AI models such as GPT-3.

[0123] Voice changer software: DeepTalk, VoxCeleb2, etc.

[0124] Facial animation technology: Synthetic video techniques such as DeepFake.

[0125] Specific step-by-step instructions

[0126] 1. Data Collection

[0127] The user provides the server with permission to access the deceased person's social media accounts and call history.

[0128] The server uses the provided API key and login information to collect communication history data from the specified account.

[0129] 2. Data extraction

[0130] The server converts the collected data into a text format, for example, extracting the message content from JSON-formatted tweet data into plain text.

[0131] 3. Data Analysis

[0132] The server uses natural language processing techniques to analyze the text data and extract keywords, specifically by identifying frequently occurring phrases and their contexts using libraries such as SpaCy and NLTK.

[0133] 4. Personality Model Generation

[0134] The server uses the identified linguistic features to generate a person model that includes the deceased person's speech patterns and intonation, using a generative AI model such as GPT-3.

[0135] 5. Confirmation of consent

[0136] The user confirms the deceased person's consent to use the service during their lifetime, and if necessary, uploads a scanned copy of the consent form to the server.

[0137] 6. Audio Reproduction

[0138] The device uses a personality model downloaded from the server and reproduces the voice using voice changer software (e.g., DeepTalk).

[0139] 7. Video Frame Generation

[0140] The device generates video frames based on the reproduced audio using facial animation technology (e.g., DeepFake).

[0141] Specific examples

[0142] For example, if the deceased frequently used the greeting "Good morning," the server can identify that expression and the device can use a voice changer to generate a video message saying "Good morning." The user can feel as if the deceased is greeting them directly through the device screen.

[0143] Prompt Sentence Examples

[0144] "Collect all tweets from user ID 'example_user'."

[0145] "Extract tweet content in text format from JSON format data."

[0146] "Extract frequently occurring words and phrases from text data."

[0147] "Use the extracted features to create a model that reproduces the speech patterns of the deceased."

[0148] "Please scan the consent form and upload it through the web portal."

[0149] "Load the model data and recreate the voice using voice changer software."

[0150] "Create video frames with facial animation technology based on the generated audio."

[0151] Although the embodiment for carrying out the present invention has been specifically described above, the present invention is not limited to this, and other embodiments and modifications are possible.

[0152] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0153] Step 1: Data collection

[0154] Users provide their social networking service API keys and login information, which the server uses to collect communication data.

[0155] Input: User's API key or login information

[0156] Output: Social media posts, call history, message history data

[0157] Specific operation: The server uses the Twitter API to retrieve the tweet history of the specified user ID in JSON format. As an example, enter the prompt "Please collect all tweets from user ID 'example_user'."

[0158] Step 2: Data extraction

[0159] The server extracts text data from the collected data.

[0160] Input: Collected social media posts, call history, and message history data

[0161] Output: Text data

[0162] Specific operation: The server extracts the text message content from the JSON format data and converts it to plain text. Specifically, it inputs the prompt "Please extract the tweet content in text format from the JSON format data."

[0163] Step 3: Data analysis

[0164] The server uses natural language processing technology to analyze the text data, extract keywords, and analyze frequently occurring phrases and writing styles.

[0165] Input: Extracted text data

[0166] Output: Linguistic features (keywords, frequent phrases, contextual information)

[0167] Specific behavior: The server uses an NLP library (e.g., SpaCy or NLTK) to extract frequent words and phrases from the text and identify the context. The server inputs the prompt: "Extract frequent words and phrases from the text data."

[0168] Step 4: Personality model generation

[0169] The server generates a personality model based on the identified linguistic features, including the deceased's speech patterns and intonation.

[0170] Input: Identified linguistic features

[0171] Output: Personality model

[0172] Specific operation: The server uses a generative AI model (e.g., GPT-3) to create a model that mimics the deceased person's speech patterns. It inputs the prompt, "Please use the extracted features to create a model that reproduces the deceased person's speech patterns."

[0173] Step 5: Confirm consent

[0174] The user checks whether the deceased person agreed to use the service before they died, and uploads a consent form to the server if necessary.

[0175] Input: Consent form (possibly scanned)

[0176] Output: Service use authorization

[0177] Specific operation: The user scans the consent form and uploads it to the server through the web portal. The server verifies and records the uploaded consent form. It then instructs the user to "scan the consent form and upload it through the web portal."

[0178] Step 6: Audio Reproduction

[0179] The device downloads a personality model from the server and reproduces the voice using voice changer software.

[0180] Input: Personality model data

[0181] Output: Reproduced audio

[0182] Specific operation: The device loads the model data downloaded from the server and generates voice using voice changer software such as DeepTalk or VoxCeleb2. The device instructs the user to "load the model data and reproduce the voice using voice changer software."

[0183] Step 7: Video Frame Generation

[0184] The terminal generates video frames in accordance with the reproduced audio and presents them to the user.

[0185] Input: Reproduced audio

[0186] Output: Video frame

[0187] Specific operation: The device uses facial animation technology (e.g., DeepFake) to create a video frame that corresponds to the reproduced audio. It then prompts the device to "create a video frame using facial animation technology based on the generated audio."

[0188] (Application example 1)

[0189] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0190] When a loved one passes away, there are limited ways to cherish memories of the deceased. Conventional technology lacks the means to reproduce the words and conversations of the deceased, in addition to visual records such as photographs and videos, making it difficult to recreate the emotional connection with the deceased. Furthermore, there is a demand for providing a natural conversation experience that includes the unique phrases and intonations used by the deceased, but no concrete means for achieving this have been established.

[0191] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0192] In this invention, the server includes means for collecting the communication history of the deceased, means for extracting text data from the communication history, and means for analyzing the text data using natural language processing technology to identify the linguistic characteristics of the deceased. This allows a personality model of the deceased to be generated based on the analyzed data, enabling the user to experience a simulated conversation with the deceased. Furthermore, the content of the conversation provided based on the generated personality model is more realistic, allowing the user to feel a deeper emotional connection with the deceased. This allows for a more realistic reproduction of memories with the deceased.

[0193] "Communication history" is a general term for records such as messages, call records, and social media posts sent and received through electronic communication means used by the deceased during their lifetime.

[0194] "Text data" refers to character information and sentences extracted from the communication history.

[0195] "Natural language processing technology" is a computer processing technology for understanding, analyzing, and generating human language, and includes a wide range of algorithms and methods for mechanically processing natural language.

[0196] "Linguistic features" refer to the characteristics of language use possessed by a particular writer or speaker, such as frequent phrases, lexical choice, and stylistic patterns.

[0197] A "personality model" is a computer model constructed to recreate the language usage and speech patterns of a deceased person based on their linguistic characteristics.

[0198] "Voice reproduction means" refers to techniques or devices that use a personality model to reproduce the voice, intonation, and distinctive speaking style of the deceased.

[0199] A "video frame" is an individual image that reproduces a specific moment in time as a video, and is the basic unit for expressing movement as a video when displayed continuously.

[0200] "Pseudo-dialogue" refers to providing a user with an experience that makes them feel as if they are conversing with the deceased person using the generated personality model.

[0201] "Dialogue content" refers to the content of words and messages exchanged between the user and the personality model of the deceased person in a simulated dialogue.

[0202] This invention is a system that collects and analyzes the communication history of a deceased person, and generates and recreates a personality model based on the linguistic features of the communication history. This system allows users to enjoy a virtual conversation experience in which the deceased person feels as if they are alive. This system incorporates key elements such as "communication history," "text data," "natural language processing technology," "linguistic features," "personality model," "means for reproducing voice," "video frames," "simulated dialogue," and "dialogue content."

[0203] Hardware and software used

[0204] Hardware: Smartphone (iPhone or Android)

[0205] Software: Python, Flask, TensorFlow / Keras, Pandas, NLTK, Twilio API, Google Cloud Text-to-Speech API

[0206] Program processing overview

[0207] 1. Data collection method (server)

[0208] The server receives the API key of the deceased person's social media account from the user. Using this API key, the server accesses the deceased person's social media account and collects their communication history, including Twitter and Facebook posts, messages, and call records.

[0209] 2. Data extraction method (server)

[0210] The server converts the collected social media data into an appropriate format and extracts it as text data. The collected data is provided in JSON format, but it is converted into an easily handled text format.

[0211] 3. Data analysis method (server)

[0212] The server analyzes the extracted text data using natural language processing libraries such as Pandas and NLTK to identify the linguistic characteristics of the deceased and extract frequent phrases, expressions, and writing style.

[0213] 4. Personality model generation means (server)

[0214] Based on the analyzed linguistic features, the server uses TensorFlow and Keras to generate a personality model of the deceased, a computer model that reproduces the speech patterns and intonation of the deceased.

[0215] 5. Consent confirmation means (user)

[0216] Users must verify through a web portal whether the deceased person consented to use the service before they died, and then upload a scanned copy of the consent form.

[0217] 6. Audio reproduction means (terminal)

[0218] The device uses a personality model downloaded from the server to reproduce the voice of the deceased person using voice synthesis technology such as the Google Cloud Text-to-Speech API.

[0219] 7. Video frame generation means (terminal)

[0220] As the audio plays, the device uses synthetic video technology to generate a video of the deceased, providing the user with realistic video frames synchronized with the audio.

[0221] Specific examples

[0222] For example, if a deceased person frequently used the greeting "Good morning," the server can identify this phrase and generate a video message in which the device says "Good morning" through a voice changer. The following is an example of a prompt sentence:

[0223] Example prompt sentence:

[0224] Prompt: "Good morning. What are you planning to do today?"

[0225] Response: The generated personality model responds, "Good morning! I'm planning to go for a walk today. How about you?"

[0226] In this way, the user can experience memories of the deceased more realistically through simulated dialogue.

[0227] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0228] Step 1:

[0229] Data collection method (server)

[0230] The server receives the API key and login information of the deceased person's SNS from the user. The server then collects the deceased person's communication history (messages, posts, call records, etc.) from various SNS platforms (e.g., Twitter, Facebook). The input data is the API key and user ID, and the output is communication history data in JSON format.

[0231] Step 2:

[0232] Data extraction method (server)

[0233] The server converts the collected SNS data from an appropriate format (for example, JSON format) into text data. Here, the tweet text and message content are extracted and organized as text data. The input is communication history data in JSON format, and the output is formatted text data. Specifically, the server extracts the necessary information from the data obtained from each SNS platform and converts it into text format.

[0234] Step 3:

[0235] Data analysis method (server)

[0236] The server analyzes the text data using natural language processing techniques (e.g., Pandas or NLTK). The analysis identifies the linguistic characteristics of the deceased (frequent phrases, writing style, intonation, etc.). The input is text data, and the output is metadata about the linguistic characteristics of the deceased. Specific operations include word frequency analysis, morphological analysis, and contextual analysis.

[0237] Step 4:

[0238] Personality model generation means (server)

[0239] Based on the analysis results, the server uses TensorFlow and Keras to generate a personality model of the deceased. This is a machine learning model that learns the deceased's speech patterns. The input is metadata about the linguistic features of the deceased, and the output is a trained personality model. Specifically, it trains an LSTM model using the deceased's speech data.

[0240] Step 5:

[0241] Consent confirmation method (user)

[0242] The user checks through the web portal whether the deceased person had consented to use the service before they died, and then uploads a scanned copy of the consent form to the server. The input is the scanned consent form (electronic file), and the output is the result of uploading the consent form to the server. Specifically, the user selects the consent form on the portal site and presses the upload button.

[0243] Step 6:

[0244] Audio reproduction means (terminal)

[0245] The device uses the personality model downloaded from the server and generates the voice of the deceased person using the Google Cloud Text-to-Speech API, etc. The input is the personality model and text data, and the output is an audio file. Specifically, it converts the text-generated dialogue into audio.

[0246] Step 7:

[0247] Video frame generation means (terminal)

[0248] The device uses synthetic video technology to generate a video of the deceased person in sync with the audio. The input is the generated audio and video data of the deceased person while they were alive, and the output is video frames synchronized with the audio. Specifically, the device generates and edits a video sequence corresponding to the audio.

[0249] This allows the user to experience a simulated conversation with the deceased, allowing memories of the deceased to be recreated more realistically.

[0250] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0251] This system collects and analyzes the communication history of a deceased person, and generates and recreates a personality model based on the linguistic features. This system also incorporates an emotion engine that recognizes the user's emotions, providing dialogue that responds to the user's emotions, thereby realizing a more realistic communication experience.

[0252] System configuration

[0253] The system of the present invention comprises the following components:

[0254] 1. Data collection method (server)

[0255] 2. Data extraction method (server)

[0256] 3. Data analysis method (server)

[0257] 4. Personality model generation means (server)

[0258] 5. Consent confirmation means (user)

[0259] 6. Audio reproduction means (terminal)

[0260] 7. Video frame generation means (terminal)

[0261] 8. Emotion Recognition Means (Emotion Engine)

[0262] Program processing overview

[0263] Data collection method (server)

[0264] The user provides the server with the API key and login information of the social networking site, which the server uses to collect the deceased person's communication history, including social networking site posts, call records, and message history.

[0265] Example: The server uses the Twitter API to retrieve the tweet history of a specified user ID.

[0266] Data extraction method (server)

[0267] The server converts the collected data into an appropriate format and extracts it as text data.

[0268] Example: The server converts JSON-formatted tweet data into text format and extracts each sentence.

[0269] Data analysis method (server)

[0270] The server uses natural language processing technology to analyze the text data, extract keywords and analyze writing style, and identify linguistic features.

[0271] Example: A server uses an NLP library to extract frequently used phrases and keywords from text.

[0272] Personality model generation means (server)

[0273] The server generates a personality model based on the identified linguistic features, including the deceased's speech patterns and intonation.

[0274] Example: The server models typical response patterns during a conversation based on the analyzed features.

[0275] Consent confirmation method (user)

[0276] A document is uploaded to the server to confirm that the user has agreed to use the service.

[0277] Example: A user uploads a scanned consent form through a web portal.

[0278] Audio reproduction means (terminal)

[0279] The device downloads the generated personality model from the server and reproduces the voice using voice changer technology.

[0280] Example: The device uses voice changer software to generate the voice of a deceased person.

[0281] Video frame generation means (terminal)

[0282] The terminal generates video frames based on the reproduced audio and presents them to the user.

[0283] Example: A device uses synthetic video techniques to create video that accompanies the reproduced audio.

[0284] Emotion recognition means (emotion engine)

[0285] The emotion engine analyzes real-time data obtained from the user's camera and microphone to identify the user's emotional state.

[0286] Example: An emotion engine analyzes a user's facial expressions and tone of voice to determine whether they are happy, sad, or surprised.

[0287] How the Emotion Engine Works

[0288] The emotion engine provides a means to recognize the user's emotional state in real time and adjust the audio and video frames of the deceased based on that. For example, if the emotion engine determines that the user is sad, it can generate a soothing voice or comforting message for the deceased. This functionality allows the user to feel more connected to the deceased.

[0289] Specific examples

[0290] For example, if the deceased frequently used the greeting "good morning," the emotion engine may determine that the user is feeling unwell. The emotion engine will adjust the reproduced voice to be more gentle and comforting depending on the situation, and change the video frame accordingly. Looking at the device screen, the user will feel as if the deceased is gently encouraging them.

[0291] In this way, the system of the present invention can generate and provide users with realistic interactive video frames that incorporate the memories and characteristics of the deceased, enabling personalized interactions based on the user's emotional state, helping to alleviate the grief of losing a loved one.

[0292] The processing flow will be explained below.

[0293] Step 1:

[0294] The user provides the server with the API key and login information of the SNS. Specifically, the user enters the API key of Twitter or other SNS through a web form and submits it.

[0295] Step 2:

[0296] The server uses the provided API key to collect the post data of the deceased person from the SNS. Specifically, the server calls the Twitter API to obtain the tweet history of the specified user ID.

[0297] Step 3:

[0298] The server converts the collected data into an appropriate format and extracts it as text data. Specifically, the server converts the JSON-formatted tweet data into text format and extracts each sentence.

[0299] Step 4:

[0300] The server analyzes the extracted text data using natural language processing techniques. Specifically, the server uses an NLP library to extract frequently used phrases and keywords from the text.

[0301] Step 5:

[0302] The server then uses the text data analyzed in the previous step to identify the deceased's linguistic characteristics, such as writing style, tone, specific phrases, and intonation patterns.

[0303] Step 6:

[0304] Based on the identified characteristics, the server generates a personality model that includes the deceased's speech patterns and intonation. Specifically, it uses the analysis results to incorporate the deceased's frequently used phrases and reaction patterns into the model.

[0305] Step 7:

[0306] The user uploads a document to the server to confirm that they have agreed to use the service. Specifically, the user uploads a scanned copy of the consent form via a web portal.

[0307] Step 8:

[0308] The device downloads the generated personality model from the server and uses voice changer technology to recreate the voice of the deceased. Specifically, the device obtains the personality model via the internet and uses voice synthesis software to generate the voice of the deceased.

[0309] Step 9:

[0310] The device generates video frames based on the reproduced audio. Specifically, it combines the synthesized audio with footage of the deceased (previous video clips or photos) to create natural-sounding video frames.

[0311] Step 10:

[0312] The terminal displays the video frames to the user, specifically, displays the video frames on the terminal's display so that the user can view them.

[0313] Step 11:

[0314] The emotion engine captures real-time data from the user's camera and microphone and analyzes their emotions by analyzing their facial expressions, tone of voice, and word choice to determine their current emotional state.

[0315] Step 12:

[0316] Based on the analysis, the emotion engine adjusts the content of the deceased person's audio and video frames. For example, if the user is sad, the message of the deceased will be changed to be more comforting.

[0317] Step 13:

[0318] The device plays the adjusted content and displays it to the user. Specifically, the content generated based on instructions from the emotion engine is displayed on the display, and the user watches and listens to it.

[0319] Example 2

[0320] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0321] Although there are systems that can recreate the memories and characteristics of the deceased and allow users to have a realistic conversation with them, these systems lack the ability to recognize the user's emotional state and dynamically adjust the conversation content based on that emotion. This limits the conversational experiences provided, making it difficult for users to feel a deeper connection with the deceased.

[0322] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0323] In this invention, the server includes means for collecting the communication history of the deceased, means for extracting text data from the communication history, means for analyzing the text data using natural language processing technology and identifying linguistic characteristics of the deceased, means for generating a personality model of the deceased based on the identified linguistic characteristics, means for downloading the personality model of the deceased from the server to a terminal, means for reproducing voice based on the personality model, means for generating video frames based on the reproduced voice, means for recognizing the emotional state of the user, and means for adjusting the voice and video frames based on the emotional state of the user. This allows the user to not only realistically experience memories and characteristics of the deceased, but also to have a personalized interaction experience according to emotions.

[0324] "Communication history" refers to a series of communication data such as social media posts, call records, and message history made by the deceased person while they were alive.

[0325] "Text data" refers to data in sentence format extracted from the communication history.

[0326] "Natural language processing technology" refers to technology used by computers to understand, analyze, and generate human language.

[0327] "Linguistic characteristics" refer to linguistic characteristics and patterns, such as the particular words, writing style, and frequently used phrases used by a particular person.

[0328] A "personality model" is a data model designed to reproduce the linguistic characteristics, speech patterns, intonation, etc. of a deceased person.

[0329] "Voice reproduction" refers to the technique or process of recreating the voice of a deceased person based on a generated personality model.

[0330] A "video frame" is a frame of video that corresponds to the reproduced audio and shows the image and gestures of the deceased.

[0331] "Emotional state" refers to a user's current psychological and sensory state, which is typically determined through facial expressions, tone of voice, etc.

[0332] "Emotion recognition" refers to the technology of analyzing a user's emotional state and thereby identifying a specific emotion.

[0333] A "personalized interaction experience" is an interaction that is optimized for a user's unique characteristics and situation, and is adjusted based on their emotional state.

[0334] This system collects and analyzes the communication history of a deceased person, and generates and recreates a personality model based on the linguistic features. This system also incorporates an emotion engine that recognizes the user's emotions, providing dialogue that responds to the user's emotions, thereby realizing a more realistic communication experience.

[0335] System configuration

[0336] The system of the present invention comprises the following components:

[0337] 1. Data collection method (server)

[0338] 2. Data extraction method (server)

[0339] 3. Data analysis method (server)

[0340] 4. Personality model generation means (server)

[0341] 5. Consent confirmation means (user)

[0342] 6. Audio reproduction means (terminal)

[0343] 7. Video frame generation means (terminal)

[0344] 8. Emotion Recognition Means (Emotion Engine)

[0345] Detailed system configuration and processing method

[0346] Data collection method (server)

[0347] The user enters their social networking service API key and login information, which is then sent in a secure format to the server, which then uses this information to collect the deceased person's communication history, including social networking posts, call logs, and message history. This process uses standard APIs, such as the Twitter API and Facebook API.

[0348] Example: The server uses the Twitter API to retrieve the tweet history of a specified user ID.

[0349] Data extraction method (server)

[0350] The server analyzes the collected data and extracts it as text data. At this time, it converts the necessary parts of the data in JSON or XML format into text format and extracts each sentence and word.

[0351] Example: The server converts JSON-formatted tweet data into text format and extracts each sentence.

[0352] Data analysis method (server)

[0353] The server uses natural language processing (NLP) techniques to analyze the extracted text data. During the analysis, it identifies linguistic features such as keywords, frequent phrases, and writing style. This process uses NLP libraries (e.g., NLTK, spacy).

[0354] Example: The server uses an NLP library to extract frequently used phrases and keywords from the text.

[0355] Personality model generation means (server)

[0356] The server generates a personality model based on the analyzed linguistic features, including the deceased's speech patterns and intonation, which is then used to simulate conversations similar to those of the deceased.

[0357] Example: The server uses the analyzed features to model typical response patterns during a conversation.

[0358] Consent confirmation method (user)

[0359] The user uploads a consent form for using the service to the server in PDF format, etc. The server then checks the uploaded consent form and determines whether or not the use of the service is permitted.

[0360] Example: User uploads scanned consent form through web portal.

[0361] Audio reproduction means (terminal)

[0362] The device uses a personality model downloaded from the server and uses voice changer technology (e.g., Voicemod, iZotope) to recreate the voice of the deceased person, which is used during interactions with the user.

[0363] Example: The device uses voice changer software to generate the voice of the deceased person.

[0364] Video frame generation means (terminal)

[0365] Based on the reproduced audio, the device uses synthesis technology (e.g., D-ID, Reallusion) to generate video frames that display the deceased person's face and gestures in sync with the audio.

[0366] Example: The device uses synthetic video techniques to create a video that matches the reproduced audio.

[0367] Emotion recognition means (emotion engine)

[0368] The emotion engine analyzes real-time data captured from the user's camera and microphone. It uses facial expression analysis software (e.g., Affectiva, Microsoft Azure Face API) and voice analysis tools to identify the user's emotional state.

[0369] Example: An emotion engine analyzes a user's facial expressions and tone of voice to determine whether they are sad, happy, or surprised.

[0370] Examples of concrete examples and prompts

[0371] For example, if the deceased frequently used the greeting "Good morning," the emotion engine may determine that the user's emotional state was poor. In this case, the emotion engine will adjust the reproduced voice to be calm and comforting, and change the video frame accordingly. Looking at the device screen, the user will feel as if the deceased is gently encouraging them.

[0372] Example prompt sentence:

[0373] It uses phrases that allow users to start a natural conversation, such as "Hi, how are you doing?" or "How was your day?"

[0374] In this way, the system of the present invention can realistically recreate the memories and characteristics of the deceased and provide a personalized interaction experience for the user.

[0375] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0376] Step 1: Enter user information

[0377] The user enters the SNS API key and login information, which is then sent to the server in a secure format. Specifically, the user enters the information on the system's login screen and presses the send button. The input data is the SNS API key and login information, and the output is the API key and login information sent to the server.

[0378] Step 2: Data collection

[0379] The server uses the API key and login information sent by the user to collect the deceased person's communication history, such as social media posts, call records, and message history. Specifically, the server uses an appropriate API (e.g., Twitter API, Facebook API) to obtain data for the specified user ID. The input data is the API key and login information provided by the user in the previous step, and the output is the collected communication history (e.g., social media post data, call records).

[0380] Step 3: Data extraction

[0381] The server converts the collected communication history into an appropriate format and extracts it as text data. Specifically, the server converts SNS data in JSON or XML format into text format and extracts each sentence. The input data is the communication history (e.g., tweet data in JSON format), and the output is the extracted text data (e.g., single sentences of text).

[0382] Step 4: Data analysis

[0383] The server uses natural language processing (NLP) techniques to analyze the extracted text data. During the analysis process, it identifies linguistic features such as keywords, frequently occurring phrases, and writing style. Specifically, the server uses an NLP library (e.g., NLTK, spacy) to extract frequently used phrases and keywords from the text. The input data is the text data, and the output is the analyzed linguistic features (e.g., a list of keywords and phrases).

[0384] Step 5: Personality model generation

[0385] The server generates a personality model based on the analyzed linguistic features, including the deceased's speech patterns and intonation. Specifically, the server uses a machine learning algorithm to model typical response patterns during conversations based on the analyzed features. The input data are the linguistic features, and the output is the generated personality model.

[0386] Step 6: Confirm consent

[0387] The user uploads the consent form for using the service to the server in PDF format or other format. Specifically, the user scans the consent form and uploads it through a web portal. The input data is the consent form in PDF format, and the output is a consent confirmation flag on the server.

[0388] Step 7: Audio Reproduction

[0389] The device uses a personality model downloaded from the server and voice changer technology (e.g., Voicemod, iZotope) to recreate the deceased's voice. Specifically, the device converts the deceased's voice pattern into audio data based on the downloaded personality model. The input data is the personality model, and the output is the recreated audio data.

[0390] Step 8: Video Frame Generation

[0391] The device generates video frames based on the reproduced audio. Specifically, the device uses synthesis technology (e.g., D-ID, Reallusion) to create video that matches the reproduced audio. The input data is the reproduced audio data, and the output is the generated video frames.

[0392] Step 9: Emotion Recognition

[0393] The emotion engine analyzes real-time data acquired from the user's camera and microphone. Specifically, the emotion engine uses facial expression analysis software (e.g., Affectiva, Microsoft Azure Face API) and voice analysis tools to identify the user's emotional state. The input data is the real-time data acquired from the camera and microphone, and the output is the identified user's emotional state.

[0394] Step 10: Dialogue Adjustment

[0395] The emotion engine appropriately adjusts the content and tone of the dialogue based on the emotion data acquired in the previous stage. For example, if the user is sad, the emotion engine generates a calm voice and comforting words, and optimizes the video frames accordingly. The input data is the user's emotional state, and the output is the adjusted voice data and video frames.

[0396] (Application example 2)

[0397] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0398] There is a need for a means to maintain the emotional connection between the deceased and those left behind, particularly to ease the grief of losing the deceased. However, conventional technologies only reproduce the personality of the deceased, and are unable to engage in dialogue that reflects the user's emotional state, making it difficult to provide a realistic communication experience. Furthermore, there is no system in physical stores that allows users to feel emotionally reassured, making it difficult to alleviate the user's psychological burden.

[0399] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0400] In this invention, the server includes means for collecting the communication history of the deceased, means for extracting text data from the communication history, means for analyzing the text data using natural language processing technology to identify the linguistic characteristics of the deceased, means for generating a personality model of the deceased based on the identified linguistic characteristics, means for reproducing voice based on the personality model, means for generating video frames using the reproduced voice, means for recognizing the emotional state of a user in real time, means for adjusting the content of the dialogue in accordance with the recognized emotional state, and means for reproducing a dialogue with the deceased via a tablet terminal or smartphone in a physical store. This allows for a realistic reproduction of the personality of the deceased while enabling personalized dialogue in accordance with the user's emotional state, thereby providing users with a sense of emotional security even in a physical store.

[0401] "Communication history" is a record of past interactions recorded as digital data, such as text messages, call logs, and social media posts, across the mediums used by individuals to communicate.

[0402] "Natural language processing technology" is a field of artificial intelligence that enables computers to understand, interpret, and generate human language, and is a technology that analyzes text and extracts linguistic features.

[0403] A "personality model" is a digital recreation of a specific individual's linguistic characteristics, speech patterns, and presence, constructed from the deceased's text data and call records.

[0404] "Voice reproduction" is a technology that generates voice by imitating the speaking style and voice quality of a deceased person based on a generated personality model.

[0405] A "video frame" is a video frame generated in accordance with the reproduced audio, which digitally reproduces the image of the deceased.

[0406] "Emotional state" refers to the emotions the user is currently feeling, such as joy, sadness, surprise, etc., and is measured in real time using a camera, microphone, etc.

[0407] A "tablet device" is a portable, flat-shaped electronic device that is equipped with a touch screen and supports the operation of applications.

[0408] A "smartphone" is a telephone terminal that has advanced computer functions in addition to calling functions, and can be used to realize a variety of functions by installing applications.

[0409] "Personalized dialogue" refers to conversation content that is customized based on the user's individual emotional state and behavior, and is dynamically generated by a personality model of the deceased person.

[0410] A "physical store" is a commercial facility that exists physically and provides a place for consumers to visit and receive goods or services.

[0411] The system that realizes this application example collects and analyzes the communication history of the deceased person to generate a personality model and provide dialogue that corresponds to the user's emotional state. This system is realized using the following hardware and software components.

[0412] System Components

[0413] 1. Data collection method (server)

[0414] The server uses the SNS API key and login information provided by the user to collect the deceased person's communication history, such as SNS posts, call records, and message history. For example, the server obtains data using the Twitter API or Facebook Graph API.

[0415] 2. Data extraction method (server)

[0416] The server converts the collected data into an appropriate format and extracts it as text data. In this step, the JSON format data is converted into text format and each sentence is extracted.

[0417] 3. Data analysis method (server)

[0418] The server analyzes the text data using natural language processing techniques, extracts keywords and analyzes writing style, and identifies the linguistic characteristics of the deceased. Specifically, it uses Python NLP libraries (e.g., spaCy and NLTK).

[0419] 4. Personality model generation means (server)

[0420] The server generates a personality model of the deceased person, including their speech patterns and intonation, based on the identified linguistic features, using OpenAI's GPT-based generative AI model.

[0421] 5. Audio reproduction means (terminal)

[0422] The device uses a personality model downloaded from a server to reproduce the voice using voice changer technology (e.g., Respeecher or Google's Tacotron).

[0423] 6. Video frame generation means (terminal)

[0424] The device generates video frames based on the reproduced audio and provides them to the user. The video is generated using synthetic video technology (e.g., Deepfake technology).

[0425] 7. Emotion Recognition Means (Emotion Engine)

[0426] The emotion engine analyzes real-time data obtained from the user's camera and microphone to identify the user's emotional state, and is implemented using Microsoft's Azure Face API and Emotion API.

[0427] 8. Dialogue Content Adjustment Method (Server)

[0428] The server generates personalized dialogue based on the recognized emotional state, using a generative AI model to dynamically create conversational content that matches the user's emotions.

[0429] 9. Physical store devices (tablets and smartphones)

[0430] The system operates via a tablet device installed in a brick-and-mortar store or a user's smartphone, which displays audio and video to recreate a conversation with the deceased.

[0431] Specific examples

[0432] A user visiting a brick-and-mortar store (e.g., a long-established inn or restaurant) accesses a tablet device installed in the store. The server collects the communication history of the deceased person provided by the user and generates a personality model. When the user visits a specific location (e.g., a place frequently visited by the deceased), the emotion engine uses facial recognition technology to detect the user's emotional state. If the server recognizes that the user is crying, it uses the generative AI model to generate dialogue content to comfort the user and displays it on the device.

[0433] Prompt Sentence Examples

[0434] Generate it as follows: You arrive at the lobby of an inn. Noticing your tears, your deceased father speaks to you kindly, saying, "Cheer up, I'm always watching over you."

[0435] This system allows users to experience a realistic reunion with their deceased loved ones and gain emotional comfort, thus recreating memories of the deceased in a physical store and reducing the psychological burden on users.

[0436] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0437] Step 1:

[0438] The user logs in to a tablet device installed in a physical store and gives permission to collect the communication history of the deceased. The input is the SNS API key and login information provided by the user. The server uses this information to collect data such as the deceased's SNS posts, call records, and message history. The output is the acquired communication history data.

[0439] Step 2:

[0440] The server converts the collected communication history data into a text format. The input is communication history data stored in JSON format or other structured data format. The server runs a data conversion program to extract the text data. The output is a list of the actual conversations in text format.

[0441] Step 3:

[0442] The server analyzes the text data using natural language processing techniques. The input is the extracted text data. The server uses Python NLP libraries (e.g., spaCy or NLTK) to extract keywords, analyze writing style, and identify frequently occurring phrases. The output is analyzed data containing the linguistic features of the deceased.

[0443] Step 4:

[0444] The server generates a personality model of the deceased person based on the analyzed data. The input is the analyzed data, including linguistic features. The server uses OpenAI's GPT-based generative AI model to build a personality model, including speech patterns and intonation. The output is a personality model of the deceased person.

[0445] Step 5:

[0446] The device recreates the voice using a personality model downloaded from a server. The input is the personality model. The device uses voice changer technology (e.g., Respeecher or Google's Tacotron) to generate the deceased person's voice. The output is the recreated voice.

[0447] Step 6:

[0448] The device generates video frames based on the reproduced audio. The input is the reproduced audio and video data of the deceased person before their death. The device uses synthetic video technology (e.g., Deepfake technology) to generate video that matches the audio. The output is video frames.

[0449] Step 7:

[0450] The emotion engine analyzes real-time data acquired from the user's camera and microphone to identify the user's emotional state. The input is data acquired in real time of the user's facial expressions and tone of voice. The emotion engine performs emotion recognition using Microsoft's Azure Face API and Emotion API. The output is data indicating the user's emotional state.

[0451] Step 8:

[0452] The server generates dialogue content according to the recognized emotional state. The input is the user's emotional state data and a personality model of the deceased. The server uses a generative AI model to create personalized dialogue content tailored to the user's emotions. The output is the generated dialogue content.

[0453] Step 9:

[0454] The device recreates the conversation with the deceased based on the generated conversation content. The input is the generated conversation content and video frames. The device displays this to the user, providing a realistic conversation with the deceased. The output is the video and audio of the deceased displayed on the user's screen.

[0455] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0456] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0457] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0458] [Second embodiment]

[0459] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0460] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0461] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0462] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0463] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0464] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0465] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0466] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0467] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0468] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0469] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0470] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0471] This system collects and analyzes the communication history of a deceased person, and generates and recreates a personality model based on the linguistic features of the person. This system allows users to virtually experience communication with the deceased person through video frames that make the deceased appear as if they were alive.

[0472] System configuration

[0473] The system of the present invention comprises the following components:

[0474] 1. Data collection method (server)

[0475] 2. Data extraction method (server)

[0476] 3. Data analysis method (server)

[0477] 4. Personality model generation means (server)

[0478] 5. Consent confirmation means (user)

[0479] 6. Audio reproduction means (terminal)

[0480] 7. Video frame generation means (terminal)

[0481] Program processing overview

[0482] Data collection method (server)

[0483] The user provides the server with the social networking service's API key and other login information, which the server uses to collect the deceased person's communication history, including social networking posts, call records, and message history.

[0484] Example: The server uses the Twitter API to retrieve the tweet history of a specified user ID.

[0485] Data extraction method (server)

[0486] The server converts the collected data into an appropriate format and extracts it as text data.

[0487] Example: A server converts JSON-formatted tweet data into a text format.

[0488] Data analysis method (server)

[0489] The server uses natural language processing technology to analyze the text data, extract keywords and analyze writing style, and identify linguistic features.

[0490] Example: The server uses an NLP library to extract frequently used phrases and their context.

[0491] Personality model generation means (server)

[0492] The server generates a personality model based on the identified linguistic features, including the deceased's speech patterns and intonation.

[0493] Example: The server models typical response patterns during a conversation based on the analyzed features.

[0494] Consent confirmation method (user)

[0495] The user must confirm whether the deceased person has agreed to use the service before they die, and if necessary, upload a consent form to the server.

[0496] Example: A user uploads a scanned consent form through a web portal.

[0497] Audio reproduction means (terminal)

[0498] The device downloads a personality model from the server and reproduces the voice using a voice changer.

[0499] Example: The device uses voice changer software to generate the voice of a deceased person.

[0500] Video frame generation means (terminal)

[0501] The terminal generates video frames based on the reproduced audio and presents them to the user.

[0502] Example: A device uses synthetic video techniques to create video that accompanies the reproduced audio.

[0503] Specific examples

[0504] For example, if a deceased person frequently used the greeting "Good morning," the server can identify that phrase and generate a video message saying "Good morning" through a voice changer on the device. Looking at the device screen, the user can feel as if the deceased person was greeting them in person.

[0505] In this way, the system of the present invention can generate and provide users with realistic interactive video frames that incorporate the memories and characteristics of the deceased, thereby helping to ease the grief of losing a loved one.

[0506] The processing flow will be explained below.

[0507] Step 1:

[0508] The user provides the server with the API key and login information of the SNS. Specifically, the user enters and submits the API key of Twitter or other SNS through a web form.

[0509] Step 2:

[0510] The server uses the provided API key to collect the post data of the deceased person from the SNS. Specifically, the server calls the Twitter API to obtain the tweet history of the specified user ID.

[0511] Step 3:

[0512] The server converts the collected data into an appropriate format and extracts it as text data. Specifically, the server converts the JSON-formatted tweet data into text format and extracts each sentence.

[0513] Step 4:

[0514] The server analyzes the extracted text data using natural language processing techniques. Specifically, the server uses an NLP library to extract frequently used phrases and keywords from the text.

[0515] Step 5:

[0516] The server then uses the text data analyzed in the previous step to identify the deceased's linguistic characteristics, such as writing style, tone, specific phrases, and intonation patterns.

[0517] Step 6:

[0518] Based on the identified features, the server generates a personality model that includes the deceased's speech patterns and intonation. Specifically, the analysis results are used to create a model that incorporates the deceased's typical speech patterns and reactions.

[0519] Step 7:

[0520] The user uploads a document to the server to confirm that they have agreed to use the service. Specifically, the user uploads a scanned copy of the consent form via a web portal.

[0521] Step 8:

[0522] The device downloads the generated personality model from the server and reproduces the voice using voice changer technology. Specifically, the device obtains the personality model via the Internet and reproduces it using voice synthesis software.

[0523] Step 9:

[0524] The device generates video frames based on the reproduced audio. Specifically, it combines the synthesized audio with footage of the deceased (previous video clips or photos) to create natural-sounding video frames.

[0525] Step 10:

[0526] The terminal provides the generated video frames to the user, specifically, displays the video frames on the terminal screen so that the user can view them.

[0527] Example 1

[0528] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0529] In modern society, there is a lack of ways to continuously remember the deceased, and it is particularly difficult to remember them through conversation or audio. Furthermore, there is no technology that can reproduce the specific linguistic characteristics and mannerisms of the deceased and imitate their personality, making it impossible to continue interacting with them. Therefore, there is a need for a way for users to experience memories of the deceased more realistically and maintain an emotional connection.

[0530] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0531] In this invention, the server includes means for collecting communication data of a user, means for extracting text data from the communication data, means for analyzing the text data using natural language processing technology and identifying linguistic characteristics of the user, means for generating a character model of the user based on the identified linguistic characteristics, means for reproducing audio based on the character model, and means for generating video frames using the reproduced audio. This allows the user to experience realistic conversations and audio with the deceased, enabling them to maintain an emotional connection with the deceased.

[0532] "User" refers to an individual or group of people who use the System.

[0533] "Communication data" refers to historical information about a user's digital communications, such as social media posts, call records, and message history.

[0534] "Text data" refers to text information extracted from communication data.

[0535] "Natural language processing technology" refers to computer technology for analyzing text data, extracting keywords, analyzing writing style, etc.

[0536] "Linguistic features" refer to the phrases, expressions, style, and other characteristics that users frequently use in a particular language.

[0537] "Person model" refers to a digital model that is generated based on a user's linguistic characteristics and mimics the user's speech patterns and intonation.

[0538] "Means for reproducing voice" refers to technologies and tools for generating voice that reproduces the characteristics of a user using a person model.

[0539] "Means for generating video frames" refers to techniques and tools for generating video frames in accordance with the reproduced audio.

[0540] "Service Use Agreement" refers to an agreement indicating that a user gives permission to use the system before death.

[0541] "Pronunciation characteristics" refers to the characteristics of a user's voice, including intonation, accent, frequently used expressions, etc.

[0542] "Frequent expressions" refer to phrases and expressions that users use frequently.

[0543] "Images" refers to your visual content, such as still images and videos.

[0544] "Synchronization" refers to the technique of playing audio and video in sync.

[0545] The present invention is a system for collecting, analyzing, and reproducing the communication history of a deceased person. Detailed procedures for implementing the present invention will be described below.

[0546] System configuration

[0547] The system of the present invention comprises the following components:

[0548] 1. Data collection means (server): Collects user communication data.

[0549] 2. Data extraction means (server): Extracts text data from the collected communication data.

[0550] 3. Data analysis means (server): Using natural language processing technology, the text data is analyzed and the user's linguistic characteristics are identified.

[0551] 4. Personality model generation means (server): Generates a personality model of the user based on the identified linguistic features.

[0552] 5. Consent confirmation means (user): Confirm whether the user agreed to use the service before death.

[0553] 6. Voice reproduction means (terminal): Reproduces the voice using the personality model downloaded from the server.

[0554] 7. Video frame generation means (terminal): Generates video frames based on the reproduced audio.

[0555] Hardware and software used

[0556] Servers: High-performance servers and cloud computing platforms (e.g., AWS, Google Cloud) are used.

[0557] Natural Language Processing library: SpaCy, NLTK, or other NLP library.

[0558] Generative AI models: Advanced AI models such as GPT-3.

[0559] Voice changer software: DeepTalk, VoxCeleb2, etc.

[0560] Facial animation technology: Synthetic video techniques such as DeepFake.

[0561] Specific step-by-step instructions

[0562] 1. Data Collection

[0563] The user provides the server with permission to access the deceased person's social media accounts and call history.

[0564] The server uses the provided API key and login information to collect communication history data from the specified account.

[0565] 2. Data extraction

[0566] The server converts the collected data into a text format, for example, extracting the message content from JSON-formatted tweet data into plain text.

[0567] 3. Data Analysis

[0568] The server uses natural language processing techniques to analyze the text data and extract keywords, specifically by identifying frequently occurring phrases and their contexts using libraries such as SpaCy and NLTK.

[0569] 4. Personality Model Generation

[0570] The server uses the identified linguistic features to generate a person model that includes the deceased person's speech patterns and intonation, using a generative AI model such as GPT-3.

[0571] 5. Confirmation of consent

[0572] The user confirms the deceased person's consent to use the service during their lifetime, and if necessary, uploads a scanned copy of the consent form to the server.

[0573] 6. Audio Reproduction

[0574] The device uses a personality model downloaded from the server and reproduces the voice using voice changer software (e.g., DeepTalk).

[0575] 7. Video Frame Generation

[0576] The device generates video frames based on the reproduced audio using facial animation technology (e.g., DeepFake).

[0577] Specific examples

[0578] For example, if the deceased frequently used the greeting "Good morning," the server can identify that expression and the device can use a voice changer to generate a video message saying "Good morning." The user can feel as if the deceased is greeting them directly through the device screen.

[0579] Prompt Sentence Examples

[0580] "Collect all tweets from user ID 'example_user'."

[0581] "Extract tweet content in text format from JSON format data."

[0582] "Extract frequently occurring words and phrases from text data."

[0583] "Use the extracted features to create a model that reproduces the speech patterns of the deceased."

[0584] "Please scan the consent form and upload it through the web portal."

[0585] "Load the model data and recreate the voice using voice changer software."

[0586] "Create video frames with facial animation technology based on the generated audio."

[0587] Although the embodiment for carrying out the present invention has been specifically described above, the present invention is not limited to this, and other embodiments and modifications are possible.

[0588] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0589] Step 1: Data collection

[0590] Users provide their social networking service API keys and login information, which the server uses to collect communication data.

[0591] Input: User's API key or login information

[0592] Output: Social media posts, call history, message history data

[0593] Specific operation: The server uses the Twitter API to retrieve the tweet history of the specified user ID in JSON format. As an example, enter the prompt "Please collect all tweets from user ID 'example_user'."

[0594] Step 2: Data extraction

[0595] The server extracts text data from the collected data.

[0596] Input: Collected social media posts, call history, and message history data

[0597] Output: Text data

[0598] Specific operation: The server extracts the text message content from the JSON format data and converts it to plain text. Specifically, it inputs the prompt "Please extract the tweet content in text format from the JSON format data."

[0599] Step 3: Data analysis

[0600] The server uses natural language processing technology to analyze the text data, extract keywords, and analyze frequently occurring phrases and writing styles.

[0601] Input: Extracted text data

[0602] Output: Linguistic features (keywords, frequent phrases, contextual information)

[0603] Specific behavior: The server uses an NLP library (e.g., SpaCy or NLTK) to extract frequent words and phrases from the text and identify the context. The server inputs the prompt: "Extract frequent words and phrases from the text data."

[0604] Step 4: Personality model generation

[0605] The server generates a personality model based on the identified linguistic features, including the deceased's speech patterns and intonation.

[0606] Input: Identified linguistic features

[0607] Output: Personality model

[0608] Specific operation: The server uses a generative AI model (e.g., GPT-3) to create a model that mimics the deceased person's speech patterns. It inputs the prompt, "Please use the extracted features to create a model that reproduces the deceased person's speech patterns."

[0609] Step 5: Confirm consent

[0610] The user checks whether the deceased person agreed to use the service before they died, and uploads a consent form to the server if necessary.

[0611] Input: Consent form (possibly scanned)

[0612] Output: Service use authorization

[0613] Specific operation: The user scans the consent form and uploads it to the server through the web portal. The server verifies and records the uploaded consent form. It then instructs the user to "scan the consent form and upload it through the web portal."

[0614] Step 6: Audio Reproduction

[0615] The device downloads a personality model from the server and reproduces the voice using voice changer software.

[0616] Input: Personality model data

[0617] Output: Reproduced audio

[0618] Specific operation: The device loads the model data downloaded from the server and generates voice using voice changer software such as DeepTalk or VoxCeleb2. The device instructs the user to "load the model data and reproduce the voice using voice changer software."

[0619] Step 7: Video Frame Generation

[0620] The terminal generates video frames in accordance with the reproduced audio and presents them to the user.

[0621] Input: Reproduced audio

[0622] Output: Video frame

[0623] Specific operation: The device uses facial animation technology (e.g., DeepFake) to create a video frame that corresponds to the reproduced audio. It then prompts the device to "create a video frame using facial animation technology based on the generated audio."

[0624] (Application example 1)

[0625] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0626] When a loved one passes away, there are limited ways to cherish memories of the deceased. Conventional technology lacks the means to reproduce the words and conversations of the deceased, in addition to visual records such as photographs and videos, making it difficult to recreate the emotional connection with the deceased. Furthermore, there is a demand for providing a natural conversation experience that includes the unique phrases and intonations used by the deceased, but no concrete means for achieving this have been established.

[0627] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0628] In this invention, the server includes means for collecting the communication history of the deceased, means for extracting text data from the communication history, and means for analyzing the text data using natural language processing technology to identify the linguistic characteristics of the deceased. This allows a personality model of the deceased to be generated based on the analyzed data, enabling the user to experience a simulated conversation with the deceased. Furthermore, the content of the conversation provided based on the generated personality model is more realistic, allowing the user to feel a deeper emotional connection with the deceased. This allows for a more realistic reproduction of memories with the deceased.

[0629] "Communication history" is a general term for records such as messages, call records, and social media posts sent and received through electronic communication means used by the deceased during their lifetime.

[0630] "Text data" refers to character information and sentences extracted from the communication history.

[0631] "Natural language processing technology" is a computer processing technology for understanding, analyzing, and generating human language, and includes a wide range of algorithms and methods for mechanically processing natural language.

[0632] "Linguistic features" refer to the characteristics of language use possessed by a particular writer or speaker, such as frequent phrases, lexical choice, and stylistic patterns.

[0633] A "personality model" is a computer model constructed to recreate the language usage and speech patterns of a deceased person based on their linguistic characteristics.

[0634] "Voice reproduction means" refers to techniques or devices that use a personality model to reproduce the voice, intonation, and distinctive speaking style of the deceased.

[0635] A "video frame" is an individual image that reproduces a specific moment in time as a video, and is the basic unit for expressing movement as a video when displayed continuously.

[0636] "Pseudo-dialogue" refers to providing a user with an experience that makes them feel as if they are conversing with the deceased person using the generated personality model.

[0637] "Dialogue content" refers to the content of words and messages exchanged between the user and the personality model of the deceased person in a simulated dialogue.

[0638] This invention is a system that collects and analyzes the communication history of a deceased person, and generates and recreates a personality model based on the linguistic features of the communication history. This system allows users to enjoy a virtual conversation experience in which the deceased person feels as if they are alive. This system incorporates key elements such as "communication history," "text data," "natural language processing technology," "linguistic features," "personality model," "means for reproducing voice," "video frames," "simulated dialogue," and "dialogue content."

[0639] Hardware and software used

[0640] Hardware: Smartphone (iPhone or Android)

[0641] Software: Python, Flask, TensorFlow / Keras, Pandas, NLTK, Twilio API, Google Cloud Text-to-Speech API

[0642] Program processing overview

[0643] 1. Data collection method (server)

[0644] The server receives the API key of the deceased person's social media account from the user. Using this API key, the server accesses the deceased person's social media account and collects their communication history, including Twitter and Facebook posts, messages, and call records.

[0645] 2. Data extraction method (server)

[0646] The server converts the collected social media data into an appropriate format and extracts it as text data. The collected data is provided in JSON format, but it is converted into an easily handled text format.

[0647] 3. Data analysis method (server)

[0648] The server analyzes the extracted text data using natural language processing libraries such as Pandas and NLTK to identify the linguistic characteristics of the deceased and extract frequent phrases, expressions, and writing style.

[0649] 4. Personality model generation means (server)

[0650] Based on the analyzed linguistic features, the server uses TensorFlow and Keras to generate a personality model of the deceased, a computer model that reproduces the speech patterns and intonation of the deceased.

[0651] 5. Consent confirmation means (user)

[0652] Users must verify through a web portal whether the deceased person consented to use the service before they died, and then upload a scanned copy of the consent form.

[0653] 6. Audio reproduction means (terminal)

[0654] The device uses a personality model downloaded from the server to reproduce the voice of the deceased person using voice synthesis technology such as the Google Cloud Text-to-Speech API.

[0655] 7. Video frame generation means (terminal)

[0656] As the audio plays, the device uses synthetic video technology to generate a video of the deceased, providing the user with realistic video frames synchronized with the audio.

[0657] Specific examples

[0658] For example, if a deceased person frequently used the greeting "Good morning," the server can identify this phrase and generate a video message in which the device says "Good morning" through a voice changer. The following is an example of a prompt sentence:

[0659] Example prompt sentence:

[0660] Prompt: "Good morning. What are you planning to do today?"

[0661] Response: The generated personality model responds, "Good morning! I'm planning to go for a walk today. How about you?"

[0662] In this way, the user can experience memories of the deceased more realistically through simulated dialogue.

[0663] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0664] Step 1:

[0665] Data collection method (server)

[0666] The server receives the API key and login information of the deceased person's SNS from the user. The server then collects the deceased person's communication history (messages, posts, call records, etc.) from various SNS platforms (e.g., Twitter, Facebook). The input data is the API key and user ID, and the output is communication history data in JSON format.

[0667] Step 2:

[0668] Data extraction method (server)

[0669] The server converts the collected SNS data from an appropriate format (for example, JSON format) into text data. Here, the tweet text and message content are extracted and organized as text data. The input is communication history data in JSON format, and the output is formatted text data. Specifically, the server extracts the necessary information from the data obtained from each SNS platform and converts it into text format.

[0670] Step 3:

[0671] Data analysis method (server)

[0672] The server analyzes the text data using natural language processing techniques (e.g., Pandas or NLTK). The analysis identifies the linguistic characteristics of the deceased (frequent phrases, writing style, intonation, etc.). The input is text data, and the output is metadata about the linguistic characteristics of the deceased. Specific operations include word frequency analysis, morphological analysis, and contextual analysis.

[0673] Step 4:

[0674] Personality model generation means (server)

[0675] Based on the analysis results, the server uses TensorFlow and Keras to generate a personality model of the deceased. This is a machine learning model that learns the deceased's speech patterns. The input is metadata about the linguistic features of the deceased, and the output is a trained personality model. Specifically, it trains an LSTM model using the deceased's speech data.

[0676] Step 5:

[0677] Consent confirmation method (user)

[0678] The user checks through the web portal whether the deceased person had consented to use the service before they died, and then uploads a scanned copy of the consent form to the server. The input is the scanned consent form (electronic file), and the output is the result of uploading the consent form to the server. Specifically, the user selects the consent form on the portal site and presses the upload button.

[0679] Step 6:

[0680] Audio reproduction means (terminal)

[0681] The device uses the personality model downloaded from the server and generates the voice of the deceased person using the Google Cloud Text-to-Speech API, etc. The input is the personality model and text data, and the output is an audio file. Specifically, it converts the text-generated dialogue into audio.

[0682] Step 7:

[0683] Video frame generation means (terminal)

[0684] The device uses synthetic video technology to generate a video of the deceased person in sync with the audio. The input is the generated audio and video data of the deceased person while they were alive, and the output is video frames synchronized with the audio. Specifically, the device generates and edits a video sequence corresponding to the audio.

[0685] This allows the user to experience a simulated conversation with the deceased, allowing memories of the deceased to be recreated more realistically.

[0686] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0687] This system collects and analyzes the communication history of a deceased person, and generates and recreates a personality model based on the linguistic features. This system also incorporates an emotion engine that recognizes the user's emotions, providing dialogue that responds to the user's emotions, thereby realizing a more realistic communication experience.

[0688] System configuration

[0689] The system of the present invention comprises the following components:

[0690] 1. Data collection method (server)

[0691] 2. Data extraction method (server)

[0692] 3. Data analysis method (server)

[0693] 4. Personality model generation means (server)

[0694] 5. Consent confirmation means (user)

[0695] 6. Audio reproduction means (terminal)

[0696] 7. Video frame generation means (terminal)

[0697] 8. Emotion Recognition Means (Emotion Engine)

[0698] Program processing overview

[0699] Data collection method (server)

[0700] The user provides the server with the API key and login information of the social networking site, which the server uses to collect the deceased person's communication history, including social networking site posts, call records, and message history.

[0701] Example: The server uses the Twitter API to retrieve the tweet history of a specified user ID.

[0702] Data extraction method (server)

[0703] The server converts the collected data into an appropriate format and extracts it as text data.

[0704] Example: The server converts JSON-formatted tweet data into text format and extracts each sentence.

[0705] Data analysis method (server)

[0706] The server uses natural language processing technology to analyze the text data, extract keywords and analyze writing style, and identify linguistic features.

[0707] Example: A server uses an NLP library to extract frequently used phrases and keywords from text.

[0708] Personality model generation means (server)

[0709] The server generates a personality model based on the identified linguistic features, including the deceased's speech patterns and intonation.

[0710] Example: The server models typical response patterns during a conversation based on the analyzed features.

[0711] Consent confirmation method (user)

[0712] A document is uploaded to the server to confirm that the user has agreed to use the service.

[0713] Example: A user uploads a scanned consent form through a web portal.

[0714] Audio reproduction means (terminal)

[0715] The device downloads the generated personality model from the server and reproduces the voice using voice changer technology.

[0716] Example: The device uses voice changer software to generate the voice of a deceased person.

[0717] Video frame generation means (terminal)

[0718] The terminal generates video frames based on the reproduced audio and presents them to the user.

[0719] Example: A device uses synthetic video techniques to create video that accompanies the reproduced audio.

[0720] Emotion recognition means (emotion engine)

[0721] The emotion engine analyzes real-time data obtained from the user's camera and microphone to identify the user's emotional state.

[0722] Example: An emotion engine analyzes a user's facial expressions and tone of voice to determine whether they are happy, sad, or surprised.

[0723] How the Emotion Engine Works

[0724] The emotion engine provides a means to recognize the user's emotional state in real time and adjust the audio and video frames of the deceased based on that. For example, if the emotion engine determines that the user is sad, it can generate a soothing voice or comforting message for the deceased. This functionality allows the user to feel more connected to the deceased.

[0725] Specific examples

[0726] For example, if the deceased frequently used the greeting "good morning," the emotion engine may determine that the user is feeling unwell. The emotion engine will adjust the reproduced voice to be more gentle and comforting depending on the situation, and change the video frame accordingly. Looking at the device screen, the user will feel as if the deceased is gently encouraging them.

[0727] In this way, the system of the present invention can generate and provide users with realistic interactive video frames that incorporate the memories and characteristics of the deceased, enabling personalized interactions based on the user's emotional state, helping to alleviate the grief of losing a loved one.

[0728] The processing flow will be explained below.

[0729] Step 1:

[0730] The user provides the server with the API key and login information of the SNS. Specifically, the user enters the API key of Twitter or other SNS through a web form and submits it.

[0731] Step 2:

[0732] The server uses the provided API key to collect the post data of the deceased person from the SNS. Specifically, the server calls the Twitter API to obtain the tweet history of the specified user ID.

[0733] Step 3:

[0734] The server converts the collected data into an appropriate format and extracts it as text data. Specifically, the server converts the JSON-formatted tweet data into text format and extracts each sentence.

[0735] Step 4:

[0736] The server analyzes the extracted text data using natural language processing techniques. Specifically, the server uses an NLP library to extract frequently used phrases and keywords from the text.

[0737] Step 5:

[0738] The server then uses the text data analyzed in the previous step to identify the deceased's linguistic characteristics, such as writing style, tone, specific phrases, and intonation patterns.

[0739] Step 6:

[0740] Based on the identified characteristics, the server generates a personality model that includes the deceased's speech patterns and intonation. Specifically, it uses the analysis results to incorporate the deceased's frequently used phrases and reaction patterns into the model.

[0741] Step 7:

[0742] The user uploads a document to the server to confirm that they have agreed to use the service. Specifically, the user uploads a scanned copy of the consent form via a web portal.

[0743] Step 8:

[0744] The device downloads the generated personality model from the server and uses voice changer technology to recreate the voice of the deceased. Specifically, the device obtains the personality model via the internet and uses voice synthesis software to generate the voice of the deceased.

[0745] Step 9:

[0746] The device generates video frames based on the reproduced audio. Specifically, it combines the synthesized audio with footage of the deceased (previous video clips or photos) to create natural-sounding video frames.

[0747] Step 10:

[0748] The terminal displays the video frames to the user, specifically, displays the video frames on the terminal's display so that the user can view them.

[0749] Step 11:

[0750] The emotion engine captures real-time data from the user's camera and microphone and analyzes their emotions by analyzing their facial expressions, tone of voice, and word choice to determine their current emotional state.

[0751] Step 12:

[0752] Based on the analysis, the emotion engine adjusts the content of the deceased person's audio and video frames. For example, if the user is sad, the message of the deceased will be changed to be more comforting.

[0753] Step 13:

[0754] The device plays the adjusted content and displays it to the user. Specifically, the content generated based on instructions from the emotion engine is displayed on the display, and the user watches and listens to it.

[0755] Example 2

[0756] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0757] Although there are systems that can recreate the memories and characteristics of the deceased and allow users to have a realistic conversation with them, these systems lack the ability to recognize the user's emotional state and dynamically adjust the conversation content based on that emotion. This limits the conversational experiences provided, making it difficult for users to feel a deeper connection with the deceased.

[0758] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0759] In this invention, the server includes means for collecting the communication history of the deceased, means for extracting text data from the communication history, means for analyzing the text data using natural language processing technology and identifying linguistic characteristics of the deceased, means for generating a personality model of the deceased based on the identified linguistic characteristics, means for downloading the personality model of the deceased from the server to a terminal, means for reproducing voice based on the personality model, means for generating video frames based on the reproduced voice, means for recognizing the emotional state of the user, and means for adjusting the voice and video frames based on the emotional state of the user. This allows the user to not only realistically experience memories and characteristics of the deceased, but also to have a personalized interaction experience according to emotions.

[0760] "Communication history" refers to a series of communication data such as social media posts, call records, and message history made by the deceased person while they were alive.

[0761] "Text data" refers to data in sentence format extracted from the communication history.

[0762] "Natural language processing technology" refers to technology used by computers to understand, analyze, and generate human language.

[0763] "Linguistic characteristics" refer to linguistic characteristics and patterns, such as the particular words, writing style, and frequently used phrases used by a particular person.

[0764] A "personality model" is a data model designed to reproduce the linguistic characteristics, speech patterns, intonation, etc. of a deceased person.

[0765] "Voice reproduction" refers to the technique or process of recreating the voice of a deceased person based on a generated personality model.

[0766] A "video frame" is a frame of video that corresponds to the reproduced audio and shows the image and gestures of the deceased.

[0767] "Emotional state" refers to a user's current psychological and sensory state, which is typically determined through facial expressions, tone of voice, etc.

[0768] "Emotion recognition" refers to the technology of analyzing a user's emotional state and thereby identifying a specific emotion.

[0769] A "personalized interaction experience" is an interaction that is optimized for a user's unique characteristics and situation, and is adjusted based on their emotional state.

[0770] This system collects and analyzes the communication history of a deceased person, and generates and recreates a personality model based on the linguistic features. This system also incorporates an emotion engine that recognizes the user's emotions, providing dialogue that responds to the user's emotions, thereby realizing a more realistic communication experience.

[0771] System configuration

[0772] The system of the present invention comprises the following components:

[0773] 1. Data collection method (server)

[0774] 2. Data extraction method (server)

[0775] 3. Data analysis method (server)

[0776] 4. Personality model generation means (server)

[0777] 5. Consent confirmation means (user)

[0778] 6. Audio reproduction means (terminal)

[0779] 7. Video frame generation means (terminal)

[0780] 8. Emotion Recognition Means (Emotion Engine)

[0781] Detailed system configuration and processing method

[0782] Data collection method (server)

[0783] The user enters their social networking service API key and login information, which is then sent in a secure format to the server, which then uses this information to collect the deceased person's communication history, including social networking posts, call logs, and message history. This process uses standard APIs, such as the Twitter API and Facebook API.

[0784] Example: The server uses the Twitter API to retrieve the tweet history of a specified user ID.

[0785] Data extraction method (server)

[0786] The server analyzes the collected data and extracts it as text data. At this time, it converts the necessary parts of the data in JSON or XML format into text format and extracts each sentence and word.

[0787] Example: The server converts JSON-formatted tweet data into text format and extracts each sentence.

[0788] Data analysis method (server)

[0789] The server uses natural language processing (NLP) techniques to analyze the extracted text data. During the analysis, it identifies linguistic features such as keywords, frequent phrases, and writing style. This process uses NLP libraries (e.g., NLTK, spacy).

[0790] Example: The server uses an NLP library to extract frequently used phrases and keywords from the text.

[0791] Personality model generation means (server)

[0792] The server generates a personality model based on the analyzed linguistic features, including the deceased's speech patterns and intonation, which is then used to simulate conversations similar to those of the deceased.

[0793] Example: The server uses the analyzed features to model typical response patterns during a conversation.

[0794] Consent confirmation method (user)

[0795] The user uploads a consent form for using the service to the server in PDF format, etc. The server then checks the uploaded consent form and determines whether or not the use of the service is permitted.

[0796] Example: User uploads scanned consent form through web portal.

[0797] Audio reproduction means (terminal)

[0798] The device uses a personality model downloaded from the server and uses voice changer technology (e.g., Voicemod, iZotope) to recreate the voice of the deceased person, which is used during interactions with the user.

[0799] Example: The device uses voice changer software to generate the voice of the deceased person.

[0800] Video frame generation means (terminal)

[0801] Based on the reproduced audio, the device uses synthesis technology (e.g., D-ID, Reallusion) to generate video frames that display the deceased person's face and gestures in sync with the audio.

[0802] Example: The device uses synthetic video techniques to create a video that matches the reproduced audio.

[0803] Emotion recognition means (emotion engine)

[0804] The emotion engine analyzes real-time data captured from the user's camera and microphone. It uses facial expression analysis software (e.g., Affectiva, Microsoft Azure Face API) and voice analysis tools to identify the user's emotional state.

[0805] Example: An emotion engine analyzes a user's facial expressions and tone of voice to determine whether they are sad, happy, or surprised.

[0806] Examples of concrete examples and prompts

[0807] For example, if the deceased frequently used the greeting "Good morning," the emotion engine may determine that the user's emotional state was poor. In this case, the emotion engine will adjust the reproduced voice to be calm and comforting, and change the video frame accordingly. Looking at the device screen, the user will feel as if the deceased is gently encouraging them.

[0808] Example prompt sentence:

[0809] It uses phrases that allow users to start a natural conversation, such as "Hi, how are you doing?" or "How was your day?"

[0810] In this way, the system of the present invention can realistically recreate the memories and characteristics of the deceased and provide a personalized interaction experience for the user.

[0811] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0812] Step 1: Enter user information

[0813] The user enters the SNS API key and login information, which is then sent to the server in a secure format. Specifically, the user enters the information on the system's login screen and presses the send button. The input data is the SNS API key and login information, and the output is the API key and login information sent to the server.

[0814] Step 2: Data collection

[0815] The server uses the API key and login information sent by the user to collect the deceased person's communication history, such as social media posts, call records, and message history. Specifically, the server uses an appropriate API (e.g., Twitter API, Facebook API) to obtain data for the specified user ID. The input data is the API key and login information provided by the user in the previous step, and the output is the collected communication history (e.g., social media post data, call records).

[0816] Step 3: Data extraction

[0817] The server converts the collected communication history into an appropriate format and extracts it as text data. Specifically, the server converts SNS data in JSON or XML format into text format and extracts each sentence. The input data is the communication history (e.g., tweet data in JSON format), and the output is the extracted text data (e.g., single sentences of text).

[0818] Step 4: Data analysis

[0819] The server uses natural language processing (NLP) techniques to analyze the extracted text data. During the analysis process, it identifies linguistic features such as keywords, frequently occurring phrases, and writing style. Specifically, the server uses an NLP library (e.g., NLTK, spacy) to extract frequently used phrases and keywords from the text. The input data is the text data, and the output is the analyzed linguistic features (e.g., a list of keywords and phrases).

[0820] Step 5: Personality model generation

[0821] The server generates a personality model based on the analyzed linguistic features, including the deceased's speech patterns and intonation. Specifically, the server uses a machine learning algorithm to model typical response patterns during conversations based on the analyzed features. The input data are the linguistic features, and the output is the generated personality model.

[0822] Step 6: Confirm consent

[0823] The user uploads the consent form for using the service to the server in PDF format or other format. Specifically, the user scans the consent form and uploads it through a web portal. The input data is the consent form in PDF format, and the output is a consent confirmation flag on the server.

[0824] Step 7: Audio Reproduction

[0825] The device uses a personality model downloaded from the server and voice changer technology (e.g., Voicemod, iZotope) to recreate the deceased's voice. Specifically, the device converts the deceased's voice pattern into audio data based on the downloaded personality model. The input data is the personality model, and the output is the recreated audio data.

[0826] Step 8: Video Frame Generation

[0827] The device generates video frames based on the reproduced audio. Specifically, the device uses synthesis technology (e.g., D-ID, Reallusion) to create video that matches the reproduced audio. The input data is the reproduced audio data, and the output is the generated video frames.

[0828] Step 9: Emotion Recognition

[0829] The emotion engine analyzes real-time data acquired from the user's camera and microphone. Specifically, the emotion engine uses facial expression analysis software (e.g., Affectiva, Microsoft Azure Face API) and voice analysis tools to identify the user's emotional state. The input data is the real-time data acquired from the camera and microphone, and the output is the identified user's emotional state.

[0830] Step 10: Dialogue Adjustment

[0831] The emotion engine appropriately adjusts the content and tone of the dialogue based on the emotion data acquired in the previous stage. For example, if the user is sad, the emotion engine generates a calm voice and comforting words, and optimizes the video frames accordingly. The input data is the user's emotional state, and the output is the adjusted voice data and video frames.

[0832] (Application example 2)

[0833] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0834] There is a need for a means to maintain the emotional connection between the deceased and those left behind, particularly to ease the grief of losing the deceased. However, conventional technologies only reproduce the personality of the deceased, and are unable to engage in dialogue that reflects the user's emotional state, making it difficult to provide a realistic communication experience. Furthermore, there is no system in physical stores that allows users to feel emotionally reassured, making it difficult to alleviate the user's psychological burden.

[0835] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0836] In this invention, the server includes means for collecting the communication history of the deceased, means for extracting text data from the communication history, means for analyzing the text data using natural language processing technology to identify the linguistic characteristics of the deceased, means for generating a personality model of the deceased based on the identified linguistic characteristics, means for reproducing voice based on the personality model, means for generating video frames using the reproduced voice, means for recognizing the emotional state of a user in real time, means for adjusting the content of the dialogue in accordance with the recognized emotional state, and means for reproducing a dialogue with the deceased via a tablet terminal or smartphone in a physical store. This allows for a realistic reproduction of the personality of the deceased while enabling personalized dialogue in accordance with the user's emotional state, thereby providing users with a sense of emotional security even in a physical store.

[0837] "Communication history" is a record of past interactions recorded as digital data, such as text messages, call logs, and social media posts, across the mediums used by individuals to communicate.

[0838] "Natural language processing technology" is a field of artificial intelligence that enables computers to understand, interpret, and generate human language, and is a technology that analyzes text and extracts linguistic features.

[0839] A "personality model" is a digital recreation of a specific individual's linguistic characteristics, speech patterns, and presence, constructed from the deceased's text data and call records.

[0840] "Voice reproduction" is a technology that generates voice by imitating the speaking style and voice quality of a deceased person based on a generated personality model.

[0841] A "video frame" is a video frame generated in accordance with the reproduced audio, which digitally reproduces the image of the deceased.

[0842] "Emotional state" refers to the emotions the user is currently feeling, such as joy, sadness, surprise, etc., and is measured in real time using a camera, microphone, etc.

[0843] A "tablet device" is a portable, flat-shaped electronic device that is equipped with a touch screen and supports the operation of applications.

[0844] A "smartphone" is a telephone terminal that has advanced computer functions in addition to calling functions, and can be used to realize a variety of functions by installing applications.

[0845] "Personalized dialogue" refers to conversation content that is customized based on the user's individual emotional state and behavior, and is dynamically generated by a personality model of the deceased person.

[0846] A "physical store" is a commercial facility that exists physically and provides a place for consumers to visit and receive goods or services.

[0847] The system that realizes this application example collects and analyzes the communication history of the deceased person to generate a personality model and provide dialogue that corresponds to the user's emotional state. This system is realized using the following hardware and software components.

[0848] System Components

[0849] 1. Data collection method (server)

[0850] The server uses the SNS API key and login information provided by the user to collect the deceased person's communication history, such as SNS posts, call records, and message history. For example, the server obtains data using the Twitter API or Facebook Graph API.

[0851] 2. Data extraction method (server)

[0852] The server converts the collected data into an appropriate format and extracts it as text data. In this step, the JSON format data is converted into text format and each sentence is extracted.

[0853] 3. Data analysis method (server)

[0854] The server analyzes the text data using natural language processing techniques, extracts keywords and analyzes writing style, and identifies the linguistic characteristics of the deceased. Specifically, it uses Python NLP libraries (e.g., spaCy and NLTK).

[0855] 4. Personality model generation means (server)

[0856] The server generates a personality model of the deceased person, including their speech patterns and intonation, based on the identified linguistic features, using OpenAI's GPT-based generative AI model.

[0857] 5. Audio reproduction means (terminal)

[0858] The device uses a personality model downloaded from a server to reproduce the voice using voice changer technology (e.g., Respeecher or Google's Tacotron).

[0859] 6. Video frame generation means (terminal)

[0860] The device generates video frames based on the reproduced audio and provides them to the user. The video is generated using synthetic video technology (e.g., Deepfake technology).

[0861] 7. Emotion Recognition Means (Emotion Engine)

[0862] The emotion engine analyzes real-time data obtained from the user's camera and microphone to identify the user's emotional state, and is implemented using Microsoft's Azure Face API and Emotion API.

[0863] 8. Dialogue Content Adjustment Method (Server)

[0864] The server generates personalized dialogue based on the recognized emotional state, using a generative AI model to dynamically create conversational content that matches the user's emotions.

[0865] 9. Physical store devices (tablets and smartphones)

[0866] The system operates via a tablet device installed in a brick-and-mortar store or a user's smartphone, which displays audio and video to recreate a conversation with the deceased.

[0867] Specific examples

[0868] A user visiting a brick-and-mortar store (e.g., a long-established inn or restaurant) accesses a tablet device installed in the store. The server collects the communication history of the deceased person provided by the user and generates a personality model. When the user visits a specific location (e.g., a place frequently visited by the deceased), the emotion engine uses facial recognition technology to detect the user's emotional state. If the server recognizes that the user is crying, it uses the generative AI model to generate dialogue content to comfort the user and displays it on the device.

[0869] Prompt Sentence Examples

[0870] Generate it as follows: You arrive at the lobby of an inn. Noticing your tears, your deceased father speaks to you kindly, saying, "Cheer up, I'm always watching over you."

[0871] This system allows users to experience a realistic reunion with their deceased loved ones and gain emotional comfort, thus recreating memories of the deceased in a physical store and reducing the psychological burden on users.

[0872] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0873] Step 1:

[0874] The user logs in to a tablet device installed in a physical store and gives permission to collect the communication history of the deceased. The input is the SNS API key and login information provided by the user. The server uses this information to collect data such as the deceased's SNS posts, call records, and message history. The output is the acquired communication history data.

[0875] Step 2:

[0876] The server converts the collected communication history data into a text format. The input is communication history data stored in JSON format or other structured data format. The server runs a data conversion program to extract the text data. The output is a list of the actual conversations in text format.

[0877] Step 3:

[0878] The server analyzes the text data using natural language processing techniques. The input is the extracted text data. The server uses Python NLP libraries (e.g., spaCy or NLTK) to extract keywords, analyze writing style, and identify frequently occurring phrases. The output is analyzed data containing the linguistic features of the deceased.

[0879] Step 4:

[0880] The server generates a personality model of the deceased person based on the analyzed data. The input is the analyzed data, including linguistic features. The server uses OpenAI's GPT-based generative AI model to build a personality model, including speech patterns and intonation. The output is a personality model of the deceased person.

[0881] Step 5:

[0882] The device recreates the voice using a personality model downloaded from a server. The input is the personality model. The device uses voice changer technology (e.g., Respeecher or Google's Tacotron) to generate the deceased person's voice. The output is the recreated voice.

[0883] Step 6:

[0884] The device generates video frames based on the reproduced audio. The input is the reproduced audio and video data of the deceased person before their death. The device uses synthetic video technology (e.g., Deepfake technology) to generate video that matches the audio. The output is video frames.

[0885] Step 7:

[0886] The emotion engine analyzes real-time data acquired from the user's camera and microphone to identify the user's emotional state. The input is data acquired in real time of the user's facial expressions and tone of voice. The emotion engine performs emotion recognition using Microsoft's Azure Face API and Emotion API. The output is data indicating the user's emotional state.

[0887] Step 8:

[0888] The server generates dialogue content according to the recognized emotional state. The input is the user's emotional state data and a personality model of the deceased. The server uses a generative AI model to create personalized dialogue content tailored to the user's emotions. The output is the generated dialogue content.

[0889] Step 9:

[0890] The device recreates the conversation with the deceased based on the generated conversation content. The input is the generated conversation content and video frames. The device displays this to the user, providing a realistic conversation with the deceased. The output is the video and audio of the deceased displayed on the user's screen.

[0891] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0892] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0893] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0894] [Third embodiment]

[0895] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0896] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.

[0897] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0898] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0899] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0900] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0901] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0902] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0903] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0904] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0905] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0906] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0907] This system collects and analyzes the communication history of a deceased person, and generates and recreates a personality model based on the linguistic features of the person. This system allows users to virtually experience communication with the deceased person through video frames that make the deceased appear as if they were alive.

[0908] System configuration

[0909] The system of the present invention comprises the following components:

[0910] 1. Data collection method (server)

[0911] 2. Data extraction method (server)

[0912] 3. Data analysis method (server)

[0913] 4. Personality model generation means (server)

[0914] 5. Consent confirmation means (user)

[0915] 6. Audio reproduction means (terminal)

[0916] 7. Video frame generation means (terminal)

[0917] Program processing overview

[0918] Data collection method (server)

[0919] The user provides the server with the social networking service's API key and other login information, which the server uses to collect the deceased person's communication history, including social networking posts, call records, and message history.

[0920] Example: The server uses the Twitter API to retrieve the tweet history of a specified user ID.

[0921] Data extraction method (server)

[0922] The server converts the collected data into an appropriate format and extracts it as text data.

[0923] Example: A server converts JSON-formatted tweet data into a text format.

[0924] Data analysis method (server)

[0925] The server uses natural language processing technology to analyze the text data, extract keywords and analyze writing style, and identify linguistic features.

[0926] Example: The server uses an NLP library to extract frequently used phrases and their context.

[0927] Personality model generation means (server)

[0928] The server generates a personality model based on the identified linguistic features, including the deceased's speech patterns and intonation.

[0929] Example: The server models typical response patterns during a conversation based on the analyzed features.

[0930] Consent confirmation method (user)

[0931] The user must confirm whether the deceased person has agreed to use the service before they die, and if necessary, upload a consent form to the server.

[0932] Example: A user uploads a scanned consent form through a web portal.

[0933] Audio reproduction means (terminal)

[0934] The device downloads a personality model from the server and reproduces the voice using a voice changer.

[0935] Example: The device uses voice changer software to generate the voice of a deceased person.

[0936] Video frame generation means (terminal)

[0937] The terminal generates video frames based on the reproduced audio and presents them to the user.

[0938] Example: A device uses synthetic video techniques to create video that accompanies the reproduced audio.

[0939] Specific examples

[0940] For example, if a deceased person frequently used the greeting "Good morning," the server can identify that phrase and generate a video message saying "Good morning" through a voice changer on the device. Looking at the device screen, the user can feel as if the deceased person was greeting them in person.

[0941] In this way, the system of the present invention can generate and provide users with realistic interactive video frames that incorporate the memories and characteristics of the deceased, thereby helping to ease the grief of losing a loved one.

[0942] The processing flow will be explained below.

[0943] Step 1:

[0944] The user provides the server with the API key and login information of the SNS. Specifically, the user enters and submits the API key of Twitter or other SNS through a web form.

[0945] Step 2:

[0946] The server uses the provided API key to collect the post data of the deceased person from the SNS. Specifically, the server calls the Twitter API to obtain the tweet history of the specified user ID.

[0947] Step 3:

[0948] The server converts the collected data into an appropriate format and extracts it as text data. Specifically, the server converts the JSON-formatted tweet data into text format and extracts each sentence.

[0949] Step 4:

[0950] The server analyzes the extracted text data using natural language processing techniques. Specifically, the server uses an NLP library to extract frequently used phrases and keywords from the text.

[0951] Step 5:

[0952] The server then uses the text data analyzed in the previous step to identify the deceased's linguistic characteristics, such as writing style, tone, specific phrases, and intonation patterns.

[0953] Step 6:

[0954] Based on the identified features, the server generates a personality model that includes the deceased's speech patterns and intonation. Specifically, the analysis results are used to create a model that incorporates the deceased's typical speech patterns and reactions.

[0955] Step 7:

[0956] The user uploads a document to the server to confirm that they have agreed to use the service. Specifically, the user uploads a scanned copy of the consent form via a web portal.

[0957] Step 8:

[0958] The device downloads the generated personality model from the server and reproduces the voice using voice changer technology. Specifically, the device obtains the personality model via the Internet and reproduces it using voice synthesis software.

[0959] Step 9:

[0960] The device generates video frames based on the reproduced audio. Specifically, it combines the synthesized audio with footage of the deceased (previous video clips or photos) to create natural-sounding video frames.

[0961] Step 10:

[0962] The terminal provides the generated video frames to the user, specifically, displays the video frames on the terminal screen so that the user can view them.

[0963] Example 1

[0964] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0965] In modern society, there is a lack of ways to continuously remember the deceased, and it is particularly difficult to remember them through conversation or audio. Furthermore, there is no technology that can reproduce the specific linguistic characteristics and mannerisms of the deceased and imitate their personality, making it impossible to continue interacting with them. Therefore, there is a need for a way for users to experience memories of the deceased more realistically and maintain an emotional connection.

[0966] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0967] In this invention, the server includes means for collecting communication data of a user, means for extracting text data from the communication data, means for analyzing the text data using natural language processing technology and identifying linguistic characteristics of the user, means for generating a character model of the user based on the identified linguistic characteristics, means for reproducing audio based on the character model, and means for generating video frames using the reproduced audio. This allows the user to experience realistic conversations and audio with the deceased, enabling them to maintain an emotional connection with the deceased.

[0968] "User" refers to an individual or group of people who use the System.

[0969] "Communication data" refers to historical information about a user's digital communications, such as social media posts, call records, and message history.

[0970] "Text data" refers to text information extracted from communication data.

[0971] "Natural language processing technology" refers to computer technology for analyzing text data, extracting keywords, analyzing writing style, etc.

[0972] "Linguistic features" refer to the phrases, expressions, style, and other characteristics that users frequently use in a particular language.

[0973] "Person model" refers to a digital model that is generated based on a user's linguistic characteristics and mimics the user's speech patterns and intonation.

[0974] "Means for reproducing voice" refers to technologies and tools for generating voice that reproduces the characteristics of a user using a person model.

[0975] "Means for generating video frames" refers to techniques and tools for generating video frames in accordance with the reproduced audio.

[0976] "Service Use Agreement" refers to an agreement indicating that a user gives permission to use the system before death.

[0977] "Pronunciation characteristics" refers to the characteristics of a user's voice, including intonation, accent, frequently used expressions, etc.

[0978] "Frequent expressions" refer to phrases and expressions that users use frequently.

[0979] "Images" refers to your visual content, such as still images and videos.

[0980] "Synchronization" refers to the technique of playing audio and video in sync.

[0981] The present invention is a system for collecting, analyzing, and reproducing the communication history of a deceased person. Detailed procedures for implementing the present invention will be described below.

[0982] System configuration

[0983] The system of the present invention comprises the following components:

[0984] 1. Data collection means (server): Collects user communication data.

[0985] 2. Data extraction means (server): Extracts text data from the collected communication data.

[0986] 3. Data analysis means (server): Using natural language processing technology, the text data is analyzed and the user's linguistic characteristics are identified.

[0987] 4. Personality model generation means (server): Generates a personality model of the user based on the identified linguistic features.

[0988] 5. Consent confirmation means (user): Confirm whether the user agreed to use the service before death.

[0989] 6. Voice reproduction means (terminal): Reproduces the voice using the personality model downloaded from the server.

[0990] 7. Video frame generation means (terminal): Generates video frames based on the reproduced audio.

[0991] Hardware and software used

[0992] Servers: High-performance servers and cloud computing platforms (e.g., AWS, Google Cloud) are used.

[0993] Natural Language Processing library: SpaCy, NLTK, or other NLP library.

[0994] Generative AI models: Advanced AI models such as GPT-3.

[0995] Voice changer software: DeepTalk, VoxCeleb2, etc.

[0996] Facial animation technology: Synthetic video techniques such as DeepFake.

[0997] Specific step-by-step instructions

[0998] 1. Data Collection

[0999] The user provides the server with permission to access the deceased person's social media accounts and call history.

[1000] The server uses the provided API key and login information to collect communication history data from the specified account.

[1001] 2. Data extraction

[1002] The server converts the collected data into a text format, for example, extracting the message content from JSON-formatted tweet data into plain text.

[1003] 3. Data Analysis

[1004] The server uses natural language processing techniques to analyze the text data and extract keywords, specifically by identifying frequently occurring phrases and their contexts using libraries such as SpaCy and NLTK.

[1005] 4. Personality Model Generation

[1006] The server uses the identified linguistic features to generate a person model that includes the deceased person's speech patterns and intonation, using a generative AI model such as GPT-3.

[1007] 5. Confirmation of consent

[1008] The user confirms the deceased person's consent to use the service during their lifetime, and if necessary, uploads a scanned copy of the consent form to the server.

[1009] 6. Audio Reproduction

[1010] The device uses a personality model downloaded from the server and reproduces the voice using voice changer software (e.g., DeepTalk).

[1011] 7. Video Frame Generation

[1012] The device generates video frames based on the reproduced audio using facial animation technology (e.g., DeepFake).

[1013] Specific examples

[1014] For example, if the deceased frequently used the greeting "Good morning," the server can identify that expression and the device can use a voice changer to generate a video message saying "Good morning." The user can feel as if the deceased is greeting them directly through the device screen.

[1015] Prompt Sentence Examples

[1016] "Collect all tweets from user ID 'example_user'."

[1017] "Extract tweet content in text format from JSON format data."

[1018] "Extract frequently occurring words and phrases from text data."

[1019] "Use the extracted features to create a model that reproduces the speech patterns of the deceased."

[1020] "Please scan the consent form and upload it through the web portal."

[1021] "Load the model data and recreate the voice using voice changer software."

[1022] "Create video frames with facial animation technology based on the generated audio."

[1023] Although the embodiment for carrying out the present invention has been specifically described above, the present invention is not limited to this, and other embodiments and modifications are possible.

[1024] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1025] Step 1: Data collection

[1026] Users provide their social networking service API keys and login information, which the server uses to collect communication data.

[1027] Input: User's API key or login information

[1028] Output: Social media posts, call history, message history data

[1029] Specific operation: The server uses the Twitter API to retrieve the tweet history of the specified user ID in JSON format. As an example, enter the prompt "Please collect all tweets from user ID 'example_user'."

[1030] Step 2: Data extraction

[1031] The server extracts text data from the collected data.

[1032] Input: Collected social media posts, call history, and message history data

[1033] Output: Text data

[1034] Specific operation: The server extracts the text message content from the JSON format data and converts it to plain text. Specifically, it inputs the prompt "Please extract the tweet content in text format from the JSON format data."

[1035] Step 3: Data analysis

[1036] The server uses natural language processing technology to analyze the text data, extract keywords, and analyze frequently occurring phrases and writing styles.

[1037] Input: Extracted text data

[1038] Output: Linguistic features (keywords, frequent phrases, contextual information)

[1039] Specific behavior: The server uses an NLP library (e.g., SpaCy or NLTK) to extract frequent words and phrases from the text and identify the context. The server inputs the prompt: "Extract frequent words and phrases from the text data."

[1040] Step 4: Personality model generation

[1041] The server generates a personality model based on the identified linguistic features, including the deceased's speech patterns and intonation.

[1042] Input: Identified linguistic features

[1043] Output: Personality model

[1044] Specific operation: The server uses a generative AI model (e.g., GPT-3) to create a model that mimics the deceased person's speech patterns. It inputs the prompt, "Please use the extracted features to create a model that reproduces the deceased person's speech patterns."

[1045] Step 5: Confirm consent

[1046] The user checks whether the deceased person agreed to use the service before they died, and uploads a consent form to the server if necessary.

[1047] Input: Consent form (possibly scanned)

[1048] Output: Service use authorization

[1049] Specific operation: The user scans the consent form and uploads it to the server through the web portal. The server verifies and records the uploaded consent form. It then instructs the user to "scan the consent form and upload it through the web portal."

[1050] Step 6: Audio Reproduction

[1051] The device downloads a personality model from the server and reproduces the voice using voice changer software.

[1052] Input: Personality model data

[1053] Output: Reproduced audio

[1054] Specific operation: The device loads the model data downloaded from the server and generates voice using voice changer software such as DeepTalk or VoxCeleb2. The device instructs the user to "load the model data and reproduce the voice using voice changer software."

[1055] Step 7: Video Frame Generation

[1056] The terminal generates video frames in accordance with the reproduced audio and presents them to the user.

[1057] Input: Reproduced audio

[1058] Output: Video frame

[1059] Specific operation: The device uses facial animation technology (e.g., DeepFake) to create a video frame that corresponds to the reproduced audio. It then prompts the device to "create a video frame using facial animation technology based on the generated audio."

[1060] (Application example 1)

[1061] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1062] When a loved one passes away, there are limited ways to cherish memories of the deceased. Conventional technology lacks the means to reproduce the words and conversations of the deceased, in addition to visual records such as photographs and videos, making it difficult to recreate the emotional connection with the deceased. Furthermore, there is a demand for providing a natural conversation experience that includes the unique phrases and intonations used by the deceased, but no concrete means for achieving this have been established.

[1063] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1064] In this invention, the server includes means for collecting the communication history of the deceased, means for extracting text data from the communication history, and means for analyzing the text data using natural language processing technology to identify the linguistic characteristics of the deceased. This allows a personality model of the deceased to be generated based on the analyzed data, enabling the user to experience a simulated conversation with the deceased. Furthermore, the content of the conversation provided based on the generated personality model is more realistic, allowing the user to feel a deeper emotional connection with the deceased. This allows for a more realistic reproduction of memories with the deceased.

[1065] "Communication history" is a general term for records such as messages, call records, and social media posts sent and received through electronic communication means used by the deceased during their lifetime.

[1066] "Text data" refers to character information and sentences extracted from the communication history.

[1067] "Natural language processing technology" is a computer processing technology for understanding, analyzing, and generating human language, and includes a wide range of algorithms and methods for mechanically processing natural language.

[1068] "Linguistic features" refer to the characteristics of language use possessed by a particular writer or speaker, such as frequent phrases, lexical choice, and stylistic patterns.

[1069] A "personality model" is a computer model constructed to recreate the language usage and speech patterns of a deceased person based on their linguistic characteristics.

[1070] "Voice reproduction means" refers to techniques or devices that use a personality model to reproduce the voice, intonation, and distinctive speaking style of the deceased.

[1071] A "video frame" is an individual image that reproduces a specific moment in time as a video, and is the basic unit for expressing movement as a video when displayed continuously.

[1072] "Pseudo-dialogue" refers to providing a user with an experience that makes them feel as if they are conversing with the deceased person using the generated personality model.

[1073] "Dialogue content" refers to the content of words and messages exchanged between the user and the personality model of the deceased person in a simulated dialogue.

[1074] This invention is a system that collects and analyzes the communication history of a deceased person, and generates and recreates a personality model based on the linguistic features of the communication history. This system allows users to enjoy a virtual conversation experience in which the deceased person feels as if they are alive. This system incorporates key elements such as "communication history," "text data," "natural language processing technology," "linguistic features," "personality model," "means for reproducing voice," "video frames," "simulated dialogue," and "dialogue content."

[1075] Hardware and software used

[1076] Hardware: Smartphone (iPhone or Android)

[1077] Software: Python, Flask, TensorFlow / Keras, Pandas, NLTK, Twilio API, Google Cloud Text-to-Speech API

[1078] Program processing overview

[1079] 1. Data collection method (server)

[1080] The server receives the API key of the deceased person's social media account from the user. Using this API key, the server accesses the deceased person's social media account and collects their communication history, including Twitter and Facebook posts, messages, and call records.

[1081] 2. Data extraction method (server)

[1082] The server converts the collected social media data into an appropriate format and extracts it as text data. The collected data is provided in JSON format, but it is converted into an easily handled text format.

[1083] 3. Data analysis method (server)

[1084] The server analyzes the extracted text data using natural language processing libraries such as Pandas and NLTK to identify the linguistic characteristics of the deceased and extract frequent phrases, expressions, and writing style.

[1085] 4. Personality model generation means (server)

[1086] Based on the analyzed linguistic features, the server uses TensorFlow and Keras to generate a personality model of the deceased, a computer model that reproduces the speech patterns and intonation of the deceased.

[1087] 5. Consent confirmation means (user)

[1088] Users must verify through a web portal whether the deceased person consented to use the service before they died, and then upload a scanned copy of the consent form.

[1089] 6. Audio reproduction means (terminal)

[1090] The device uses a personality model downloaded from the server to reproduce the voice of the deceased person using voice synthesis technology such as the Google Cloud Text-to-Speech API.

[1091] 7. Video frame generation means (terminal)

[1092] As the audio plays, the device uses synthetic video technology to generate a video of the deceased, providing the user with realistic video frames synchronized with the audio.

[1093] Specific examples

[1094] For example, if a deceased person frequently used the greeting "Good morning," the server can identify this phrase and generate a video message in which the device says "Good morning" through a voice changer. The following is an example of a prompt sentence:

[1095] Example prompt sentence:

[1096] Prompt: "Good morning. What are you planning to do today?"

[1097] Response: The generated personality model responds, "Good morning! I'm planning to go for a walk today. How about you?"

[1098] In this way, the user can experience memories of the deceased more realistically through simulated dialogue.

[1099] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1100] Step 1:

[1101] Data collection method (server)

[1102] The server receives the API key and login information of the deceased person's SNS from the user. The server then collects the deceased person's communication history (messages, posts, call records, etc.) from various SNS platforms (e.g., Twitter, Facebook). The input data is the API key and user ID, and the output is communication history data in JSON format.

[1103] Step 2:

[1104] Data extraction method (server)

[1105] The server converts the collected SNS data from an appropriate format (for example, JSON format) into text data. Here, the tweet text and message content are extracted and organized as text data. The input is communication history data in JSON format, and the output is formatted text data. Specifically, the server extracts the necessary information from the data obtained from each SNS platform and converts it into text format.

[1106] Step 3:

[1107] Data analysis method (server)

[1108] The server analyzes the text data using natural language processing techniques (e.g., Pandas or NLTK). The analysis identifies the linguistic characteristics of the deceased (frequent phrases, writing style, intonation, etc.). The input is text data, and the output is metadata about the linguistic characteristics of the deceased. Specific operations include word frequency analysis, morphological analysis, and contextual analysis.

[1109] Step 4:

[1110] Personality model generation means (server)

[1111] Based on the analysis results, the server uses TensorFlow and Keras to generate a personality model of the deceased. This is a machine learning model that learns the deceased's speech patterns. The input is metadata about the linguistic features of the deceased, and the output is a trained personality model. Specifically, it trains an LSTM model using the deceased's speech data.

[1112] Step 5:

[1113] Consent confirmation method (user)

[1114] The user checks through the web portal whether the deceased person had consented to use the service before they died, and then uploads a scanned copy of the consent form to the server. The input is the scanned consent form (electronic file), and the output is the result of uploading the consent form to the server. Specifically, the user selects the consent form on the portal site and presses the upload button.

[1115] Step 6:

[1116] Audio reproduction means (terminal)

[1117] The device uses the personality model downloaded from the server and generates the voice of the deceased person using the Google Cloud Text-to-Speech API, etc. The input is the personality model and text data, and the output is an audio file. Specifically, it converts the text-generated dialogue into audio.

[1118] Step 7:

[1119] Video frame generation means (terminal)

[1120] The device uses synthetic video technology to generate a video of the deceased person in sync with the audio. The input is the generated audio and video data of the deceased person while they were alive, and the output is video frames synchronized with the audio. Specifically, the device generates and edits a video sequence corresponding to the audio.

[1121] This allows the user to experience a simulated conversation with the deceased, allowing memories of the deceased to be recreated more realistically.

[1122] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1123] This system collects and analyzes the communication history of a deceased person, and generates and recreates a personality model based on the linguistic features. This system also incorporates an emotion engine that recognizes the user's emotions, providing dialogue that responds to the user's emotions, thereby realizing a more realistic communication experience.

[1124] System configuration

[1125] The system of the present invention comprises the following components:

[1126] 1. Data collection method (server)

[1127] 2. Data extraction method (server)

[1128] 3. Data analysis method (server)

[1129] 4. Personality model generation means (server)

[1130] 5. Consent confirmation means (user)

[1131] 6. Audio reproduction means (terminal)

[1132] 7. Video frame generation means (terminal)

[1133] 8. Emotion Recognition Means (Emotion Engine)

[1134] Program processing overview

[1135] Data collection method (server)

[1136] The user provides the server with the API key and login information of the social networking site, which the server uses to collect the deceased person's communication history, including social networking site posts, call records, and message history.

[1137] Example: The server uses the Twitter API to retrieve the tweet history of a specified user ID.

[1138] Data extraction method (server)

[1139] The server converts the collected data into an appropriate format and extracts it as text data.

[1140] Example: The server converts JSON-formatted tweet data into text format and extracts each sentence.

[1141] Data analysis method (server)

[1142] The server uses natural language processing technology to analyze the text data, extract keywords and analyze writing style, and identify linguistic features.

[1143] Example: A server uses an NLP library to extract frequently used phrases and keywords from text.

[1144] Personality model generation means (server)

[1145] The server generates a personality model based on the identified linguistic features, including the deceased's speech patterns and intonation.

[1146] Example: The server models typical response patterns during a conversation based on the analyzed features.

[1147] Consent confirmation method (user)

[1148] A document is uploaded to the server to confirm that the user has agreed to use the service.

[1149] Example: A user uploads a scanned consent form through a web portal.

[1150] Audio reproduction means (terminal)

[1151] The device downloads the generated personality model from the server and reproduces the voice using voice changer technology.

[1152] Example: The device uses voice changer software to generate the voice of a deceased person.

[1153] Video frame generation means (terminal)

[1154] The terminal generates video frames based on the reproduced audio and presents them to the user.

[1155] Example: A device uses synthetic video techniques to create video that accompanies the reproduced audio.

[1156] Emotion recognition means (emotion engine)

[1157] The emotion engine analyzes real-time data obtained from the user's camera and microphone to identify the user's emotional state.

[1158] Example: An emotion engine analyzes a user's facial expressions and tone of voice to determine whether they are happy, sad, or surprised.

[1159] How the Emotion Engine Works

[1160] The emotion engine provides a means to recognize the user's emotional state in real time and adjust the audio and video frames of the deceased based on that. For example, if the emotion engine determines that the user is sad, it can generate a soothing voice or comforting message for the deceased. This functionality allows the user to feel more connected to the deceased.

[1161] Specific examples

[1162] For example, if the deceased frequently used the greeting "good morning," the emotion engine may determine that the user is feeling unwell. The emotion engine will adjust the reproduced voice to be more gentle and comforting depending on the situation, and change the video frame accordingly. Looking at the device screen, the user will feel as if the deceased is gently encouraging them.

[1163] In this way, the system of the present invention can generate and provide users with realistic interactive video frames that incorporate the memories and characteristics of the deceased, enabling personalized interactions based on the user's emotional state, helping to alleviate the grief of losing a loved one.

[1164] The processing flow will be explained below.

[1165] Step 1:

[1166] The user provides the server with the API key and login information of the SNS. Specifically, the user enters the API key of Twitter or other SNS through a web form and submits it.

[1167] Step 2:

[1168] The server uses the provided API key to collect the post data of the deceased person from the SNS. Specifically, the server calls the Twitter API to obtain the tweet history of the specified user ID.

[1169] Step 3:

[1170] The server converts the collected data into an appropriate format and extracts it as text data. Specifically, the server converts the JSON-formatted tweet data into text format and extracts each sentence.

[1171] Step 4:

[1172] The server analyzes the extracted text data using natural language processing techniques. Specifically, the server uses an NLP library to extract frequently used phrases and keywords from the text.

[1173] Step 5:

[1174] The server then uses the text data analyzed in the previous step to identify the deceased's linguistic characteristics, such as writing style, tone, specific phrases, and intonation patterns.

[1175] Step 6:

[1176] Based on the identified characteristics, the server generates a personality model that includes the deceased's speech patterns and intonation. Specifically, it uses the analysis results to incorporate the deceased's frequently used phrases and reaction patterns into the model.

[1177] Step 7:

[1178] The user uploads a document to the server to confirm that they have agreed to use the service. Specifically, the user uploads a scanned copy of the consent form via a web portal.

[1179] Step 8:

[1180] The device downloads the generated personality model from the server and uses voice changer technology to recreate the voice of the deceased. Specifically, the device obtains the personality model via the internet and uses voice synthesis software to generate the voice of the deceased.

[1181] Step 9:

[1182] The device generates video frames based on the reproduced audio. Specifically, it combines the synthesized audio with footage of the deceased (previous video clips or photos) to create natural-sounding video frames.

[1183] Step 10:

[1184] The terminal displays the video frames to the user, specifically, displays the video frames on the terminal's display so that the user can view them.

[1185] Step 11:

[1186] The emotion engine captures real-time data from the user's camera and microphone and analyzes their emotions by analyzing their facial expressions, tone of voice, and word choice to determine their current emotional state.

[1187] Step 12:

[1188] Based on the analysis, the emotion engine adjusts the content of the deceased person's audio and video frames. For example, if the user is sad, the message of the deceased will be changed to be more comforting.

[1189] Step 13:

[1190] The device plays the adjusted content and displays it to the user. Specifically, the content generated based on instructions from the emotion engine is displayed on the display, and the user watches and listens to it.

[1191] Example 2

[1192] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1193] Although there are systems that can recreate the memories and characteristics of the deceased and allow users to have a realistic conversation with them, these systems lack the ability to recognize the user's emotional state and dynamically adjust the conversation content based on that emotion. This limits the conversational experiences provided, making it difficult for users to feel a deeper connection with the deceased.

[1194] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1195] In this invention, the server includes means for collecting the communication history of the deceased, means for extracting text data from the communication history, means for analyzing the text data using natural language processing technology and identifying linguistic characteristics of the deceased, means for generating a personality model of the deceased based on the identified linguistic characteristics, means for downloading the personality model of the deceased from the server to a terminal, means for reproducing voice based on the personality model, means for generating video frames based on the reproduced voice, means for recognizing the emotional state of the user, and means for adjusting the voice and video frames based on the emotional state of the user. This allows the user to not only realistically experience memories and characteristics of the deceased, but also to have a personalized interaction experience according to emotions.

[1196] "Communication history" refers to a series of communication data such as social media posts, call records, and message history made by the deceased person while they were alive.

[1197] "Text data" refers to data in sentence format extracted from the communication history.

[1198] "Natural language processing technology" refers to technology used by computers to understand, analyze, and generate human language.

[1199] "Linguistic characteristics" refer to linguistic characteristics and patterns, such as the particular words, writing style, and frequently used phrases used by a particular person.

[1200] A "personality model" is a data model designed to reproduce the linguistic characteristics, speech patterns, intonation, etc. of a deceased person.

[1201] "Voice reproduction" refers to the technique or process of recreating the voice of a deceased person based on a generated personality model.

[1202] A "video frame" is a frame of video that corresponds to the reproduced audio and shows the image and gestures of the deceased.

[1203] "Emotional state" refers to a user's current psychological and sensory state, which is typically determined through facial expressions, tone of voice, etc.

[1204] "Emotion recognition" refers to the technology of analyzing a user's emotional state and thereby identifying a specific emotion.

[1205] A "personalized interaction experience" is an interaction that is optimized for a user's unique characteristics and situation, and is adjusted based on their emotional state.

[1206] This system collects and analyzes the communication history of a deceased person, and generates and recreates a personality model based on the linguistic features. This system also incorporates an emotion engine that recognizes the user's emotions, providing dialogue that responds to the user's emotions, thereby realizing a more realistic communication experience.

[1207] System configuration

[1208] The system of the present invention comprises the following components:

[1209] 1. Data collection method (server)

[1210] 2. Data extraction method (server)

[1211] 3. Data analysis method (server)

[1212] 4. Personality model generation means (server)

[1213] 5. Consent confirmation means (user)

[1214] 6. Audio reproduction means (terminal)

[1215] 7. Video frame generation means (terminal)

[1216] 8. Emotion Recognition Means (Emotion Engine)

[1217] Detailed system configuration and processing method

[1218] Data collection method (server)

[1219] The user enters their social networking service API key and login information, which is then sent in a secure format to the server, which then uses this information to collect the deceased person's communication history, including social networking posts, call logs, and message history. This process uses standard APIs, such as the Twitter API and Facebook API.

[1220] Example: The server uses the Twitter API to retrieve the tweet history of a specified user ID.

[1221] Data extraction method (server)

[1222] The server analyzes the collected data and extracts it as text data. At this time, it converts the necessary parts of the data in JSON or XML format into text format and extracts each sentence and word.

[1223] Example: The server converts JSON-formatted tweet data into text format and extracts each sentence.

[1224] Data analysis method (server)

[1225] The server uses natural language processing (NLP) techniques to analyze the extracted text data. During the analysis, it identifies linguistic features such as keywords, frequent phrases, and writing style. This process uses NLP libraries (e.g., NLTK, spacy).

[1226] Example: The server uses an NLP library to extract frequently used phrases and keywords from the text.

[1227] Personality model generation means (server)

[1228] The server generates a personality model based on the analyzed linguistic features, including the deceased's speech patterns and intonation, which is then used to simulate conversations similar to those of the deceased.

[1229] Example: The server uses the analyzed features to model typical response patterns during a conversation.

[1230] Consent confirmation method (user)

[1231] The user uploads a consent form for using the service to the server in PDF format, etc. The server then checks the uploaded consent form and determines whether or not the use of the service is permitted.

[1232] Example: User uploads scanned consent form through web portal.

[1233] Audio reproduction means (terminal)

[1234] The device uses a personality model downloaded from the server and uses voice changer technology (e.g., Voicemod, iZotope) to recreate the voice of the deceased person, which is used during interactions with the user.

[1235] Example: The device uses voice changer software to generate the voice of the deceased person.

[1236] Video frame generation means (terminal)

[1237] Based on the reproduced audio, the device uses synthesis technology (e.g., D-ID, Reallusion) to generate video frames that display the deceased person's face and gestures in sync with the audio.

[1238] Example: The device uses synthetic video techniques to create a video that matches the reproduced audio.

[1239] Emotion recognition means (emotion engine)

[1240] The emotion engine analyzes real-time data captured from the user's camera and microphone. It uses facial expression analysis software (e.g., Affectiva, Microsoft Azure Face API) and voice analysis tools to identify the user's emotional state.

[1241] Example: An emotion engine analyzes a user's facial expressions and tone of voice to determine whether they are sad, happy, or surprised.

[1242] Examples of concrete examples and prompts

[1243] For example, if the deceased frequently used the greeting "Good morning," the emotion engine may determine that the user's emotional state was poor. In this case, the emotion engine will adjust the reproduced voice to be calm and comforting, and change the video frame accordingly. Looking at the device screen, the user will feel as if the deceased is gently encouraging them.

[1244] Example prompt sentence:

[1245] It uses phrases that allow users to start a natural conversation, such as "Hi, how are you doing?" or "How was your day?"

[1246] In this way, the system of the present invention can realistically recreate the memories and characteristics of the deceased and provide a personalized interaction experience for the user.

[1247] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1248] Step 1: Enter user information

[1249] The user enters the SNS API key and login information, which is then sent to the server in a secure format. Specifically, the user enters the information on the system's login screen and presses the send button. The input data is the SNS API key and login information, and the output is the API key and login information sent to the server.

[1250] Step 2: Data collection

[1251] The server uses the API key and login information sent by the user to collect the deceased person's communication history, such as social media posts, call records, and message history. Specifically, the server uses an appropriate API (e.g., Twitter API, Facebook API) to obtain data for the specified user ID. The input data is the API key and login information provided by the user in the previous step, and the output is the collected communication history (e.g., social media post data, call records).

[1252] Step 3: Data extraction

[1253] The server converts the collected communication history into an appropriate format and extracts it as text data. Specifically, the server converts SNS data in JSON or XML format into text format and extracts each sentence. The input data is the communication history (e.g., tweet data in JSON format), and the output is the extracted text data (e.g., single sentences of text).

[1254] Step 4: Data analysis

[1255] The server uses natural language processing (NLP) techniques to analyze the extracted text data. During the analysis process, it identifies linguistic features such as keywords, frequently occurring phrases, and writing style. Specifically, the server uses an NLP library (e.g., NLTK, spacy) to extract frequently used phrases and keywords from the text. The input data is the text data, and the output is the analyzed linguistic features (e.g., a list of keywords and phrases).

[1256] Step 5: Personality model generation

[1257] The server generates a personality model based on the analyzed linguistic features, including the deceased's speech patterns and intonation. Specifically, the server uses a machine learning algorithm to model typical response patterns during conversations based on the analyzed features. The input data are the linguistic features, and the output is the generated personality model.

[1258] Step 6: Confirm consent

[1259] The user uploads the consent form for using the service to the server in PDF format or other format. Specifically, the user scans the consent form and uploads it through a web portal. The input data is the consent form in PDF format, and the output is a consent confirmation flag on the server.

[1260] Step 7: Audio Reproduction

[1261] The device uses a personality model downloaded from the server and voice changer technology (e.g., Voicemod, iZotope) to recreate the deceased's voice. Specifically, the device converts the deceased's voice pattern into audio data based on the downloaded personality model. The input data is the personality model, and the output is the recreated audio data.

[1262] Step 8: Video Frame Generation

[1263] The device generates video frames based on the reproduced audio. Specifically, the device uses synthesis technology (e.g., D-ID, Reallusion) to create video that matches the reproduced audio. The input data is the reproduced audio data, and the output is the generated video frames.

[1264] Step 9: Emotion Recognition

[1265] The emotion engine analyzes real-time data acquired from the user's camera and microphone. Specifically, the emotion engine uses facial expression analysis software (e.g., Affectiva, Microsoft Azure Face API) and voice analysis tools to identify the user's emotional state. The input data is the real-time data acquired from the camera and microphone, and the output is the identified user's emotional state.

[1266] Step 10: Dialogue Adjustment

[1267] The emotion engine appropriately adjusts the content and tone of the dialogue based on the emotion data acquired in the previous stage. For example, if the user is sad, the emotion engine generates a calm voice and comforting words, and optimizes the video frames accordingly. The input data is the user's emotional state, and the output is the adjusted voice data and video frames.

[1268] (Application example 2)

[1269] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1270] There is a need for a means to maintain the emotional connection between the deceased and those left behind, particularly to ease the grief of losing the deceased. However, conventional technologies only reproduce the personality of the deceased, and are unable to engage in dialogue that reflects the user's emotional state, making it difficult to provide a realistic communication experience. Furthermore, there is no system in physical stores that allows users to feel emotionally reassured, making it difficult to alleviate the user's psychological burden.

[1271] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1272] In this invention, the server includes means for collecting the communication history of the deceased, means for extracting text data from the communication history, means for analyzing the text data using natural language processing technology to identify the linguistic characteristics of the deceased, means for generating a personality model of the deceased based on the identified linguistic characteristics, means for reproducing voice based on the personality model, means for generating video frames using the reproduced voice, means for recognizing the emotional state of a user in real time, means for adjusting the content of the dialogue in accordance with the recognized emotional state, and means for reproducing a dialogue with the deceased via a tablet terminal or smartphone in a physical store. This allows for a realistic reproduction of the personality of the deceased while enabling personalized dialogue in accordance with the user's emotional state, thereby providing users with a sense of emotional security even in a physical store.

[1273] "Communication history" is a record of past interactions recorded as digital data, such as text messages, call logs, and social media posts, across the mediums used by individuals to communicate.

[1274] "Natural language processing technology" is a field of artificial intelligence that enables computers to understand, interpret, and generate human language, and is a technology that analyzes text and extracts linguistic features.

[1275] A "personality model" is a digital recreation of a specific individual's linguistic characteristics, speech patterns, and presence, constructed from the deceased's text data and call records.

[1276] "Voice reproduction" is a technology that generates voice by imitating the speaking style and voice quality of a deceased person based on a generated personality model.

[1277] A "video frame" is a video frame generated in accordance with the reproduced audio, which digitally reproduces the image of the deceased.

[1278] "Emotional state" refers to the emotions the user is currently feeling, such as joy, sadness, surprise, etc., and is measured in real time using a camera, microphone, etc.

[1279] A "tablet device" is a portable, flat-shaped electronic device that is equipped with a touch screen and supports the operation of applications.

[1280] A "smartphone" is a telephone terminal that has advanced computer functions in addition to calling functions, and can be used to realize a variety of functions by installing applications.

[1281] "Personalized dialogue" refers to conversation content that is customized based on the user's individual emotional state and behavior, and is dynamically generated by a personality model of the deceased person.

[1282] A "physical store" is a commercial facility that exists physically and provides a place for consumers to visit and receive goods or services.

[1283] The system that realizes this application example collects and analyzes the communication history of the deceased person to generate a personality model and provide dialogue that corresponds to the user's emotional state. This system is realized using the following hardware and software components.

[1284] System Components

[1285] 1. Data collection method (server)

[1286] The server uses the SNS API key and login information provided by the user to collect the deceased person's communication history, such as SNS posts, call records, and message history. For example, the server obtains data using the Twitter API or Facebook Graph API.

[1287] 2. Data extraction method (server)

[1288] The server converts the collected data into an appropriate format and extracts it as text data. In this step, the JSON format data is converted into text format and each sentence is extracted.

[1289] 3. Data analysis method (server)

[1290] The server analyzes the text data using natural language processing techniques, extracts keywords and analyzes writing style, and identifies the linguistic characteristics of the deceased. Specifically, it uses Python NLP libraries (e.g., spaCy and NLTK).

[1291] 4. Personality model generation means (server)

[1292] The server generates a personality model of the deceased person, including their speech patterns and intonation, based on the identified linguistic features, using OpenAI's GPT-based generative AI model.

[1293] 5. Audio reproduction means (terminal)

[1294] The device uses a personality model downloaded from a server to reproduce the voice using voice changer technology (e.g., Respeecher or Google's Tacotron).

[1295] 6. Video frame generation means (terminal)

[1296] The device generates video frames based on the reproduced audio and provides them to the user. The video is generated using synthetic video technology (e.g., Deepfake technology).

[1297] 7. Emotion Recognition Means (Emotion Engine)

[1298] The emotion engine analyzes real-time data obtained from the user's camera and microphone to identify the user's emotional state, and is implemented using Microsoft's Azure Face API and Emotion API.

[1299] 8. Dialogue Content Adjustment Method (Server)

[1300] The server generates personalized dialogue based on the recognized emotional state, using a generative AI model to dynamically create conversational content that matches the user's emotions.

[1301] 9. Physical store devices (tablets and smartphones)

[1302] The system operates via a tablet device installed in a brick-and-mortar store or a user's smartphone, which displays audio and video to recreate a conversation with the deceased.

[1303] Specific examples

[1304] A user visiting a brick-and-mortar store (e.g., a long-established inn or restaurant) accesses a tablet device installed in the store. The server collects the communication history of the deceased person provided by the user and generates a personality model. When the user visits a specific location (e.g., a place frequently visited by the deceased), the emotion engine uses facial recognition technology to detect the user's emotional state. If the server recognizes that the user is crying, it uses the generative AI model to generate dialogue content to comfort the user and displays it on the device.

[1305] Prompt Sentence Examples

[1306] Generate it as follows: You arrive at the lobby of an inn. Noticing your tears, your deceased father speaks to you kindly, saying, "Cheer up, I'm always watching over you."

[1307] This system allows users to experience a realistic reunion with their deceased loved ones and gain emotional comfort, thus recreating memories of the deceased in a physical store and reducing the psychological burden on users.

[1308] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1309] Step 1:

[1310] The user logs in to a tablet device installed in a physical store and gives permission to collect the communication history of the deceased. The input is the SNS API key and login information provided by the user. The server uses this information to collect data such as the deceased's SNS posts, call records, and message history. The output is the acquired communication history data.

[1311] Step 2:

[1312] The server converts the collected communication history data into a text format. The input is communication history data stored in JSON format or other structured data format. The server runs a data conversion program to extract the text data. The output is a list of the actual conversations in text format.

[1313] Step 3:

[1314] The server analyzes the text data using natural language processing techniques. The input is the extracted text data. The server uses Python NLP libraries (e.g., spaCy or NLTK) to extract keywords, analyze writing style, and identify frequently occurring phrases. The output is analyzed data containing the linguistic features of the deceased.

[1315] Step 4:

[1316] The server generates a personality model of the deceased person based on the analyzed data. The input is the analyzed data, including linguistic features. The server uses OpenAI's GPT-based generative AI model to build a personality model, including speech patterns and intonation. The output is a personality model of the deceased person.

[1317] Step 5:

[1318] The device recreates the voice using a personality model downloaded from a server. The input is the personality model. The device uses voice changer technology (e.g., Respeecher or Google's Tacotron) to generate the deceased person's voice. The output is the recreated voice.

[1319] Step 6:

[1320] The device generates video frames based on the reproduced audio. The input is the reproduced audio and video data of the deceased person before their death. The device uses synthetic video technology (e.g., Deepfake technology) to generate video that matches the audio. The output is video frames.

[1321] Step 7:

[1322] The emotion engine analyzes real-time data acquired from the user's camera and microphone to identify the user's emotional state. The input is data acquired in real time of the user's facial expressions and tone of voice. The emotion engine performs emotion recognition using Microsoft's Azure Face API and Emotion API. The output is data indicating the user's emotional state.

[1323] Step 8:

[1324] The server generates dialogue content according to the recognized emotional state. The input is the user's emotional state data and a personality model of the deceased. The server uses a generative AI model to create personalized dialogue content tailored to the user's emotions. The output is the generated dialogue content.

[1325] Step 9:

[1326] The device recreates the conversation with the deceased based on the generated conversation content. The input is the generated conversation content and video frames. The device displays this to the user, providing a realistic conversation with the deceased. The output is the video and audio of the deceased displayed on the user's screen.

[1327] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1328] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1329] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1330] [Fourth embodiment]

[1331] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1332] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1333] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1334] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1335] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1336] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1337] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1338] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1339] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1340] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1341] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1342] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1343] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1344] This system collects and analyzes the communication history of a deceased person, and generates and recreates a personality model based on the linguistic features of the person. This system allows users to virtually experience communication with the deceased person through video frames that make the deceased appear as if they were alive.

[1345] System configuration

[1346] The system of the present invention comprises the following components:

[1347] 1. Data collection method (server)

[1348] 2. Data extraction method (server)

[1349] 3. Data analysis method (server)

[1350] 4. Personality model generation means (server)

[1351] 5. Consent confirmation means (user)

[1352] 6. Audio reproduction means (terminal)

[1353] 7. Video frame generation means (terminal)

[1354] Program processing overview

[1355] Data collection method (server)

[1356] The user provides the server with the social networking service's API key and other login information, which the server uses to collect the deceased person's communication history, including social networking posts, call records, and message history.

[1357] Example: The server uses the Twitter API to retrieve the tweet history of a specified user ID.

[1358] Data extraction method (server)

[1359] The server converts the collected data into an appropriate format and extracts it as text data.

[1360] Example: A server converts JSON-formatted tweet data into a text format.

[1361] Data analysis method (server)

[1362] The server uses natural language processing technology to analyze the text data, extract keywords and analyze writing style, and identify linguistic features.

[1363] Example: The server uses an NLP library to extract frequently used phrases and their context.

[1364] Personality model generation means (server)

[1365] The server generates a personality model based on the identified linguistic features, including the deceased's speech patterns and intonation.

[1366] Example: The server models typical response patterns during a conversation based on the analyzed features.

[1367] Consent confirmation method (user)

[1368] The user must confirm whether the deceased person has agreed to use the service before they die, and if necessary, upload a consent form to the server.

[1369] Example: A user uploads a scanned consent form through a web portal.

[1370] Audio reproduction means (terminal)

[1371] The device downloads a personality model from the server and reproduces the voice using a voice changer.

[1372] Example: The device uses voice changer software to generate the voice of a deceased person.

[1373] Video frame generation means (terminal)

[1374] The terminal generates video frames based on the reproduced audio and presents them to the user.

[1375] Example: A device uses synthetic video techniques to create video that accompanies the reproduced audio.

[1376] Specific examples

[1377] For example, if a deceased person frequently used the greeting "Good morning," the server can identify that phrase and generate a video message saying "Good morning" through a voice changer on the device. Looking at the device screen, the user can feel as if the deceased person was greeting them in person.

[1378] In this way, the system of the present invention can generate and provide users with realistic interactive video frames that incorporate the memories and characteristics of the deceased, thereby helping to ease the grief of losing a loved one.

[1379] The processing flow will be explained below.

[1380] Step 1:

[1381] The user provides the server with the API key and login information of the SNS. Specifically, the user enters and submits the API key of Twitter or other SNS through a web form.

[1382] Step 2:

[1383] The server uses the provided API key to collect the post data of the deceased person from the SNS. Specifically, the server calls the Twitter API to obtain the tweet history of the specified user ID.

[1384] Step 3:

[1385] The server converts the collected data into an appropriate format and extracts it as text data. Specifically, the server converts the JSON-formatted tweet data into text format and extracts each sentence.

[1386] Step 4:

[1387] The server analyzes the extracted text data using natural language processing techniques. Specifically, the server uses an NLP library to extract frequently used phrases and keywords from the text.

[1388] Step 5:

[1389] The server then uses the text data analyzed in the previous step to identify the deceased's linguistic characteristics, such as writing style, tone, specific phrases, and intonation patterns.

[1390] Step 6:

[1391] Based on the identified features, the server generates a personality model that includes the deceased's speech patterns and intonation. Specifically, the analysis results are used to create a model that incorporates the deceased's typical speech patterns and reactions.

[1392] Step 7:

[1393] The user uploads a document to the server to confirm that they have agreed to use the service. Specifically, the user uploads a scanned copy of the consent form via a web portal.

[1394] Step 8:

[1395] The device downloads the generated personality model from the server and reproduces the voice using voice changer technology. Specifically, the device obtains the personality model via the Internet and reproduces it using voice synthesis software.

[1396] Step 9:

[1397] The device generates video frames based on the reproduced audio. Specifically, it combines the synthesized audio with footage of the deceased (previous video clips or photos) to create natural-sounding video frames.

[1398] Step 10:

[1399] The terminal provides the generated video frames to the user, specifically, displays the video frames on the terminal screen so that the user can view them.

[1400] Example 1

[1401] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1402] In modern society, there is a lack of ways to continuously remember the deceased, and it is particularly difficult to remember them through conversation or audio. Furthermore, there is no technology that can reproduce the specific linguistic characteristics and mannerisms of the deceased and imitate their personality, making it impossible to continue interacting with them. Therefore, there is a need for a way for users to experience memories of the deceased more realistically and maintain an emotional connection.

[1403] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1404] In this invention, the server includes means for collecting communication data of a user, means for extracting text data from the communication data, means for analyzing the text data using natural language processing technology and identifying linguistic characteristics of the user, means for generating a character model of the user based on the identified linguistic characteristics, means for reproducing audio based on the character model, and means for generating video frames using the reproduced audio. This allows the user to experience realistic conversations and audio with the deceased, enabling them to maintain an emotional connection with the deceased.

[1405] "User" refers to an individual or group of people who use the System.

[1406] "Communication data" refers to historical information about a user's digital communications, such as social media posts, call records, and message history.

[1407] "Text data" refers to text information extracted from communication data.

[1408] "Natural language processing technology" refers to computer technology for analyzing text data, extracting keywords, analyzing writing style, etc.

[1409] "Linguistic features" refer to the phrases, expressions, style, and other characteristics that users frequently use in a particular language.

[1410] "Person model" refers to a digital model that is generated based on a user's linguistic characteristics and mimics the user's speech patterns and intonation.

[1411] "Means for reproducing voice" refers to technologies and tools for generating voice that reproduces the characteristics of a user using a person model.

[1412] "Means for generating video frames" refers to techniques and tools for generating video frames in accordance with the reproduced audio.

[1413] "Service Use Agreement" refers to an agreement indicating that a user gives permission to use the system before death.

[1414] "Pronunciation characteristics" refers to the characteristics of a user's voice, including intonation, accent, frequently used expressions, etc.

[1415] "Frequent expressions" refer to phrases and expressions that users use frequently.

[1416] "Images" refers to your visual content, such as still images and videos.

[1417] "Synchronization" refers to the technique of playing audio and video in sync.

[1418] The present invention is a system for collecting, analyzing, and reproducing the communication history of a deceased person. Detailed procedures for implementing the present invention will be described below.

[1419] System configuration

[1420] The system of the present invention comprises the following components:

[1421] 1. Data collection means (server): Collects user communication data.

[1422] 2. Data extraction means (server): Extracts text data from the collected communication data.

[1423] 3. Data analysis means (server): Using natural language processing technology, the text data is analyzed and the user's linguistic characteristics are identified.

[1424] 4. Personality model generation means (server): Generates a personality model of the user based on the identified linguistic features.

[1425] 5. Consent confirmation means (user): Confirm whether the user agreed to use the service before death.

[1426] 6. Voice reproduction means (terminal): Reproduces the voice using the personality model downloaded from the server.

[1427] 7. Video frame generation means (terminal): Generates video frames based on the reproduced audio.

[1428] Hardware and software used

[1429] Servers: High-performance servers and cloud computing platforms (e.g., AWS, Google Cloud) are used.

[1430] Natural Language Processing library: SpaCy, NLTK, or other NLP library.

[1431] Generative AI models: Advanced AI models such as GPT-3.

[1432] Voice changer software: DeepTalk, VoxCeleb2, etc.

[1433] Facial animation technology: Synthetic video techniques such as DeepFake.

[1434] Specific step-by-step instructions

[1435] 1. Data Collection

[1436] The user provides the server with permission to access the deceased person's social media accounts and call history.

[1437] The server uses the provided API key and login information to collect communication history data from the specified account.

[1438] 2. Data extraction

[1439] The server converts the collected data into a text format, for example, extracting the message content from JSON-formatted tweet data into plain text.

[1440] 3. Data Analysis

[1441] The server uses natural language processing techniques to analyze the text data and extract keywords, specifically by identifying frequently occurring phrases and their contexts using libraries such as SpaCy and NLTK.

[1442] 4. Personality Model Generation

[1443] The server uses the identified linguistic features to generate a person model that includes the deceased person's speech patterns and intonation, using a generative AI model such as GPT-3.

[1444] 5. Confirmation of consent

[1445] The user confirms the deceased person's consent to use the service during their lifetime, and if necessary, uploads a scanned copy of the consent form to the server.

[1446] 6. Audio Reproduction

[1447] The device uses a personality model downloaded from the server and reproduces the voice using voice changer software (e.g., DeepTalk).

[1448] 7. Video Frame Generation

[1449] The device generates video frames based on the reproduced audio using facial animation technology (e.g., DeepFake).

[1450] Specific examples

[1451] For example, if the deceased frequently used the greeting "Good morning," the server can identify that expression and the device can use a voice changer to generate a video message saying "Good morning." The user can feel as if the deceased is greeting them directly through the device screen.

[1452] Prompt Sentence Examples

[1453] "Collect all tweets from user ID 'example_user'."

[1454] "Extract tweet content in text format from JSON format data."

[1455] "Extract frequently occurring words and phrases from text data."

[1456] "Use the extracted features to create a model that reproduces the speech patterns of the deceased."

[1457] "Please scan the consent form and upload it through the web portal."

[1458] "Load the model data and recreate the voice using voice changer software."

[1459] "Create video frames with facial animation technology based on the generated audio."

[1460] Although the embodiment for carrying out the present invention has been specifically described above, the present invention is not limited to this, and other embodiments and modifications are possible.

[1461] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1462] Step 1: Data collection

[1463] Users provide their social networking service API keys and login information, which the server uses to collect communication data.

[1464] Input: User's API key or login information

[1465] Output: Social media posts, call history, message history data

[1466] Specific operation: The server uses the Twitter API to retrieve the tweet history of the specified user ID in JSON format. As an example, enter the prompt "Please collect all tweets from user ID 'example_user'."

[1467] Step 2: Data extraction

[1468] The server extracts text data from the collected data.

[1469] Input: Collected social media posts, call history, and message history data

[1470] Output: Text data

[1471] Specific operation: The server extracts the text message content from the JSON format data and converts it to plain text. Specifically, it inputs the prompt "Please extract the tweet content in text format from the JSON format data."

[1472] Step 3: Data analysis

[1473] The server uses natural language processing technology to analyze the text data, extract keywords, and analyze frequently occurring phrases and writing styles.

[1474] Input: Extracted text data

[1475] Output: Linguistic features (keywords, frequent phrases, contextual information)

[1476] Specific behavior: The server uses an NLP library (e.g., SpaCy or NLTK) to extract frequent words and phrases from the text and identify the context. The server inputs the prompt: "Extract frequent words and phrases from the text data."

[1477] Step 4: Personality model generation

[1478] The server generates a personality model based on the identified linguistic features, including the deceased's speech patterns and intonation.

[1479] Input: Identified linguistic features

[1480] Output: Personality model

[1481] Specific operation: The server uses a generative AI model (e.g., GPT-3) to create a model that mimics the deceased person's speech patterns. It inputs the prompt, "Please use the extracted features to create a model that reproduces the deceased person's speech patterns."

[1482] Step 5: Confirm consent

[1483] The user checks whether the deceased person agreed to use the service before they died, and uploads a consent form to the server if necessary.

[1484] Input: Consent form (possibly scanned)

[1485] Output: Service use authorization

[1486] Specific operation: The user scans the consent form and uploads it to the server through the web portal. The server verifies and records the uploaded consent form. It then instructs the user to "scan the consent form and upload it through the web portal."

[1487] Step 6: Audio Reproduction

[1488] The device downloads a personality model from the server and reproduces the voice using voice changer software.

[1489] Input: Personality model data

[1490] Output: Reproduced audio

[1491] Specific operation: The device loads the model data downloaded from the server and generates voice using voice changer software such as DeepTalk or VoxCeleb2. The device instructs the user to "load the model data and reproduce the voice using voice changer software."

[1492] Step 7: Video Frame Generation

[1493] The terminal generates video frames in accordance with the reproduced audio and presents them to the user.

[1494] Input: Reproduced audio

[1495] Output: Video frame

[1496] Specific operation: The device uses facial animation technology (e.g., DeepFake) to create a video frame that corresponds to the reproduced audio. It then prompts the device to "create a video frame using facial animation technology based on the generated audio."

[1497] (Application example 1)

[1498] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1499] When a loved one passes away, there are limited ways to cherish memories of the deceased. Conventional technology lacks the means to reproduce the words and conversations of the deceased, in addition to visual records such as photographs and videos, making it difficult to recreate the emotional connection with the deceased. Furthermore, there is a demand for providing a natural conversation experience that includes the unique phrases and intonations used by the deceased, but no concrete means for achieving this have been established.

[1500] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1501] In this invention, the server includes means for collecting the communication history of the deceased, means for extracting text data from the communication history, and means for analyzing the text data using natural language processing technology to identify the linguistic characteristics of the deceased. This allows a personality model of the deceased to be generated based on the analyzed data, enabling the user to experience a simulated conversation with the deceased. Furthermore, the content of the conversation provided based on the generated personality model is more realistic, allowing the user to feel a deeper emotional connection with the deceased. This allows for a more realistic reproduction of memories with the deceased.

[1502] "Communication history" is a general term for records such as messages, call records, and social media posts sent and received through electronic communication means used by the deceased during their lifetime.

[1503] "Text data" refers to character information and sentences extracted from the communication history.

[1504] "Natural language processing technology" is a computer processing technology for understanding, analyzing, and generating human language, and includes a wide range of algorithms and methods for mechanically processing natural language.

[1505] "Linguistic features" refer to the characteristics of language use possessed by a particular writer or speaker, such as frequent phrases, lexical choice, and stylistic patterns.

[1506] A "personality model" is a computer model constructed to recreate the language usage and speech patterns of a deceased person based on their linguistic characteristics.

[1507] "Voice reproduction means" refers to techniques or devices that use a personality model to reproduce the voice, intonation, and distinctive speaking style of the deceased.

[1508] A "video frame" is an individual image that reproduces a specific moment in time as a video, and is the basic unit for expressing movement as a video when displayed continuously.

[1509] "Pseudo-dialogue" refers to providing a user with an experience that makes them feel as if they are conversing with the deceased person using the generated personality model.

[1510] "Dialogue content" refers to the content of words and messages exchanged between the user and the personality model of the deceased person in a simulated dialogue.

[1511] This invention is a system that collects and analyzes the communication history of a deceased person, and generates and recreates a personality model based on the linguistic features of the communication history. This system allows users to enjoy a virtual conversation experience in which the deceased person feels as if they are alive. This system incorporates key elements such as "communication history," "text data," "natural language processing technology," "linguistic features," "personality model," "means for reproducing voice," "video frames," "simulated dialogue," and "dialogue content."

[1512] Hardware and software used

[1513] Hardware: Smartphone (iPhone or Android)

[1514] Software: Python, Flask, TensorFlow / Keras, Pandas, NLTK, Twilio API, Google Cloud Text-to-Speech API

[1515] Program processing overview

[1516] 1. Data collection method (server)

[1517] The server receives the API key of the deceased person's social media account from the user. Using this API key, the server accesses the deceased person's social media account and collects their communication history, including Twitter and Facebook posts, messages, and call records.

[1518] 2. Data extraction method (server)

[1519] The server converts the collected social media data into an appropriate format and extracts it as text data. The collected data is provided in JSON format, but it is converted into an easily handled text format.

[1520] 3. Data analysis method (server)

[1521] The server analyzes the extracted text data using natural language processing libraries such as Pandas and NLTK to identify the linguistic characteristics of the deceased and extract frequent phrases, expressions, and writing style.

[1522] 4. Personality model generation means (server)

[1523] Based on the analyzed linguistic features, the server uses TensorFlow and Keras to generate a personality model of the deceased, a computer model that reproduces the speech patterns and intonation of the deceased.

[1524] 5. Consent confirmation means (user)

[1525] Users must verify through a web portal whether the deceased person consented to use the service before they died, and then upload a scanned copy of the consent form.

[1526] 6. Audio reproduction means (terminal)

[1527] The device uses a personality model downloaded from the server to reproduce the voice of the deceased person using voice synthesis technology such as the Google Cloud Text-to-Speech API.

[1528] 7. Video frame generation means (terminal)

[1529] As the audio plays, the device uses synthetic video technology to generate a video of the deceased, providing the user with realistic video frames synchronized with the audio.

[1530] Specific examples

[1531] For example, if a deceased person frequently used the greeting "Good morning," the server can identify this phrase and generate a video message in which the device says "Good morning" through a voice changer. The following is an example of a prompt sentence:

[1532] Example prompt sentence:

[1533] Prompt: "Good morning. What are you planning to do today?"

[1534] Response: The generated personality model responds, "Good morning! I'm planning to go for a walk today. How about you?"

[1535] In this way, the user can experience memories of the deceased more realistically through simulated dialogue.

[1536] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1537] Step 1:

[1538] Data collection method (server)

[1539] The server receives the API key and login information of the deceased person's SNS from the user. The server then collects the deceased person's communication history (messages, posts, call records, etc.) from various SNS platforms (e.g., Twitter, Facebook). The input data is the API key and user ID, and the output is communication history data in JSON format.

[1540] Step 2:

[1541] Data extraction method (server)

[1542] The server converts the collected SNS data from an appropriate format (for example, JSON format) into text data. Here, the tweet text and message content are extracted and organized as text data. The input is communication history data in JSON format, and the output is formatted text data. Specifically, the server extracts the necessary information from the data obtained from each SNS platform and converts it into text format.

[1543] Step 3:

[1544] Data analysis method (server)

[1545] The server analyzes the text data using natural language processing techniques (e.g., Pandas or NLTK). The analysis identifies the linguistic characteristics of the deceased (frequent phrases, writing style, intonation, etc.). The input is text data, and the output is metadata about the linguistic characteristics of the deceased. Specific operations include word frequency analysis, morphological analysis, and contextual analysis.

[1546] Step 4:

[1547] Personality model generation means (server)

[1548] Based on the analysis results, the server uses TensorFlow and Keras to generate a personality model of the deceased. This is a machine learning model that learns the deceased's speech patterns. The input is metadata about the linguistic features of the deceased, and the output is a trained personality model. Specifically, it trains an LSTM model using the deceased's speech data.

[1549] Step 5:

[1550] Consent confirmation method (user)

[1551] The user checks through the web portal whether the deceased person had consented to use the service before they died, and then uploads a scanned copy of the consent form to the server. The input is the scanned consent form (electronic file), and the output is the result of uploading the consent form to the server. Specifically, the user selects the consent form on the portal site and presses the upload button.

[1552] Step 6:

[1553] Audio reproduction means (terminal)

[1554] The device uses the personality model downloaded from the server and generates the voice of the deceased person using the Google Cloud Text-to-Speech API, etc. The input is the personality model and text data, and the output is an audio file. Specifically, it converts the text-generated dialogue into audio.

[1555] Step 7:

[1556] Video frame generation means (terminal)

[1557] The device uses synthetic video technology to generate a video of the deceased person in sync with the audio. The input is the generated audio and video data of the deceased person while they were alive, and the output is video frames synchronized with the audio. Specifically, the device generates and edits a video sequence corresponding to the audio.

[1558] This allows the user to experience a simulated conversation with the deceased, allowing memories of the deceased to be recreated more realistically.

[1559] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1560] This system collects and analyzes the communication history of a deceased person, and generates and recreates a personality model based on the linguistic features. This system also incorporates an emotion engine that recognizes the user's emotions, providing dialogue that responds to the user's emotions, thereby realizing a more realistic communication experience.

[1561] System configuration

[1562] The system of the present invention comprises the following components:

[1563] 1. Data collection method (server)

[1564] 2. Data extraction method (server)

[1565] 3. Data analysis method (server)

[1566] 4. Personality model generation means (server)

[1567] 5. Consent confirmation means (user)

[1568] 6. Audio reproduction means (terminal)

[1569] 7. Video frame generation means (terminal)

[1570] 8. Emotion Recognition Means (Emotion Engine)

[1571] Program processing overview

[1572] Data collection method (server)

[1573] The user provides the server with the API key and login information of the social networking site, which the server uses to collect the deceased person's communication history, including social networking site posts, call records, and message history.

[1574] Example: The server uses the Twitter API to retrieve the tweet history of a specified user ID.

[1575] Data extraction method (server)

[1576] The server converts the collected data into an appropriate format and extracts it as text data.

[1577] Example: The server converts JSON-formatted tweet data into text format and extracts each sentence.

[1578] Data analysis method (server)

[1579] The server uses natural language processing technology to analyze the text data, extract keywords and analyze writing style, and identify linguistic features.

[1580] Example: A server uses an NLP library to extract frequently used phrases and keywords from text.

[1581] Personality model generation means (server)

[1582] The server generates a personality model based on the identified linguistic features, including the deceased's speech patterns and intonation.

[1583] Example: The server models typical response patterns during a conversation based on the analyzed features.

[1584] Consent confirmation method (user)

[1585] A document is uploaded to the server to confirm that the user has agreed to use the service.

[1586] Example: A user uploads a scanned consent form through a web portal.

[1587] Audio reproduction means (terminal)

[1588] The device downloads the generated personality model from the server and reproduces the voice using voice changer technology.

[1589] Example: The device uses voice changer software to generate the voice of a deceased person.

[1590] Video frame generation means (terminal)

[1591] The terminal generates video frames based on the reproduced audio and presents them to the user.

[1592] Example: A device uses synthetic video techniques to create video that accompanies the reproduced audio.

[1593] Emotion recognition means (emotion engine)

[1594] The emotion engine analyzes real-time data obtained from the user's camera and microphone to identify the user's emotional state.

[1595] Example: An emotion engine analyzes a user's facial expressions and tone of voice to determine whether they are happy, sad, or surprised.

[1596] How the Emotion Engine Works

[1597] The emotion engine provides a means to recognize the user's emotional state in real time and adjust the audio and video frames of the deceased based on that. For example, if the emotion engine determines that the user is sad, it can generate a soothing voice or comforting message for the deceased. This functionality allows the user to feel more connected to the deceased.

[1598] Specific examples

[1599] For example, if the deceased frequently used the greeting "good morning," the emotion engine may determine that the user is feeling unwell. The emotion engine will adjust the reproduced voice to be more gentle and comforting depending on the situation, and change the video frame accordingly. Looking at the device screen, the user will feel as if the deceased is gently encouraging them.

[1600] In this way, the system of the present invention can generate and provide users with realistic interactive video frames that incorporate the memories and characteristics of the deceased, enabling personalized interactions based on the user's emotional state, helping to alleviate the grief of losing a loved one.

[1601] The processing flow will be explained below.

[1602] Step 1:

[1603] The user provides the server with the API key and login information of the SNS. Specifically, the user enters the API key of Twitter or other SNS through a web form and submits it.

[1604] Step 2:

[1605] The server uses the provided API key to collect the post data of the deceased person from the SNS. Specifically, the server calls the Twitter API to obtain the tweet history of the specified user ID.

[1606] Step 3:

[1607] The server converts the collected data into an appropriate format and extracts it as text data. Specifically, the server converts the JSON-formatted tweet data into text format and extracts each sentence.

[1608] Step 4:

[1609] The server analyzes the extracted text data using natural language processing techniques. Specifically, the server uses an NLP library to extract frequently used phrases and keywords from the text.

[1610] Step 5:

[1611] The server then uses the text data analyzed in the previous step to identify the deceased's linguistic characteristics, such as writing style, tone, specific phrases, and intonation patterns.

[1612] Step 6:

[1613] Based on the identified characteristics, the server generates a personality model that includes the deceased's speech patterns and intonation. Specifically, it uses the analysis results to incorporate the deceased's frequently used phrases and reaction patterns into the model.

[1614] Step 7:

[1615] The user uploads a document to the server to confirm that they have agreed to use the service. Specifically, the user uploads a scanned copy of the consent form via a web portal.

[1616] Step 8:

[1617] The device downloads the generated personality model from the server and uses voice changer technology to recreate the voice of the deceased. Specifically, the device obtains the personality model via the internet and uses voice synthesis software to generate the voice of the deceased.

[1618] Step 9:

[1619] The device generates video frames based on the reproduced audio. Specifically, it combines the synthesized audio with footage of the deceased (previous video clips or photos) to create natural-sounding video frames.

[1620] Step 10:

[1621] The terminal displays the video frames to the user, specifically, displays the video frames on the terminal's display so that the user can view them.

[1622] Step 11:

[1623] The emotion engine captures real-time data from the user's camera and microphone and analyzes their emotions by analyzing their facial expressions, tone of voice, and word choice to determine their current emotional state.

[1624] Step 12:

[1625] Based on the analysis, the emotion engine adjusts the content of the deceased person's audio and video frames. For example, if the user is sad, the message of the deceased will be changed to be more comforting.

[1626] Step 13:

[1627] The device plays the adjusted content and displays it to the user. Specifically, the content generated based on instructions from the emotion engine is displayed on the display, and the user watches and listens to it.

[1628] Example 2

[1629] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1630] Although there are systems that can recreate the memories and characteristics of the deceased and allow users to have a realistic conversation with them, these systems lack the ability to recognize the user's emotional state and dynamically adjust the conversation content based on that emotion. This limits the conversational experiences provided, making it difficult for users to feel a deeper connection with the deceased.

[1631] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1632] In this invention, the server includes means for collecting the communication history of the deceased, means for extracting text data from the communication history, means for analyzing the text data using natural language processing technology and identifying linguistic characteristics of the deceased, means for generating a personality model of the deceased based on the identified linguistic characteristics, means for downloading the personality model of the deceased from the server to a terminal, means for reproducing voice based on the personality model, means for generating video frames based on the reproduced voice, means for recognizing the emotional state of the user, and means for adjusting the voice and video frames based on the emotional state of the user. This allows the user to not only realistically experience memories and characteristics of the deceased, but also to have a personalized interaction experience according to emotions.

[1633] "Communication history" refers to a series of communication data such as social media posts, call records, and message history made by the deceased person while they were alive.

[1634] "Text data" refers to data in sentence format extracted from the communication history.

[1635] "Natural language processing technology" refers to technology used by computers to understand, analyze, and generate human language.

[1636] "Linguistic characteristics" refer to linguistic characteristics and patterns, such as the particular words, writing style, and frequently used phrases used by a particular person.

[1637] A "personality model" is a data model designed to reproduce the linguistic characteristics, speech patterns, intonation, etc. of a deceased person.

[1638] "Voice reproduction" refers to the technique or process of recreating the voice of a deceased person based on a generated personality model.

[1639] A "video frame" is a frame of video that corresponds to the reproduced audio and shows the image and gestures of the deceased.

[1640] "Emotional state" refers to a user's current psychological and sensory state, which is typically determined through facial expressions, tone of voice, etc.

[1641] "Emotion recognition" refers to the technology of analyzing a user's emotional state and thereby identifying a specific emotion.

[1642] A "personalized interaction experience" is an interaction that is optimized for a user's unique characteristics and situation, and is adjusted based on their emotional state.

[1643] This system collects and analyzes the communication history of a deceased person, and generates and recreates a personality model based on the linguistic features. This system also incorporates an emotion engine that recognizes the user's emotions, providing dialogue that responds to the user's emotions, thereby realizing a more realistic communication experience.

[1644] System configuration

[1645] The system of the present invention comprises the following components:

[1646] 1. Data collection method (server)

[1647] 2. Data extraction method (server)

[1648] 3. Data analysis method (server)

[1649] 4. Personality model generation means (server)

[1650] 5. Consent confirmation means (user)

[1651] 6. Audio reproduction means (terminal)

[1652] 7. Video frame generation means (terminal)

[1653] 8. Emotion Recognition Means (Emotion Engine)

[1654] Detailed system configuration and processing method

[1655] Data collection method (server)

[1656] The user enters their social networking service API key and login information, which is then sent in a secure format to the server, which then uses this information to collect the deceased person's communication history, including social networking posts, call logs, and message history. This process uses standard APIs, such as the Twitter API and Facebook API.

[1657] Example: The server uses the Twitter API to retrieve the tweet history of a specified user ID.

[1658] Data extraction method (server)

[1659] The server analyzes the collected data and extracts it as text data. At this time, it converts the necessary parts of the data in JSON or XML format into text format and extracts each sentence and word.

[1660] Example: The server converts JSON-formatted tweet data into text format and extracts each sentence.

[1661] Data analysis method (server)

[1662] The server uses natural language processing (NLP) techniques to analyze the extracted text data. During the analysis, it identifies linguistic features such as keywords, frequent phrases, and writing style. This process uses NLP libraries (e.g., NLTK, spacy).

[1663] Example: The server uses an NLP library to extract frequently used phrases and keywords from the text.

[1664] Personality model generation means (server)

[1665] The server generates a personality model based on the analyzed linguistic features, including the deceased's speech patterns and intonation, which is then used to simulate conversations similar to those of the deceased.

[1666] Example: The server uses the analyzed features to model typical response patterns during a conversation.

[1667] Consent confirmation method (user)

[1668] The user uploads a consent form for using the service to the server in PDF format, etc. The server then checks the uploaded consent form and determines whether or not the use of the service is permitted.

[1669] Example: User uploads scanned consent form through web portal.

[1670] Audio reproduction means (terminal)

[1671] The device uses a personality model downloaded from the server and uses voice changer technology (e.g., Voicemod, iZotope) to recreate the voice of the deceased person, which is used during interactions with the user.

[1672] Example: The device uses voice changer software to generate the voice of the deceased person.

[1673] Video frame generation means (terminal)

[1674] Based on the reproduced audio, the device uses synthesis technology (e.g., D-ID, Reallusion) to generate video frames that display the deceased person's face and gestures in sync with the audio.

[1675] Example: The device uses synthetic video techniques to create a video that matches the reproduced audio.

[1676] Emotion recognition means (emotion engine)

[1677] The emotion engine analyzes real-time data captured from the user's camera and microphone. It uses facial expression analysis software (e.g., Affectiva, Microsoft Azure Face API) and voice analysis tools to identify the user's emotional state.

[1678] Example: An emotion engine analyzes a user's facial expressions and tone of voice to determine whether they are sad, happy, or surprised.

[1679] Examples of concrete examples and prompts

[1680] For example, if the deceased frequently used the greeting "Good morning," the emotion engine may determine that the user's emotional state was poor. In this case, the emotion engine will adjust the reproduced voice to be calm and comforting, and change the video frame accordingly. Looking at the device screen, the user will feel as if the deceased is gently encouraging them.

[1681] Example prompt sentence:

[1682] It uses phrases that allow users to start a natural conversation, such as "Hi, how are you doing?" or "How was your day?"

[1683] In this way, the system of the present invention can realistically recreate the memories and characteristics of the deceased and provide a personalized interaction experience for the user.

[1684] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1685] Step 1: Enter user information

[1686] The user enters the SNS API key and login information, which is then sent to the server in a secure format. Specifically, the user enters the information on the system's login screen and presses the send button. The input data is the SNS API key and login information, and the output is the API key and login information sent to the server.

[1687] Step 2: Data collection

[1688] The server uses the API key and login information sent by the user to collect the deceased person's communication history, such as social media posts, call records, and message history. Specifically, the server uses an appropriate API (e.g., Twitter API, Facebook API) to obtain data for the specified user ID. The input data is the API key and login information provided by the user in the previous step, and the output is the collected communication history (e.g., social media post data, call records).

[1689] Step 3: Data extraction

[1690] The server converts the collected communication history into an appropriate format and extracts it as text data. Specifically, the server converts SNS data in JSON or XML format into text format and extracts each sentence. The input data is the communication history (e.g., tweet data in JSON format), and the output is the extracted text data (e.g., single sentences of text).

[1691] Step 4: Data analysis

[1692] The server uses natural language processing (NLP) techniques to analyze the extracted text data. During the analysis process, it identifies linguistic features such as keywords, frequently occurring phrases, and writing style. Specifically, the server uses an NLP library (e.g., NLTK, spacy) to extract frequently used phrases and keywords from the text. The input data is the text data, and the output is the analyzed linguistic features (e.g., a list of keywords and phrases).

[1693] Step 5: Personality model generation

[1694] The server generates a personality model based on the analyzed linguistic features, including the deceased's speech patterns and intonation. Specifically, the server uses a machine learning algorithm to model typical response patterns during conversations based on the analyzed features. The input data are the linguistic features, and the output is the generated personality model.

[1695] Step 6: Confirm consent

[1696] The user uploads the consent form for using the service to the server in PDF format or other format. Specifically, the user scans the consent form and uploads it through a web portal. The input data is the consent form in PDF format, and the output is a consent confirmation flag on the server.

[1697] Step 7: Audio Reproduction

[1698] The device uses a personality model downloaded from the server and voice changer technology (e.g., Voicemod, iZotope) to recreate the deceased's voice. Specifically, the device converts the deceased's voice pattern into audio data based on the downloaded personality model. The input data is the personality model, and the output is the recreated audio data.

[1699] Step 8: Video Frame Generation

[1700] The device generates video frames based on the reproduced audio. Specifically, the device uses synthesis technology (e.g., D-ID, Reallusion) to create video that matches the reproduced audio. The input data is the reproduced audio data, and the output is the generated video frames.

[1701] Step 9: Emotion Recognition

[1702] The emotion engine analyzes real-time data acquired from the user's camera and microphone. Specifically, the emotion engine uses facial expression analysis software (e.g., Affectiva, Microsoft Azure Face API) and voice analysis tools to identify the user's emotional state. The input data is the real-time data acquired from the camera and microphone, and the output is the identified user's emotional state.

[1703] Step 10: Dialogue Adjustment

[1704] The emotion engine appropriately adjusts the content and tone of the dialogue based on the emotion data acquired in the previous stage. For example, if the user is sad, the emotion engine generates a calm voice and comforting words, and optimizes the video frames accordingly. The input data is the user's emotional state, and the output is the adjusted voice data and video frames.

[1705] (Application example 2)

[1706] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1707] There is a need for a means to maintain the emotional connection between the deceased and those left behind, particularly to ease the grief of losing the deceased. However, conventional technologies only reproduce the personality of the deceased, and are unable to engage in dialogue that reflects the user's emotional state, making it difficult to provide a realistic communication experience. Furthermore, there is no system in physical stores that allows users to feel emotionally reassured, making it difficult to alleviate the user's psychological burden.

[1708] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1709] In this invention, the server includes means for collecting the communication history of the deceased, means for extracting text data from the communication history, means for analyzing the text data using natural language processing technology to identify the linguistic characteristics of the deceased, means for generating a personality model of the deceased based on the identified linguistic characteristics, means for reproducing voice based on the personality model, means for generating video frames using the reproduced voice, means for recognizing the emotional state of a user in real time, means for adjusting the content of the dialogue in accordance with the recognized emotional state, and means for reproducing a dialogue with the deceased via a tablet terminal or smartphone in a physical store. This allows for a realistic reproduction of the personality of the deceased while enabling personalized dialogue in accordance with the user's emotional state, thereby providing users with a sense of emotional security even in a physical store.

[1710] "Communication history" is a record of past interactions recorded as digital data, such as text messages, call logs, and social media posts, across the mediums used by individuals to communicate.

[1711] "Natural language processing technology" is a field of artificial intelligence that enables computers to understand, interpret, and generate human language, and is a technology that analyzes text and extracts linguistic features.

[1712] A "personality model" is a digital recreation of a specific individual's linguistic characteristics, speech patterns, and presence, constructed from the deceased's text data and call records.

[1713] "Voice reproduction" is a technology that generates voice by imitating the speaking style and voice quality of a deceased person based on a generated personality model.

[1714] A "video frame" is a video frame generated in accordance with the reproduced audio, which digitally reproduces the image of the deceased.

[1715] "Emotional state" refers to the emotions the user is currently feeling, such as joy, sadness, surprise, etc., and is measured in real time using a camera, microphone, etc.

[1716] A "tablet device" is a portable, flat-shaped electronic device that is equipped with a touch screen and supports the operation of applications.

[1717] A "smartphone" is a telephone terminal that has advanced computer functions in addition to calling functions, and can be used to realize a variety of functions by installing applications.

[1718] "Personalized dialogue" refers to conversation content that is customized based on the user's individual emotional state and behavior, and is dynamically generated by a personality model of the deceased person.

[1719] A "physical store" is a commercial facility that exists physically and provides a place for consumers to visit and receive goods or services.

[1720] The system that realizes this application example collects and analyzes the communication history of the deceased person to generate a personality model and provide dialogue that corresponds to the user's emotional state. This system is realized using the following hardware and software components.

[1721] System Components

[1722] 1. Data collection method (server)

[1723] The server uses the SNS API key and login information provided by the user to collect the deceased person's communication history, such as SNS posts, call records, and message history. For example, the server obtains data using the Twitter API or Facebook Graph API.

[1724] 2. Data extraction method (server)

[1725] The server converts the collected data into an appropriate format and extracts it as text data. In this step, the JSON format data is converted into text format and each sentence is extracted.

[1726] 3. Data analysis method (server)

[1727] The server analyzes the text data using natural language processing techniques, extracts keywords and analyzes writing style, and identifies the linguistic characteristics of the deceased. Specifically, it uses Python NLP libraries (e.g., spaCy and NLTK).

[1728] 4. Personality model generation means (server)

[1729] The server generates a personality model of the deceased person, including their speech patterns and intonation, based on the identified linguistic features, using OpenAI's GPT-based generative AI model.

[1730] 5. Audio reproduction means (terminal)

[1731] The device uses a personality model downloaded from a server to reproduce the voice using voice changer technology (e.g., Respeecher or Google's Tacotron).

[1732] 6. Video frame generation means (terminal)

[1733] The device generates video frames based on the reproduced audio and provides them to the user. The video is generated using synthetic video technology (e.g., Deepfake technology).

[1734] 7. Emotion Recognition Means (Emotion Engine)

[1735] The emotion engine analyzes real-time data obtained from the user's camera and microphone to identify the user's emotional state, and is implemented using Microsoft's Azure Face API and Emotion API.

[1736] 8. Dialogue Content Adjustment Method (Server)

[1737] The server generates personalized dialogue based on the recognized emotional state, using a generative AI model to dynamically create conversational content that matches the user's emotions.

[1738] 9. Physical store devices (tablets and smartphones)

[1739] The system operates via a tablet device installed in a brick-and-mortar store or a user's smartphone, which displays audio and video to recreate a conversation with the deceased.

[1740] Specific examples

[1741] A user visiting a brick-and-mortar store (e.g., a long-established inn or restaurant) accesses a tablet device installed in the store. The server collects the communication history of the deceased person provided by the user and generates a personality model. When the user visits a specific location (e.g., a place frequently visited by the deceased), the emotion engine uses facial recognition technology to detect the user's emotional state. If the server recognizes that the user is crying, it uses the generative AI model to generate dialogue content to comfort the user and displays it on the device.

[1742] Prompt Sentence Examples

[1743] Generate it as follows: You arrive at the lobby of an inn. Noticing your tears, your deceased father speaks to you kindly, saying, "Cheer up, I'm always watching over you."

[1744] This system allows users to experience a realistic reunion with their deceased loved ones and gain emotional comfort, thus recreating memories of the deceased in a physical store and reducing the psychological burden on users.

[1745] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1746] Step 1:

[1747] The user logs in to a tablet device installed in a physical store and gives permission to collect the communication history of the deceased. The input is the SNS API key and login information provided by the user. The server uses this information to collect data such as the deceased's SNS posts, call records, and message history. The output is the acquired communication history data.

[1748] Step 2:

[1749] The server converts the collected communication history data into a text format. The input is communication history data stored in JSON format or other structured data format. The server runs a data conversion program to extract the text data. The output is a list of the actual conversations in text format.

[1750] Step 3:

[1751] The server analyzes the text data using natural language processing techniques. The input is the extracted text data. The server uses Python NLP libraries (e.g., spaCy or NLTK) to extract keywords, analyze writing style, and identify frequently occurring phrases. The output is analyzed data containing the linguistic features of the deceased.

[1752] Step 4:

[1753] The server generates a personality model of the deceased person based on the analyzed data. The input is the analyzed data, including linguistic features. The server uses OpenAI's GPT-based generative AI model to build a personality model, including speech patterns and intonation. The output is a personality model of the deceased person.

[1754] Step 5:

[1755] The device recreates the voice using a personality model downloaded from a server. The input is the personality model. The device uses voice changer technology (e.g., Respeecher or Google's Tacotron) to generate the deceased person's voice. The output is the recreated voice.

[1756] Step 6:

[1757] The device generates video frames based on the reproduced audio. The input is the reproduced audio and video data of the deceased person before their death. The device uses synthetic video technology (e.g., Deepfake technology) to generate video that matches the audio. The output is video frames.

[1758] Step 7:

[1759] The emotion engine analyzes real-time data acquired from the user's camera and microphone to identify the user's emotional state. The input is data acquired in real time of the user's facial expressions and tone of voice. The emotion engine performs emotion recognition using Microsoft's Azure Face API and Emotion API. The output is data indicating the user's emotional state.

[1760] Step 8:

[1761] The server generates dialogue content according to the recognized emotional state. The input is the user's emotional state data and a personality model of the deceased. The server uses a generative AI model to create personalized dialogue content tailored to the user's emotions. The output is the generated dialogue content.

[1762] Step 9:

[1763] The device recreates the conversation with the deceased based on the generated conversation content. The input is the generated conversation content and video frames. The device displays this to the user, providing a realistic conversation with the deceased. The output is the video and audio of the deceased displayed on the user's screen.

[1764] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1765] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1766] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1767] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1768] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1769] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1770] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1771] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1772] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1773] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1774] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1775] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1776] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1777] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1778] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1779] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1780] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1781] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1782] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1783] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1784] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1785] The following is further disclosed regarding the above embodiment.

[1786] (Claim 1)

[1787] A means of collecting the deceased person's communication history;

[1788] means for extracting text data from the communication history;

[1789] A means for analyzing the text data using natural language processing technology to identify the linguistic characteristics of the deceased;

[1790] means for generating a personality model of the deceased person based on the identified linguistic features;

[1791] means for reproducing a voice based on said personality model;

[1792] means for generating video frames using the reproduced audio;

[1793] A system including:

[1794] (Claim 2)

[1795] A means to confirm the deceased person's consent to use the service during their lifetime;

[1796] means for providing the service after confirming said consent;

[1797] The system of claim 1 further comprising:

[1798] (Claim 3)

[1799] means for enabling the reproduced speech to include the intonation and frequent phrases of the deceased;

[1800] means for synchronizing said video frames with footage of the deceased person in life;

[1801] The system of claim 1 further comprising:

[1802] "Example 1"

[1803] (Claim 1)

[1804] A means for collecting user communication data;

[1805] means for extracting text data from the communication data;

[1806] A means for analyzing the text data using natural language processing technology and identifying the linguistic characteristics of the user;

[1807] means for generating a person model of the user based on the identified linguistic features;

[1808] means for reproducing a voice based on the person model;

[1809] means for generating video frames using the reproduced audio;

[1810] A system including:

[1811] (Claim 2)

[1812] A means to confirm the user's consent to use the service before death;

[1813] means for providing the service after confirming said consent;

[1814] The system of claim 1 further comprising:

[1815] (Claim 3)

[1816] means for making the reproduced speech include the pronunciation characteristics and frequently used expressions of the user;

[1817] means for synchronizing said video frames with pre-mortem images of the user;

[1818] The system of claim 1 further comprising:

[1819] "Application Example 1"

[1820] (Claim 1)

[1821] A means of collecting the deceased person's communication history;

[1822] means for extracting text data from the communication history;

[1823] A means for analyzing the text data using natural language processing technology to identify the linguistic characteristics of the deceased;

[1824] means for generating a personality model of the deceased person based on the identified linguistic features;

[1825] means for reproducing a voice based on said personality model;

[1826] means for generating video frames using the reproduced audio;

[1827] A means for a user to experience a simulated conversation with the deceased;

[1828] a means for providing the dialogue content based on a generated personality model;

[1829] A system including:

[1830] (Claim 2)

[1831] A means to confirm the deceased person's consent to use the service during their lifetime;

[1832] means for providing the service after confirming said consent;

[1833] The system of claim 1 further comprising:

[1834] (Claim 3)

[1835] means for enabling the reproduced speech to include the intonation and frequent phrases of the deceased;

[1836] means for synchronizing said video frames with footage of the deceased person in life;

[1837] The system of claim 1 further comprising:

[1838] "Example 2: Combining Emotion Engines"

[1839] (Claim 1)

[1840] A means of collecting the deceased person's communication history;

[1841] means for extracting text data from the communication history;

[1842] A means for analyzing the text data using natural language processing technology to identify the linguistic characteristics of the deceased;

[1843] means for generating a personality model of the deceased person based on the identified linguistic features;

[1844] means for downloading the personality model of the deceased person from a server to a terminal;

[1845] means for reproducing a voice based on said personality model;

[1846] means for generating video frames based on the reproduced audio;

[1847] means for recognizing the emotional state of a user;

[1848] means for adjusting audio and video frames based on the emotional state of the user;

[1849] A system including:

[1850] (Claim 2)

[1851] A means to confirm the deceased person's consent to use the service during their lifetime;

[1852] means for providing the service after confirming said consent;

[1853] The system of claim 1 further comprising:

[1854] (Claim 3)

[1855] means for enabling the reproduced speech to include the intonation and frequent phrases of the deceased;

[1856] means for synchronizing said video frames with footage of the deceased person in life;

[1857] The system of claim 1 further comprising:

[1858] "Application example 2 when combining emotion engines"

[1859] (Claim 1)

[1860] A means of collecting the deceased person's communication history;

[1861] means for extracting text data from the communication history;

[1862] A means for analyzing the text data using natural language processing technology to identify the linguistic characteristics of the deceased;

[1863] means for generating a personality model of the deceased person based on the identified linguistic features;

[1864] means for reproducing a voice based on said personality model;

[1865] means for generating video frames using the reproduced audio;

[1866] means for recognizing a user's emotional state in real time;

[1867] means for adjusting dialogue content in response to the recognized emotional state;

[1868] A means to recreate conversations with the deceased via tablets or smartphones in brick-and-mortar stores.

[1869] A system including:

[1870] (Claim 2)

[1871] A means to confirm the deceased person's consent to use the service during their lifetime;

[1872] means for providing the service after confirming said consent;

[1873] The system of claim 1 further comprising:

[1874] (Claim 3)

[1875] means for enabling the reproduced speech to include the intonation and frequent phrases of the deceased;

[1876] means for synchronizing said video frames with footage of the deceased person in life;

[1877] A means for reproducing a usage situation in the physical store;

[1878] The system of claim 1 further comprising: [Explanation of symbols]

[1879] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. A means of collecting the deceased person's communication history; means for extracting text data from the communication history; A means for analyzing the text data using natural language processing technology to identify the linguistic characteristics of the deceased; means for generating a personality model of the deceased person based on the identified linguistic features; means for reproducing a voice based on said personality model; means for generating video frames using the reproduced audio; A system including:

2. A means to confirm the deceased person's consent to use the service during their lifetime; means for providing the service after confirming said consent; The system of claim 1 further comprising:

3. means for enabling the reproduced speech to include the intonation and frequent phrases of the deceased; means for synchronizing said video frames with footage of the deceased person in life; The system of claim 1 further comprising:

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A