System

A system using a user terminal and server to analyze and generate AI voice data from deceased individuals allows bereaved families to relive conversations, offering comfort and empathy.

JP2026027104APending Publication Date: 2026-02-18SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024129525
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-08-05
Publication Date
2026-02-18

AI Technical Summary

Technical Problem

Current technology lacks the means to recreate the voices and words of the deceased, limiting the ways in which bereaved families can find healing and empathy.

Method used

A system that includes a user terminal for uploading voice data of the deceased, a server for analyzing and training a generative AI model on the extracted feature data, and a user terminal for playing back the generated voice data, allowing users to relive conversations with the deceased.

Benefits of technology

Enables bereaved family members to relive heartfelt conversations with the deceased, providing comfort and empathy through the use of AI-generated voice data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026027104000001_ABST
    Figure 2026027104000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system comprising: means for uploading a voice datum of a deceased person from a user device to a server; means for analyzing the voice datum received by the server and extracting a feature datum; means for training a generative AI model using the extracted feature datum; means for receiving a text message input by a user at the server and converting the text message into a voice datum using the trained generative AI model; and means for receiving the voice datum sent from the server at the user device and playing it to the user.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] The voices and words of the deceased fade over time, and surviving family and friends often lose the opportunity to communicate with them. Current technology lacks the means to recreate the memories of the deceased through audio, limiting the ways in which bereaved families can find healing and empathy. This invention uses AI technology to recreate the voices and words of the deceased, allowing bereaved families to relive their heartfelt conversation with the deceased. [Means for solving the problem]

[0005] The present invention allows users to hear the voice of the deceased through a system including a means for uploading voice data of the deceased from a user terminal to a server, a means for the server to analyze the received voice data and extract feature data, a means for training a generative AI model using the extracted feature data, a means for the server to receive text messages entered by the user and convert the text messages into voice data using the trained generative AI model, and a means for the user terminal to receive the voice data transmitted from the server and play it back to the user, thereby enabling bereaved family members to relive conversations with the deceased and find comfort and empathy.

[0006] A "deceased person" refers to a person who has passed away, particularly a person for whom a user provides voice data.

[0007] "Audio data" refers to data that contains a recording of the deceased person's voice and is saved in an audio file format (e.g., MP3, WAV).

[0008] "User" refers to a person who uploads audio data of a deceased person and then seeks to have a subsequent interactive experience.

[0009] A "user terminal" is a device used by a user to upload audio data or play generated audio, such as a smartphone or computer.

[0010] "Server" refers to a computer system that processes speech data, trains models, generates new speech, and other processes.

[0011] A "generative AI model" refers to an artificial intelligence algorithm that is trained on voice data to reproduce the vocal characteristics of a deceased person.

[0012] "Feature data" refers to characteristic data such as pitch, speed, and intonation of speech extracted from speech data.

[0013] A "text message" is character string information entered by the user, and represents the content that the user wants to have read aloud in the voice of the deceased.

[0014] "Voice data conversion" refers to the process of converting a text message into voice data using a generative AI model.

[0015] "Playback" refers to outputting the audio data generated by the user terminal as audio. [Brief explanation of the drawings]

[0016] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12]FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0017] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0018] First, the terms used in the following description will be explained.

[0019] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0020] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0021] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0022] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0023] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0024] [First embodiment]

[0025] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0026] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0027] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0028] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0029] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0030] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0031] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0032] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0033] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0034] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0035] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0036] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0037] This invention uses the voice data of the deceased to allow their family and friends to hear their voice again through AI. The program and specific processing flow of this system are shown below.

[0038] Program Overview

[0039] The program works through the following major steps:

[0040] 1. The user uploads the audio data of the deceased person.

[0041] 2. The device sends the audio data to the server.

[0042] 3. The server analyzes the voice data and extracts feature data.

[0043] 4. The server uses the feature data to train a generative AI model.

[0044] 5. The server receives the user's text message and generates audio data.

[0045] 6. The device receives and plays the generated audio data.

[0046] Processing flow and specific examples

[0047] User voice data input

[0048] 1. The user launches the dedicated application and clicks the "Upload audio data" button.

[0049] 2. The user selects the audio file (e.g. MP3, WAV) of the deceased person and begins uploading.

[0050] Terminal handling

[0051] 3. The device will import the selected audio file and check the format and quality of the audio file.

[0052] 4. If the format and quality match, the device sends the audio file to the server.

[0053] Validation and storage of voice data by the server

[0054] 1. The server receives the audio file sent from the device.

[0055] 2. The server analyzes the contents of the audio file to ensure there are no problems with the format or quality.

[0056] 3. If the audio file matches, the server stores the file in its database and assigns it a unique ID.

[0057] 4. The server notifies the device that the audio file has been saved.

[0058] Training generative AI models

[0059] 1. After the audio file is saved, the server performs a characteristic analysis of the deceased person's audio data.

[0060] 2. The server uses a feature extraction algorithm to analyze and extract features such as pitch, speed, and intonation of the speech.

[0061] 3. Train a generative AI model (e.g., a large-scale language model or a speech synthesis model) based on the extracted feature data.

[0062] 4. Once trained, the generative AI model is optimized to reproduce the vocal characteristics of the deceased person.

[0063] Enter and send text messages

[0064] 1. The user enters an arbitrary text message through the application (e.g., "How are you feeling today?").

[0065] 2. The user clicks the "Send" button to send the text message to the server.

[0066] Server-generated audio

[0067] 1. The server parses the received text message.

[0068] 2. The server inputs the analyzed text into a generative AI model and generates corresponding audio data.

[0069] 3. The generative AI model uses the trained data to convert the input text into audio data in the characteristic voice of the deceased.

[0070] 4. The server sends the generated voice data to the terminal.

[0071] Receiving and playing audio on the device

[0072] 1. The device analyzes the voice data received from the server.

[0073] 2. The device checks the integrity of the audio data and prepares for playback.

[0074] 3. The device plays the audio data within the application, allowing the user to hear the voice of the deceased.

[0075] Specific examples

[0076] Consider a case where a user uses a smartphone app to upload audio data of their deceased grandmother. The audio data is uploaded to the app, and the smartphone sends the data to a server. The server analyzes the audio data, extracts the characteristics of the grandmother's voice, and trains an AI model. Next, when the user types a text message into the app saying, "Tell me about your favorite flower," the server receives this message and uses the trained AI model to generate audio data in the grandmother's voice. The generated audio data is sent to the smartphone and played back by the app. The user can hear their grandmother's voice again, which can be moving and comforting.

[0077] The above is an example of the implementation of the present invention. This system allows the bereaved family to re-create a heart-to-heart conversation with the deceased, and to gain healing and empathy.

[0078] The processing flow will be explained below.

[0079] Step 1:

[0080] User

[0081] The user launches the dedicated application and clicks the "Upload audio data" button.

[0082] Step 2:

[0083] User

[0084] The user selects the deceased person's audio file (e.g. MP3, WAV) and begins uploading.

[0085] Step 3:

[0086] Terminal

[0087] The device captures the selected audio file and checks the format and quality of the audio file.

[0088] Step 4:

[0089] Terminal

[0090] If the format and quality match, the terminal sends the audio file to the server.

[0091] Step 5:

[0092] server

[0093] The server receives the audio file sent from the terminal.

[0094] Step 6:

[0095] server

[0096] The server analyzes the content of the audio file to ensure that there are no problems with the format or quality.

[0097] Step 7:

[0098] server

[0099] If the audio file matches, the server stores the file in a database and assigns it a unique ID.

[0100] Step 8:

[0101] server

[0102] The server notifies the terminal that the audio file has been saved.

[0103] Step 9:

[0104] server

[0105] The server uses the stored audio files to characterize the deceased person's audio data.

[0106] Step 10:

[0107] server

[0108] The server uses feature extraction algorithms to analyze and extract features such as pitch, speed, and intonation of the speech.

[0109] Step 11:

[0110] server

[0111] Based on the extracted feature data, generative AI models (e.g., large-scale language models or speech synthesis models) are trained.

[0112] Step 12:

[0113] server

[0114] Once trained, the generative AI model is optimized to reproduce the vocal characteristics of the deceased person.

[0115] Step 13:

[0116] User

[0117] Through the application, the user enters an arbitrary text message (e.g., "How are you feeling today?").

[0118] Step 14:

[0119] User

[0120] The user clicks the "Send" button to send the text message to the server.

[0121] Step 15:

[0122] Terminal

[0123] The terminal transmits the input text message to the server and makes a request to the server.

[0124] Step 16:

[0125] server

[0126] The server parses the received text message.

[0127] Step 17:

[0128] server

[0129] The server inputs the analyzed text into a generative AI model and generates corresponding audio data.

[0130] Step 18:

[0131] Generative AI Models

[0132] The generative AI model uses trained data to convert input text into audio data in the characteristic voice of the deceased.

[0133] Step 19:

[0134] server

[0135] The server transmits the generated voice data to the terminal.

[0136] Step 20:

[0137] Terminal

[0138] The terminal analyzes the voice data received from the server.

[0139] Step 21:

[0140] Terminal

[0141] The device checks the integrity of the audio data and prepares it for playback.

[0142] Step 22:

[0143] Terminal

[0144] The device plays the audio data within the application, allowing the user to hear the voice of the deceased.

[0145] Example 1

[0146] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0147] Technology that allows bereaved families and friends to hear the voices of the deceased again can bring emotion and healing, but conventional systems require a lot of work, such as uploading the voice data, checking the quality, and training the AI ​​model, making it difficult to operate efficiently.In addition, because the processes of checking the voice data quality, extracting features, and training the AI ​​model are complex, making it easy for users to use is a challenge.

[0148] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0149] In this invention, the server includes a means for a user to upload voice data of the deceased via a terminal, a means for the terminal to check the format and quality of the voice data and send it to the server, and a means for the server to analyze the received voice data and extract feature data, thereby enabling a user to efficiently and easily use the voice data of the deceased and reproduce the deceased's voice using an AI model.

[0150] A "user" is a person who uses the system to upload and play back audio data of a deceased person.

[0151] A "terminal" is a device operated by a user that uploads audio data, checks the format and quality, transmits it to a server, and plays back the received audio data.

[0152] A "server" is a device that analyzes received voice data, extracts feature data, trains a generative AI model, generates voice data, and transmits it to a terminal.

[0153] "Audio data" refers to a digital data file (e.g., MP3, WAV) that records the voice of a deceased person.

[0154] "Uploading" is the act of a user sending audio data to a server via a terminal.

[0155] "Format" refers to the format of the audio data file (e.g., MP3, WAV).

[0156] "Quality" refers to characteristics related to the sound quality of data, such as the sampling rate and bit depth of the audio data.

[0157] "Analysis" is the process by which the server performs detailed analysis of the received audio data to extract specific information.

[0158] "Feature data" is information that indicates voice characteristics such as pitch, speed, and intonation extracted from voice data.

[0159] A "generative AI model" is an artificial intelligence model trained to convert text messages into audio data based on feature data.

[0160] "Training" is the process of learning from feature data so that the generative AI model can accurately reproduce the voice of the deceased.

[0161] A "text message" is text information that is entered by a user and sent to a server.

[0162] "Conversion" is the process of using a generative AI model to turn a text message into audio data.

[0163] "Reception" refers to the act of the server receiving data transmitted from the user terminal, or the act of the user terminal receiving data transmitted from the server.

[0164] "Playback" refers to the act of outputting audio data received by a user terminal as audio so that the user can hear it.

[0165] This invention is a system that uses the voice data of the deceased to allow bereaved family members and friends to hear the voice of the deceased again through AI. This system consists of three main elements: a user, a terminal, and a server. The specific configuration and embodiments of this system are described below.

[0166] First, the user launches a dedicated application on their device. This device can be a smartphone, tablet, or computer, which is capable of importing and playing audio data. When the user clicks the "Upload Audio Data" button, a screen appears allowing them to select the audio file (e.g., MP3, WAV) of the deceased person. The user selects the target audio file and begins uploading.

[0167] The device then retrieves the selected audio file and checks its format and quality. This check checks whether the audio file is in MP3 or WAV format and meets certain audio quality standards (e.g., 44.1kHz, 16-bit). If the quality is acceptable, the device sends the audio file to the server. During this process, the user sees an "Uploading" progress bar.

[0168] When the server receives the audio file sent from the device, it checks the format and quality of the audio data again. If there are no problems, the server saves the audio file in its database and assigns a unique identifier (ID) to the file. The server notifies the device that the file has been saved, allowing the user to confirm that the upload was successful.

[0169] The server extracts feature data from the stored audio files. Specifically, it uses a feature extraction algorithm to analyze and extract features such as pitch, speed, and intonation. This feature data is used to train a generative AI model (e.g., a large-scale language model or a speech synthesis model). This model is optimized to reproduce the characteristics of the deceased person's voice.

[0170] If a user wants to use the voice of a deceased person, they can enter any text message through the application, for example, "How are you feeling today?" or "Tell me about a flower you liked," and click the "Send" button. This message will be sent to the server.

[0171] The server analyzes the received text message and inputs it into a generative AI model to generate corresponding audio data. This AI model converts the input text into audio data in the deceased's characteristic voice based on the trained audio feature data. The generated audio data is then sent to the device.

[0172] The audio data received by the device is checked for integrity and then prepared for playback. Finally, the audio data is played back within the application, allowing the user to hear the voice of the deceased. This recreates an emotional conversation with the deceased, bringing emotion and healing to bereaved family and friends.

[0173] Specifically, when a user uses a smartphone app to upload audio data of their deceased grandmother, the audio data is uploaded to the app, and the smartphone sends the data to a server. The server analyzes the audio data, extracts the characteristics of the grandmother's voice, and trains an AI model. When the user types a text message into the app saying, "Tell me about your favorite flower," the server receives the message and uses the trained AI model to generate audio data in the grandmother's voice. The generated audio data is sent to the smartphone and played back in the app. The user can hear their grandmother's voice again, and feel moved and comforted.

[0174] An example of a prompt is as follows:

[0175] Prompt: "Tell me about a flower you loved."

[0176] Generated text: "Hello. I love talking about flowers you liked. My favorite was roses..."

[0177] This invention allows users to relive the heartfelt conversation by listening to the voice of the deceased again, and thus find healing.

[0178] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0179] Step 1:

[0180] The user launches the dedicated application on the device and clicks the "Upload audio data" button. When the user selects an audio file to upload, this file is imported into the device. The input is the audio file selected by the user, and the output is the imported audio file.

[0181] Step 2:

[0182] The device checks the format and quality of the selected audio file. Specifically, it checks that the file is in MP3 or WAV format and that the audio quality is 44.1 kHz and 16-bit. The input is the imported audio file, and the output is the result of matching the format and quality.

[0183] Step 3:

[0184] The device checks the format and quality and, if it is found to be compatible, sends the audio file to the server. A progress bar indicating "uploading" is displayed to the user during the upload. The input is the format and quality check result and the audio file, and the output is the audio file sent to the server.

[0185] Step 4:

[0186] The server receives the audio file sent from the terminal. The input is the audio file sent from the terminal, and the output is the received audio file.

[0187] Step 5:

[0188] The server rechecks the content of the received audio file and, if there are no problems, stores it in a database. Specifically, it analyzes the format and quality of the audio data again and assigns a unique identifier. The input is the received audio file, and the output is the audio file stored in the database and its unique identifier.

[0189] Step 6:

[0190] The server performs feature analysis on the audio files stored in the database. It uses a feature extraction algorithm to analyze and extract features such as pitch, speed, and intonation of the voice. Based on this analysis, feature data is generated. The input is the audio files stored in the database, and the output is the extracted feature data.

[0191] Step 7:

[0192] The server uses the extracted feature data to train a generative AI model. Specifically, the feature data is input into the AI ​​model and the AI ​​model is optimized through an appropriate learning process. The input is the feature data, and the output is a trained generative AI model.

[0193] Step 8:

[0194] A user enters a text message through an application. The user clicks a "Send" button to send the entered text message to a server. The input is the text message entered by the user, and the output is the text message sent to the server.

[0195] Step 9:

[0196] The server analyzes the received text message. The analyzed text message is input into the generative AI model to generate corresponding audio data. The input is the text message sent to the server, and the output is the generated audio data.

[0197] Step 10:

[0198] The server sends the generated voice data to the terminal. The input is the generated voice data, and the output is the voice data sent to the terminal.

[0199] Step 11:

[0200] The device analyzes the audio data received from the server. It checks the integrity of the audio data and prepares for playback. The input is the audio data received from the server, and the output is the audio data whose integrity has been checked.

[0201] Step 12:

[0202] The device plays back the audio data within the application, allowing the user to hear the voice of the deceased. The input is the audio data whose integrity has been verified, and the output is the played-back audio.

[0203] (Application example 1)

[0204] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0205] Technology that allows users to feel moved and comforted by re-listening to the voices of the deceased is extremely valuable. However, conventional systems have difficulty using the voice data of the deceased to read text aloud or generate audiobooks. This makes it difficult for users to experience the added emotional impact of listening to books or poems in the voice of the deceased.

[0206] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0207] In this invention, the server includes: [means for uploading the deceased's voice data from the user terminal to the server;] [means for analyzing the voice data received by the server and extracting feature data;] [means for training a generative AI model using the extracted feature data;] [means for the server to receive text selected by the user and convert the text into voice data using the trained generative AI model;] [means for the user terminal to receive voice data sent from the server and play it back to the user; and [means for sending the text selected by the user to the server and generating an audiobook in the deceased's voice. This allows the user to listen to books and poems in the deceased's voice.

[0208] "Voice data of the deceased" refers to information in the form of data that contains recorded or collected voices made by the deceased during their lifetime.

[0209] A "user terminal" is an information processing device such as a computer, smartphone, or tablet that is used by a user.

[0210] A "server" is a computer system that provides services over a network and sends and receives voice data from user terminals and analyzes it.

[0211] "Feature data" is data that indicates sound characteristics such as pitch, speed, and intonation extracted from voice data.

[0212] A "generative AI model" is an artificial intelligence model that learns the characteristics of voice data and is used to convert specified text into synthetic speech.

[0213] A "text message" is character string information entered by a user, which is received by the server and converted into voice data.

[0214] A "unique ID" is a unique identifier assigned to identify a particular piece of audio data within a database.

[0215] An "audiobook" is an audio version of the contents of a book or poem, and is audio data generated using the voice of a deceased person.

[0216] MODE FOR CARRYING OUT THE INVENTION

[0217] This invention is a system that allows users to feel touched or comforted by using the voice data of the deceased. The specific configuration and processing flow are described below.

[0218] System program and processing description

[0219] This system is constructed from a user device, a server, and a generative AI model. Each component plays the following role:

[0220] Hardware and software used:

[0221] User device: A smartphone, tablet, or computer is used to run applications, upload audio data, and play audiobooks.

[0222] Server: Located on the cloud, it is responsible for analyzing voice data, extracting features, and training and generating generative AI models.

[0223] Generative AI models: Convert text to audio using large-scale language models (e.g., GPT-3) and speech synthesis models (e.g., TTS models).

[0224] Cloud storage: Amazon Web Services (AWS) S3 and other cloud storage services are used to store audio data and generated audiobooks.

[0225] Data processing and data calculations:

[0226] 1. Uploading and sending audio data

[0227] The user uploads the deceased's audio data to the server from their device. First, the user selects an audio file (e.g., MP3, WAV) via a dedicated app and sends the data to the server. The server checks the format and quality of the received audio data and stores the appropriate audio data in cloud storage.

[0228] 2. Extracting feature data and training the AI ​​model

[0229] The server analyzes the stored audio data and extracts feature data. The feature extraction algorithm analyzes characteristics such as pitch, speed, and intonation of the voice, and uses this data to train a generative AI model. The trained model is optimized to reproduce the characteristics of the deceased person's voice.

[0230] 3. Entering text messages and generating voice data

[0231] The user selects the text of a book or poem they want to read through the application and sends it to the server, which analyzes the text and uses a generative AI model to convert the text into audio data in the deceased person's voice. The converted audiobook data is then stored in cloud storage.

[0232] 4. Receiving and playing audio data

[0233] The user device receives the generated audiobook from the server, verifies its integrity, and then plays it within the application, allowing the user to listen to books and poems in the voice of the deceased.

[0234] Examples:

[0235] For example, if a user uses a smartphone app to upload audio data of a deceased person, the audio data is uploaded to the app, and the smartphone sends the data to a server. The server analyzes the audio data, extracts the voice characteristics of the deceased person (e.g., grandmother), and trains an AI model. Next, when the user selects a poetry collection and enters it into the app, the server receives the poetry collection and uses the trained AI model to generate audio data in the deceased person's (grandmother's) voice. The generated audio data is sent to the smartphone and played back in the app. The user can relisten to the poetry collection in the deceased's voice, which can be moving and comforting.

[0236] Example prompt sentence:

[0237] "Generate an audiobook of this text read in the voice of the deceased: Complete Poems"

[0238] The above is a specific embodiment for carrying out the present invention. This system allows the user to re-create a heart-to-heart conversation with the deceased, and can bring about deep emotions and healing.

[0239] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0240] Step 1:

[0241] The user launches the dedicated application and clicks the "Upload audio data" button. The user selects the deceased's audio file (e.g., MP3, WAV) and starts uploading. The input is the audio file, and the output is the device sending the audio file to the server.

[0242] Step 2:

[0243] The terminal takes the selected audio file, checks the format and quality of the audio file, and if it matches, sends the audio file to the server. The input is the audio file, and the output is the audio data whose format and quality have been checked.

[0244] Step 3:

[0245] The server receives the audio file sent from the device. It analyzes the content of the audio file and checks that there are no problems with the format or quality. It stores the matching audio file in cloud storage and assigns a unique ID. The input is the audio file, and the output is the stored audio data and the unique ID.

[0246] Step 4:

[0247] The server analyzes the stored voice data and extracts feature data. Feature extraction algorithms are used to analyze the voice's pitch, speed, intonation, etc., and the data is used to train the generative AI model. The input is the stored voice data, and the output is the feature data.

[0248] Step 5:

[0249] The server uses the extracted feature data to train a generative AI model, which is optimized to reproduce the characteristics of the deceased person's voice. The input is the feature data, and the output is the trained generative AI model.

[0250] Step 6:

[0251] The user selects the text of a book or poem they want to read through the application and sends it to the server. The input is a text message, and the output is the sent text message.

[0252] Step 7:

[0253] The server analyzes the received text message and converts the text content into audio data using a trained generative AI model. The input is the text message and the output is the generated audio data.

[0254] Step 8:

[0255] The server stores the generated voice data in cloud storage and transmits it to the user's device. The input is the generated voice data, and the output is the transmitted voice data.

[0256] Step 9:

[0257] The device analyzes the audio data received from the server and verifies its integrity. The audio data is then played back within the application, allowing the user to hear the voice of the deceased. The input is the audio data received from the server, and the output is the played audio data.

[0258] The above are the specific processing steps of the system based on the application example, through which users can create audiobooks using the voices of deceased people and listen to them again.

[0259] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0260] This invention is a system that combines the voice data of the deceased with an emotion engine that recognizes the user's emotions, allowing bereaved family and friends to hear the deceased's voice again through AI and providing adapted responses based on the user's emotions. The program and specific processing flow of this system are shown below.

[0261] Program Overview

[0262] The program works through the following major steps:

[0263] 1. The user uploads the audio data of the deceased person.

[0264] 2. The device sends the audio data to the server.

[0265] 3. The server analyzes the voice data and extracts feature data.

[0266] 4. The server uses the feature data to train a generative AI model.

[0267] 5. The emotion engine recognizes the user's emotions.

[0268] 6. The user enters a text message, and the emotion engine sends the emotion analysis results to the server.

[0269] 7. The server uses the sentiment analysis results to generate audio data based on the text message.

[0270] 8. The device receives and plays the generated audio data.

[0271] Processing flow and specific examples

[0272] User voice data input

[0273] 1. The user launches the dedicated application and clicks the "Upload audio data" button.

[0274] 2. The user selects the audio file (e.g. MP3, WAV) of the deceased person and begins uploading.

[0275] Terminal handling

[0276] 3. The device will import the selected audio file and check the format and quality of the audio file.

[0277] Sending terminal

[0278] 4. If the format and quality match, the device sends the audio file to the server.

[0279] Validation and storage of voice data by the server

[0280] 1. The server receives the audio file sent from the device.

[0281] 2. The server analyzes the contents of the audio file to ensure there are no problems with the format or quality.

[0282] 3. If the audio file matches, the server stores the file in its database and assigns it a unique ID.

[0283] 4. The server notifies the device that the audio file has been saved.

[0284] Training generative AI models

[0285] 1. The server uses the stored audio files to characterize the deceased person's audio data.

[0286] 2. The server uses a feature extraction algorithm to analyze and extract features such as pitch, speed, and intonation of the speech.

[0287] 3. Train a generative AI model (e.g., a large-scale language model or a speech synthesis model) based on the extracted feature data.

[0288] 4. Once trained, the generative AI model is optimized to reproduce the vocal characteristics of the deceased person.

[0289] Emotion Engine Operation

[0290] 1. When a user uses an application, the emotion engine acquires the user's facial expression data and voice tone data through the camera and microphone.

[0291] 2. The emotion engine analyzes the acquired data and determines the user's emotional state (e.g., joy, sadness, surprise).

[0292] Enter and send text messages

[0293] 1. The user enters a text message through the application (e.g., "How have things been lately?").

[0294] 2. The emotion engine sends the user's text message along with the analyzed emotion results to the server.

[0295] Server-generated audio

[0296] 1. The server analyzes the received text message and the sentiment analysis results.

[0297] 2. The server inputs the analyzed text and emotion data into the generative AI model and generates corresponding voice data.

[0298] 3. Using the trained data, the generative AI model converts input text into audio data in the characteristic voice of the deceased based on the user's emotional state.

[0299] 4. The server sends the generated voice data to the terminal.

[0300] Receiving and playing audio on the device

[0301] 1. The device analyzes the voice data received from the server.

[0302] 2. The device checks the integrity of the audio data and prepares for playback.

[0303] 3. The device plays the audio data within the application, allowing the user to hear the voice of the deceased.

[0304] Specific examples

[0305] Consider a scenario where a user uses a smartphone app to upload audio data of their deceased grandmother. The audio data is uploaded to the app, and the smartphone sends the data to a server. The server analyzes the audio data, extracts the characteristics of the grandmother's voice, and trains an AI model. Next, when the user types a text message into the app saying, "Tell me about your favorite flower," the emotion engine analyzes the user's emotions and sends the emotion data to the server. Based on this information, the server uses the trained AI model to generate audio data in the grandmother's voice. The generated audio data is sent to the smartphone and played back by the app. The user can hear their grandmother's voice again, and experience deeper emotion and healing from the emotionally sensitive response.

[0306] The above is an example of the implementation of the present invention. This system recreates the heart-to-heart dialogue with the deceased and provides responses tailored to the user's emotions, allowing the bereaved to receive even deeper healing and empathy.

[0307] The processing flow will be explained below.

[0308] Step 1:

[0309] User

[0310] The user launches the dedicated application and clicks the "Upload audio data" button.

[0311] Step 2:

[0312] User

[0313] The user selects the deceased person's audio file (e.g. MP3, WAV) and begins uploading.

[0314] Step 3:

[0315] Terminal

[0316] The device captures the selected audio file and checks the format and quality of the audio file.

[0317] Step 4:

[0318] Terminal

[0319] If the format and quality match, the terminal sends the audio file to the server.

[0320] Step 5:

[0321] server

[0322] The server receives the audio file sent from the terminal.

[0323] Step 6:

[0324] server

[0325] The server analyzes the content of the audio file to ensure that there are no problems with the format or quality.

[0326] Step 7:

[0327] server

[0328] If the audio file matches, the server stores the file in a database and assigns it a unique ID.

[0329] Step 8:

[0330] server

[0331] The server notifies the terminal that the audio file has been saved.

[0332] Step 9:

[0333] server

[0334] The server uses the stored audio files to characterize the deceased person's audio data.

[0335] Step 10:

[0336] server

[0337] The server uses feature extraction algorithms to analyze and extract features such as pitch, speed, and intonation of the speech.

[0338] Step 11:

[0339] server

[0340] Based on the extracted feature data, generative AI models (e.g., large-scale language models or speech synthesis models) are trained.

[0341] Step 12:

[0342] server

[0343] Once trained, the generative AI model is optimized to reproduce the vocal characteristics of the deceased person.

[0344] Step 13:

[0345] Emotion Engine

[0346] When a user uses an application, the emotion engine acquires the user's facial expression data and voice tone data through a camera and microphone.

[0347] Step 14:

[0348] Emotion Engine

[0349] The emotion engine analyzes the acquired data and determines the user's emotional state (e.g., joy, sadness, surprise).

[0350] Step 15:

[0351] User

[0352] The user enters an arbitrary text message (e.g., "How have things been lately?") through the application.

[0353] Step 16:

[0354] User

[0355] The user clicks the "Send" button to send the text message to the server.

[0356] Step 17:

[0357] Emotion Engine

[0358] The emotion engine sends the user's text message along with the analyzed emotion results to the server.

[0359] Step 18:

[0360] server

[0361] The server analyzes the received text message and the sentiment analysis results.

[0362] Step 19:

[0363] server

[0364] The server inputs the analyzed text and emotion data into a generative AI model to generate corresponding voice data.

[0365] Step 20:

[0366] Generative AI Models

[0367] The generative AI model uses trained data to convert input text and emotional data into audio data in the characteristic voice of the deceased.

[0368] Step 21:

[0369] server

[0370] The server transmits the generated voice data to the terminal.

[0371] Step 22:

[0372] Terminal

[0373] The terminal analyzes the voice data received from the server.

[0374] Step 23:

[0375] Terminal

[0376] The device checks the integrity of the audio data and prepares it for playback.

[0377] Step 24:

[0378] Terminal

[0379] The device plays the audio data within the application, allowing the user to hear the voice of the deceased.

[0380] As a concrete example, consider the case where a user uses a smartphone app to upload voice data of their deceased grandmother and engages in emotionally appropriate conversations through an emotion engine. The voice file uploaded by the user is received by a server, which analyzes and extracts features. A generative AI model is trained based on this data. The user then performs emotion analysis through the emotion engine and enters and sends a text message. The data including the emotion analysis results is passed to the server, and emotion-appropriate voice data is generated in the grandmother's voice through the generative AI model. This voice data is received and played back on the device. The user can hear the grandmother's voice emotionally, and experience deep emotion and healing.

[0381] Example 2

[0382] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0383] For bereaved family members and friends who wish to reconnect with their loved ones, recreating their voice and receiving emotionally responsive responses is extremely valuable. However, simply playing back audio data does not provide a deep conversational experience with the deceased. Furthermore, current technology lacks a system that can recognize the user's emotions and respond accordingly.

[0384] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0385] In this invention, the server includes: [means for uploading the voice data of the deceased from the user terminal to the server; [means for analyzing the voice data received by the server and extracting feature data; [means for training a generative AI model using the extracted feature data; [means for an emotion engine to recognize the user's emotions and acquire facial expression data and voice tone data; [means for the server to receive a text message entered by the user and convert the text message into voice data using the emotion analysis results; and [means for the user terminal to receive the voice data sent from the server and play it back to the user. This enables the user to have a conversation with the deceased that is sensitive to their emotions.

[0386] "Voice data of a deceased person" refers to a digital file uploaded by a user that records the voice of a deceased person during their lifetime.

[0387] "User terminal" refers to electronic devices used by users, such as smartphones, tablets, and personal computers.

[0388] "Server" refers to a high-performance computer system for analyzing, storing, and generating voice data.

[0389] "Analysis" refers to the process of examining and examining the content of received voice data in detail to extract necessary feature data.

[0390] "Feature data" refers to information that quantifies the unique characteristics of voice data, such as voice pitch, speed, and intonation.

[0391] "Generative AI model" refers to an artificial intelligence model that generates voice based on feature data extracted from the voice data of a deceased person.

[0392] "Training" refers to the process of learning using data to enable a generative AI model to reproduce the vocal characteristics of the deceased person.

[0393] An "emotion engine" refers to an algorithm and system for recognizing and analyzing user emotions in real time.

[0394] "Facial expression data" refers to data that quantifies and classifies changes in a user's facial expressions captured through a camera.

[0395] "Voice tone data" refers to data obtained by quantifying and analyzing changes in the tone of a user's speech obtained through a microphone.

[0396] "Emotion analysis results" refers to data in which the emotion engine analyzes the user's emotional state and expresses the results in numerical values ​​and categories.

[0397] "Text message" refers to text information entered by a user through an application.

[0398] "Audio Data" refers to the digital audio files generated by the generative AI model that reproduce the voice of the deceased person.

[0399] This invention is a system that allows bereaved family and friends to hear the voice of the deceased again through AI by combining the voice data of the deceased with an emotion engine that recognizes the user's emotions. The system provides adaptive responses based on the emotions. Below, we will explain in detail how this system is implemented.

[0400] System configuration and processing overview

[0401] 1. User voice data input

[0402] The user launches the dedicated application and clicks the "Upload audio data" button.

[0403] The user selects the audio file of the deceased person (e.g., MP3, WAV format) and begins uploading.

[0404] 2. Terminal Processing

[0405] The device will then import the selected audio file and check the format and quality of the audio file using an audio analysis library (e.g., FFmpeg).

[0406] 3. Sending the device

[0407] If the format and quality match, the device sends the audio file to the server via an HTTP POST request.

[0408] 4. Verification and storage of audio data by the server

[0409] The server receives the audio file sent from the device, analyzes it again, and checks that there are no problems with the format or quality. If the audio file is suitable, the server stores it in a database (e.g., MongoDB) and assigns it a unique ID.

[0410] The server sends a storage completion notification to the terminal.

[0411] 5. Training the generative AI model

[0412] The server uses the stored audio files to perform a feature analysis of the deceased person's voice data. The feature analysis of the voice data is performed using a voice analysis tool (e.g., Praat, Kaldi).

[0413] The server uses feature extraction algorithms to extract features such as pitch, speed, and intonation, and then uses them to train a generative AI model (e.g., GPT-3, Tacotron 2).

[0414] 6. Operation of the Emotion Engine

[0415] The emotion engine captures facial expression data and voice tone data via a camera or microphone when a user uses an application, and analyzes the user's emotional state in real time using facial expression recognition algorithms (e.g., OpenCV, Dlib) and voice tone analysis tools (e.g., Praat).

[0416] 7. Enter and send text messages

[0417] The user types a text message through the application (e.g., "How's it been lately?").

[0418] The emotion engine sends a text message along with the analyzed emotion results to the server.

[0419] 8. Server-generated audio

[0420] The server analyzes the received text messages and emotion analysis results, inputs them into a generative AI model, and generates audio data in the voice of the deceased.

[0421] The server transmits the generated voice data to the terminal.

[0422] 9. Receiving and Playing Audio by the Device

[0423] The device analyzes the audio data received from the server, checks its integrity, and then plays the audio data within the application.

[0424] Specific use cases

[0425] For example, consider a case where a user uploads audio data of a deceased person using a smartphone app. The user uploads the audio data to the app, and the device sends the data to a server. The server analyzes the audio data, extracts the characteristics of the deceased's voice, and trains a generative AI model. Next, when the user types "Tell me about your favorite flower" into the app, the emotion engine analyzes the user's emotions and generates the most appropriate response based on those emotions. The generated audio data is sent to the smartphone, and the user can listen to the deceased's voice again in the app and enjoy a dialogue tailored to their emotions.

[0426] Prompt Sentence Examples

[0427] "Tell me about a flower you liked."

[0428] "How was your day today?"

[0429] "Tell me about a memory of going out together."

[0430] The above is a concrete embodiment for carrying out the present invention. Through this system, the user can hear the voice of the deceased again and receive emotional responses.

[0431] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0432] Step 1:

[0433] The user launches the dedicated application and clicks the "Upload audio data" button. The user selects the audio file (e.g., MP3, WAV) of the deceased person and begins uploading.

[0434] Input: User selection of an audio file.

[0435] Output: The selected audio file.

[0436] Step 2:

[0437] The device will then import the selected audio file and check the format and quality of the audio file, using an audio analysis library (e.g. FFmpeg) to analyze the bit rate, sampling rate, etc.

[0438] Input: The selected audio file.

[0439] Output: Data about the format and quality of the audio file.

[0440] Step 3:

[0441] If the audio file format and quality are compatible, the device sends the audio file to the server via an HTTP POST request.

[0442] Input: Quality checked audio files.

[0443] Output: The audio file sent to the server.

[0444] Step 4:

[0445] The server receives the audio file sent from the device and analyzes its content. It uses an audio analysis tool (e.g., Praat, Kaldi) to check for format and quality issues. If there are no problems, it stores the audio file in a database (e.g., MongoDB) and assigns a unique ID.

[0446] Input: The audio file sent from the device.

[0447] Output: The saved audio file and its unique ID.

[0448] Step 5:

[0449] The server uses the stored audio files to extract feature data. It uses speech analysis algorithms to analyze and extract features such as pitch, speed, and intonation. It then trains a generative AI model (e.g., GPT-3, Tacotron 2) based on the extracted feature data.

[0450] Input: Saved audio file.

[0451] Output: feature data and a trained generative AI model.

[0452] Step 6:

[0453] The emotion engine captures the user's facial expression data and voice tone data via the camera and microphone when the user uses the application, and performs real-time emotion analysis using facial expression recognition algorithms (e.g., OpenCV, Dlib) and voice tone analysis tools.

[0454] Input: User's facial expression data and voice tone data.

[0455] Output: User sentiment analysis results.

[0456] Step 7:

[0457] The user inputs a text message through the application (e.g., "How have things been lately?"). The emotion engine analyzes the text message and sends it to the server along with the resulting emotion.

[0458] Input: User text message and sentiment analysis results.

[0459] Output: The text message sent to the server and the sentiment result.

[0460] Step 8:

[0461] The server analyzes the received text message and the emotion analysis results, inputs them into a generative AI model, and generates audio data in the deceased's voice. The generated audio data is then sent from the server to the device.

[0462] Input: Text message, sentiment analysis results, pre-trained generative AI model.

[0463] Output: The generated audio data.

[0464] Step 9:

[0465] The device analyzes the audio data received from the server, checks its integrity, and then plays the audio data within the application.

[0466] Input: Audio data received from the server.

[0467] Output: The audio data that is played within the application.

[0468] (Application example 2)

[0469] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0470] In recent years, there has been an increasing need to cherish memories of deceased loved ones. However, conventional technologies have difficulty reproducing the voices of the deceased, and are unable to provide responses that are particularly sensitive to the user's emotions. Furthermore, more advanced technology is required to enable users to re-experience conversations with their deceased loved ones and achieve emotional healing. Furthermore, there is a need for a system that can provide a deeper experience in both real and virtual spaces by realizing this through specific devices.

[0471] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0472] In this invention, the server includes: [means for uploading the voice data of the deceased from the user terminal to the server; [means for analyzing the voice data received by the server and extracting feature data; [means for training a generative AI model using the extracted feature data; [means for the server to receive a text message entered by the user and convert the text message into voice data using the trained generative AI model; [means for the user terminal to receive the voice data sent from the server and play it back to the user; [means for an emotion engine to recognize the user's emotions and provide a response based on those emotions; and [means for playing back the response through a specific terminal (e.g., smartphone, smart glasses, etc.)]. This allows [the user to hear the voice of the deceased again and receive a response that is in tune with their emotions, thereby achieving deeper healing and empathy].

[0473] A "user terminal" is an electronic device operated by a user, and includes devices such as smartphones, tablets, and smart glasses.

[0474] A "server" is a central system that provides data and services over a network, and performs various processes such as analyzing, storing, and generating voice data.

[0475] "Audio data" refers to data in file format that records the voice of the deceased, and refers to information that includes the characteristics of the voice.

[0476] "Analysis" is the process in which the server analyzes the received voice data and extracts feature data.

[0477] "Feature data" refers to individual characteristic information such as pitch, speed, and intonation extracted from speech data.

[0478] A "generative AI model" is an artificial intelligence model trained based on extracted feature data, and is responsible for converting text messages into voice data.

[0479] A "text message" is a textual message entered by a user and sent to a server.

[0480] "Converting to voice data" refers to the process in which a generative AI model generates voice-format data based on the input text message.

[0481] An "emotion engine" is a system that recognizes a user's emotions and generates a response according to those emotions.

[0482] "Play" refers to the operation of the user terminal playing back the audio data received from the server as audio.

[0483] This invention is a system that reproduces the voice of a deceased person and provides responses that are in tune with the user's emotions. This system is realized mainly using a user terminal, a server, an emotion engine, and a generative AI model.

[0484] System configuration

[0485] User terminal

[0486] A user terminal is an electronic device operated by a user, including a smartphone, tablet, smart glasses, etc. This terminal is used by the user to upload voice data and perform subsequent interactions.

[0487] server

[0488] The server is a central system that provides data and services over the network. It is responsible for various processes such as analyzing, storing, and generating voice data. This server receives voice data sent from user devices, extracts feature data, and trains generative AI models. It is also responsible for the process of converting text messages entered by users into voice data.

[0489] Emotion Engine

[0490] The emotion engine is a system that recognizes the user's emotions and generates responses based on those emotions. The emotion engine acquires and analyzes the user's facial expressions and voice tone data via the camera and microphone on the user's device.

[0491] Generative AI Models

[0492] A generative AI model is an artificial intelligence model trained using extracted feature data. This model generates voice data that reproduces the characteristics of the deceased based on the input text message. Examples of such models include the Wav2Vec2 model and EmotionRecognizer.

[0493] Data processing and calculation

[0494] The system performs the following data processing and calculations:

[0495] 1. Analysis of speech data: The speech data of the deceased person uploaded from the user's device to the server is analyzed and feature data is extracted. This analysis is performed using models such as Wav2Vec2Processor and Wav2Vec2ForCTC.

[0496] 2. Emotion Recognition: The emotion engine analyzes voice tone and facial expression data to recognize the user's emotions. This analysis is performed using emotion recognition models such as EmotionRecognizer.

[0497] 3. Voice data generation: The server uses the generative AI model to generate voice data that reproduces the voice of the deceased person based on the text message entered by the user, matching the emotions analyzed by the emotion engine.

[0498] Specific examples

[0499] A specific usage scenario is when a user wears smart glasses at a memorial cafe and uploads their mother's voice data. The user enters a prompt such as, "Mom, tell me about your recent trip." In this case, the emotion engine analyzes the user's emotions, and the server generates an optimal response in the mother's voice. The user's device then receives the generated voice data and plays it back to the user, allowing the user to hear their mother's voice again. This allows for deep healing and empathy that is tailored to their emotions.

[0500] Prompt Sentence Examples

[0501] "Mom, tell me about your recent trip."

[0502] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0503] Step 1:

[0504] The user uploads the audio data of the deceased.

[0505] Description: The user launches the dedicated application and clicks the "Upload audio data" button. The user selects the audio file (e.g., MP3, WAV) of the deceased person and begins uploading.

[0506] Input: Audio file (e.g. MP3, WAV)

[0507] Output: Uploaded audio data

[0508] Step 2:

[0509] The terminal transmits the voice data to the server.

[0510] Description: The terminal captures the selected audio file, checks the format and quality of the audio file, and then sends the audio file to the server if it is suitable.

[0511] Input: Uploaded audio data

[0512] Output: Audio data sent to the server

[0513] Step 3:

[0514] The server analyzes the voice data and extracts feature data.

[0515] Description: The server analyzes the received audio data and extracts features such as pitch, speed, and intonation. The analysis is performed using Wav2Vec2Processor and Wav2Vec2ForCTC models.

[0516] Input: Audio data sent to the server

[0517] Output: feature data

[0518] Step 4:

[0519] The server uses the feature data to train a generative AI model.

[0520] Description: The server uses the extracted feature data to train a generative AI model, which is then optimized to reproduce the characteristics of the deceased person's voice.

[0521] Input: feature data

[0522] Output: A trained generative AI model

[0523] Step 5:

[0524] The emotion engine recognizes the user's emotions.

[0525] Description: When a user uses an application, the emotion engine captures and analyzes the user's facial expression data and voice tone data through the camera and microphone on the user's device. The analysis uses models such as EmotionRecognizer.

[0526] Input: User's facial expression data, voice tone data

[0527] Output: User emotion data

[0528] Step 6:

[0529] The user inputs a text message, and the emotion engine sends the emotion analysis results to the server.

[0530] Description: A user inputs any text message through the application. The emotion engine analyzes the user's text message and sends it to the server along with the emotion results.

[0531] Input: Text messages, emotion data

[0532] Output: Text message and emotion data sent to the server

[0533] Step 7:

[0534] The server generates voice data using the text message and the emotion data.

[0535] Description: The server inputs the received text message and emotional data into the generative AI model and generates corresponding voice data.

[0536] Input: Text messages, emotion data, pre-trained generative AI model

[0537] Output: Generated audio data

[0538] Step 8:

[0539] The terminal receives and plays the generated audio data.

[0540] Description: The device analyzes the audio data received from the server, checks its integrity, and then plays it back.

[0541] Input: Generated audio data

[0542] Output: Played audio

[0543] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0544] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0545] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0546] [Second embodiment]

[0547] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0548] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0549] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0550] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0551] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0552] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0553] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0554] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0555] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0556] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0557] In the smart glasses 214, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0558] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0559] This invention uses the voice data of the deceased to allow their family and friends to hear their voice again through AI. The program and specific processing flow of this system are shown below.

[0560] Program Overview

[0561] The program works through the following major steps:

[0562] 1. The user uploads the audio data of the deceased person.

[0563] 2. The device sends the audio data to the server.

[0564] 3. The server analyzes the voice data and extracts feature data.

[0565] 4. The server uses the feature data to train a generative AI model.

[0566] 5. The server receives the user's text message and generates audio data.

[0567] 6. The device receives and plays the generated audio data.

[0568] Processing flow and specific examples

[0569] User voice data input

[0570] 1. The user launches the dedicated application and clicks the "Upload audio data" button.

[0571] 2. The user selects the audio file (e.g. MP3, WAV) of the deceased person and begins uploading.

[0572] Terminal handling

[0573] 3. The device will import the selected audio file and check the format and quality of the audio file.

[0574] 4. If the format and quality match, the device sends the audio file to the server.

[0575] Validation and storage of voice data by the server

[0576] 1. The server receives the audio file sent from the device.

[0577] 2. The server analyzes the contents of the audio file to ensure there are no problems with the format or quality.

[0578] 3. If the audio file matches, the server stores the file in its database and assigns it a unique ID.

[0579] 4. The server notifies the device that the audio file has been saved.

[0580] Training generative AI models

[0581] 1. After the audio file is saved, the server performs a characteristic analysis of the deceased person's audio data.

[0582] 2. The server uses a feature extraction algorithm to analyze and extract features such as pitch, speed, and intonation of the speech.

[0583] 3. Train a generative AI model (e.g., a large-scale language model or a speech synthesis model) based on the extracted feature data.

[0584] 4. Once trained, the generative AI model is optimized to reproduce the vocal characteristics of the deceased person.

[0585] Enter and send text messages

[0586] 1. The user enters an arbitrary text message through the application (e.g., "How are you feeling today?").

[0587] 2. The user clicks the "Send" button to send the text message to the server.

[0588] Server-generated audio

[0589] 1. The server parses the received text message.

[0590] 2. The server inputs the analyzed text into a generative AI model and generates corresponding audio data.

[0591] 3. The generative AI model uses the trained data to convert the input text into audio data in the characteristic voice of the deceased.

[0592] 4. The server sends the generated voice data to the terminal.

[0593] Receiving and playing audio on the device

[0594] 1. The device analyzes the voice data received from the server.

[0595] 2. The device checks the integrity of the audio data and prepares for playback.

[0596] 3. The device plays the audio data within the application, allowing the user to hear the voice of the deceased.

[0597] Specific examples

[0598] Consider a case where a user uses a smartphone app to upload audio data of their deceased grandmother. The audio data is uploaded to the app, and the smartphone sends the data to a server. The server analyzes the audio data, extracts the characteristics of the grandmother's voice, and trains an AI model. Next, when the user types a text message into the app saying, "Tell me about your favorite flower," the server receives this message and uses the trained AI model to generate audio data in the grandmother's voice. The generated audio data is sent to the smartphone and played back by the app. The user can hear their grandmother's voice again, which can be moving and comforting.

[0599] The above is an example of the implementation of the present invention. This system allows the bereaved family to re-create a heart-to-heart conversation with the deceased, and to gain healing and empathy.

[0600] The processing flow will be explained below.

[0601] Step 1:

[0602] User

[0603] The user launches the dedicated application and clicks the "Upload audio data" button.

[0604] Step 2:

[0605] User

[0606] The user selects the deceased person's audio file (e.g. MP3, WAV) and begins uploading.

[0607] Step 3:

[0608] Terminal

[0609] The device captures the selected audio file and checks the format and quality of the audio file.

[0610] Step 4:

[0611] Terminal

[0612] If the format and quality match, the terminal sends the audio file to the server.

[0613] Step 5:

[0614] server

[0615] The server receives the audio file sent from the terminal.

[0616] Step 6:

[0617] server

[0618] The server analyzes the content of the audio file to ensure that there are no problems with the format or quality.

[0619] Step 7:

[0620] server

[0621] If the audio file matches, the server stores the file in a database and assigns it a unique ID.

[0622] Step 8:

[0623] server

[0624] The server notifies the terminal that the audio file has been saved.

[0625] Step 9:

[0626] server

[0627] The server uses the stored audio files to characterize the deceased person's audio data.

[0628] Step 10:

[0629] server

[0630] The server uses feature extraction algorithms to analyze and extract features such as pitch, speed, and intonation of the speech.

[0631] Step 11:

[0632] server

[0633] Based on the extracted feature data, generative AI models (e.g., large-scale language models or speech synthesis models) are trained.

[0634] Step 12:

[0635] server

[0636] Once trained, the generative AI model is optimized to reproduce the vocal characteristics of the deceased person.

[0637] Step 13:

[0638] User

[0639] Through the application, the user enters an arbitrary text message (e.g., "How are you feeling today?").

[0640] Step 14:

[0641] User

[0642] The user clicks the "Send" button to send the text message to the server.

[0643] Step 15:

[0644] Terminal

[0645] The terminal transmits the input text message to the server and makes a request to the server.

[0646] Step 16:

[0647] server

[0648] The server parses the received text message.

[0649] Step 17:

[0650] server

[0651] The server inputs the analyzed text into a generative AI model and generates corresponding audio data.

[0652] Step 18:

[0653] Generative AI Models

[0654] The generative AI model uses trained data to convert input text into audio data in the characteristic voice of the deceased.

[0655] Step 19:

[0656] server

[0657] The server transmits the generated voice data to the terminal.

[0658] Step 20:

[0659] Terminal

[0660] The terminal analyzes the voice data received from the server.

[0661] Step 21:

[0662] Terminal

[0663] The device checks the integrity of the audio data and prepares it for playback.

[0664] Step 22:

[0665] Terminal

[0666] The device plays the audio data within the application, allowing the user to hear the voice of the deceased.

[0667] Example 1

[0668] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0669] Technology that allows bereaved families and friends to hear the voices of the deceased again can bring emotion and healing, but conventional systems require a lot of work, such as uploading the voice data, checking the quality, and training the AI ​​model, making it difficult to operate efficiently.In addition, because the processes of checking the voice data quality, extracting features, and training the AI ​​model are complex, making it easy for users to use is a challenge.

[0670] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0671] In this invention, the server includes a means for a user to upload voice data of the deceased via a terminal, a means for the terminal to check the format and quality of the voice data and send it to the server, and a means for the server to analyze the received voice data and extract feature data, thereby enabling a user to efficiently and easily use the voice data of the deceased and reproduce the deceased's voice using an AI model.

[0672] A "user" is a person who uses the system to upload and play back audio data of a deceased person.

[0673] A "terminal" is a device operated by a user that uploads audio data, checks the format and quality, transmits it to a server, and plays back the received audio data.

[0674] A "server" is a device that analyzes received voice data, extracts feature data, trains a generative AI model, generates voice data, and transmits it to a terminal.

[0675] "Audio data" refers to a digital data file (e.g., MP3, WAV) that records the voice of a deceased person.

[0676] "Uploading" is the act of a user sending audio data to a server via a terminal.

[0677] "Format" refers to the format of the audio data file (e.g., MP3, WAV).

[0678] "Quality" refers to characteristics related to the sound quality of data, such as the sampling rate and bit depth of the audio data.

[0679] "Analysis" is the process by which the server performs detailed analysis of the received audio data to extract specific information.

[0680] "Feature data" is information that indicates voice characteristics such as pitch, speed, and intonation extracted from voice data.

[0681] A "generative AI model" is an artificial intelligence model trained to convert text messages into audio data based on feature data.

[0682] "Training" is the process of learning from feature data so that the generative AI model can accurately reproduce the voice of the deceased.

[0683] A "text message" is text information that is entered by a user and sent to a server.

[0684] "Conversion" is the process of using a generative AI model to turn a text message into audio data.

[0685] "Reception" refers to the act of the server receiving data transmitted from the user terminal, or the act of the user terminal receiving data transmitted from the server.

[0686] "Playback" refers to the act of outputting audio data received by a user terminal as audio so that the user can hear it.

[0687] This invention is a system that uses the voice data of the deceased to allow bereaved family members and friends to hear the voice of the deceased again through AI. This system consists of three main elements: a user, a terminal, and a server. The specific configuration and embodiments of this system are described below.

[0688] First, the user launches a dedicated application on their device. This device can be a smartphone, tablet, or computer, which is capable of importing and playing audio data. When the user clicks the "Upload Audio Data" button, a screen appears allowing them to select the audio file (e.g., MP3, WAV) of the deceased person. The user selects the target audio file and begins uploading.

[0689] The device then retrieves the selected audio file and checks its format and quality. This check checks whether the audio file is in MP3 or WAV format and meets certain audio quality standards (e.g., 44.1kHz, 16-bit). If the quality is acceptable, the device sends the audio file to the server. During this process, the user sees an "Uploading" progress bar.

[0690] When the server receives the audio file sent from the device, it checks the format and quality of the audio data again. If there are no problems, the server saves the audio file in its database and assigns a unique identifier (ID) to the file. The server notifies the device that the file has been saved, allowing the user to confirm that the upload was successful.

[0691] The server extracts feature data from the stored audio files. Specifically, it uses a feature extraction algorithm to analyze and extract features such as pitch, speed, and intonation. This feature data is used to train a generative AI model (e.g., a large-scale language model or a speech synthesis model). This model is optimized to reproduce the characteristics of the deceased person's voice.

[0692] If a user wants to use the voice of a deceased person, they can enter any text message through the application, for example, "How are you feeling today?" or "Tell me about a flower you liked," and click the "Send" button. This message will be sent to the server.

[0693] The server analyzes the received text message and inputs it into a generative AI model to generate corresponding audio data. This AI model converts the input text into audio data in the deceased's characteristic voice based on the trained audio feature data. The generated audio data is then sent to the device.

[0694] The audio data received by the device is checked for integrity and then prepared for playback. Finally, the audio data is played back within the application, allowing the user to hear the voice of the deceased. This recreates an emotional conversation with the deceased, bringing emotion and healing to bereaved family and friends.

[0695] Specifically, when a user uses a smartphone app to upload audio data of their deceased grandmother, the audio data is uploaded to the app, and the smartphone sends the data to a server. The server analyzes the audio data, extracts the characteristics of the grandmother's voice, and trains an AI model. When the user types a text message into the app saying, "Tell me about your favorite flower," the server receives the message and uses the trained AI model to generate audio data in the grandmother's voice. The generated audio data is sent to the smartphone and played back in the app. The user can hear their grandmother's voice again, and feel moved and comforted.

[0696] An example of a prompt is as follows:

[0697] Prompt: "Tell me about a flower you loved."

[0698] Generated text: "Hello. I love talking about flowers you liked. My favorite was roses..."

[0699] This invention allows users to relive the heartfelt conversation by listening to the voice of the deceased again, and thus find healing.

[0700] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0701] Step 1:

[0702] The user launches the dedicated application on the device and clicks the "Upload audio data" button. When the user selects an audio file to upload, this file is imported into the device. The input is the audio file selected by the user, and the output is the imported audio file.

[0703] Step 2:

[0704] The device checks the format and quality of the selected audio file. Specifically, it checks that the file is in MP3 or WAV format and that the audio quality is 44.1 kHz and 16-bit. The input is the imported audio file, and the output is the result of matching the format and quality.

[0705] Step 3:

[0706] The device checks the format and quality and, if it is found to be compatible, sends the audio file to the server. A progress bar indicating "uploading" is displayed to the user during the upload. The input is the format and quality check result and the audio file, and the output is the audio file sent to the server.

[0707] Step 4:

[0708] The server receives the audio file sent from the terminal. The input is the audio file sent from the terminal, and the output is the received audio file.

[0709] Step 5:

[0710] The server rechecks the content of the received audio file and, if there are no problems, stores it in a database. Specifically, it analyzes the format and quality of the audio data again and assigns a unique identifier. The input is the received audio file, and the output is the audio file stored in the database and its unique identifier.

[0711] Step 6:

[0712] The server performs feature analysis on the audio files stored in the database. It uses a feature extraction algorithm to analyze and extract features such as pitch, speed, and intonation of the voice. Based on this analysis, feature data is generated. The input is the audio files stored in the database, and the output is the extracted feature data.

[0713] Step 7:

[0714] The server uses the extracted feature data to train a generative AI model. Specifically, the feature data is input into the AI ​​model and the AI ​​model is optimized through an appropriate learning process. The input is the feature data, and the output is a trained generative AI model.

[0715] Step 8:

[0716] A user enters a text message through an application. The user clicks a "Send" button to send the entered text message to a server. The input is the text message entered by the user, and the output is the text message sent to the server.

[0717] Step 9:

[0718] The server analyzes the received text message. The analyzed text message is input into the generative AI model to generate corresponding audio data. The input is the text message sent to the server, and the output is the generated audio data.

[0719] Step 10:

[0720] The server sends the generated voice data to the terminal. The input is the generated voice data, and the output is the voice data sent to the terminal.

[0721] Step 11:

[0722] The device analyzes the audio data received from the server. It checks the integrity of the audio data and prepares for playback. The input is the audio data received from the server, and the output is the audio data whose integrity has been checked.

[0723] Step 12:

[0724] The device plays back the audio data within the application, allowing the user to hear the voice of the deceased. The input is the audio data whose integrity has been verified, and the output is the played-back audio.

[0725] (Application example 1)

[0726] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0727] Technology that allows users to feel moved and comforted by re-listening to the voices of the deceased is extremely valuable. However, conventional systems have difficulty using the voice data of the deceased to read text aloud or generate audiobooks. This makes it difficult for users to experience the added emotional impact of listening to books or poems in the voice of the deceased.

[0728] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0729] In this invention, the server includes: [means for uploading the deceased's voice data from the user terminal to the server;] [means for analyzing the voice data received by the server and extracting feature data;] [means for training a generative AI model using the extracted feature data;] [means for the server to receive text selected by the user and convert the text into voice data using the trained generative AI model;] [means for the user terminal to receive voice data sent from the server and play it back to the user; and [means for sending the text selected by the user to the server and generating an audiobook in the deceased's voice. This allows the user to listen to books and poems in the deceased's voice.

[0730] "Voice data of the deceased" refers to information in the form of data that contains recorded or collected voices made by the deceased during their lifetime.

[0731] A "user terminal" is an information processing device such as a computer, smartphone, or tablet that is used by a user.

[0732] A "server" is a computer system that provides services over a network and sends and receives voice data from user terminals and analyzes it.

[0733] "Feature data" is data that indicates sound characteristics such as pitch, speed, and intonation extracted from voice data.

[0734] A "generative AI model" is an artificial intelligence model that learns the characteristics of voice data and is used to convert specified text into synthetic speech.

[0735] A "text message" is character string information entered by a user, which is received by the server and converted into voice data.

[0736] A "unique ID" is a unique identifier assigned to identify a particular piece of audio data within a database.

[0737] An "audiobook" is an audio version of the contents of a book or poem, and is audio data generated using the voice of a deceased person.

[0738] MODE FOR CARRYING OUT THE INVENTION

[0739] This invention is a system that allows users to feel touched or comforted by using the voice data of the deceased. The specific configuration and processing flow are described below.

[0740] System program and processing description

[0741] This system is constructed from a user device, a server, and a generative AI model. Each component plays the following role:

[0742] Hardware and software used:

[0743] User device: A smartphone, tablet, or computer is used to run applications, upload audio data, and play audiobooks.

[0744] Server: Located on the cloud, it is responsible for analyzing voice data, extracting features, and training and generating generative AI models.

[0745] Generative AI models: Convert text to audio using large-scale language models (e.g., GPT-3) and speech synthesis models (e.g., TTS models).

[0746] Cloud storage: Amazon Web Services (AWS) S3 and other cloud storage services are used to store audio data and generated audiobooks.

[0747] Data processing and data calculations:

[0748] 1. Uploading and sending audio data

[0749] The user uploads the deceased's audio data to the server from their device. First, the user selects an audio file (e.g., MP3, WAV) via a dedicated app and sends the data to the server. The server checks the format and quality of the received audio data and stores the appropriate audio data in cloud storage.

[0750] 2. Extracting feature data and training the AI ​​model

[0751] The server analyzes the stored audio data and extracts feature data. The feature extraction algorithm analyzes characteristics such as pitch, speed, and intonation of the voice, and uses this data to train a generative AI model. The trained model is optimized to reproduce the characteristics of the deceased person's voice.

[0752] 3. Entering text messages and generating voice data

[0753] The user selects the text of a book or poem they want to read through the application and sends it to the server, which analyzes the text and uses a generative AI model to convert the text into audio data in the deceased person's voice. The converted audiobook data is then stored in cloud storage.

[0754] 4. Receiving and playing audio data

[0755] The user device receives the generated audiobook from the server, verifies its integrity, and then plays it within the application, allowing the user to listen to books and poems in the voice of the deceased.

[0756] Examples:

[0757] For example, if a user uses a smartphone app to upload audio data of a deceased person, the audio data is uploaded to the app, and the smartphone sends the data to a server. The server analyzes the audio data, extracts the voice characteristics of the deceased person (e.g., grandmother), and trains an AI model. Next, when the user selects a poetry collection and enters it into the app, the server receives the poetry collection and uses the trained AI model to generate audio data in the deceased person's (grandmother's) voice. The generated audio data is sent to the smartphone and played back in the app. The user can relisten to the poetry collection in the deceased's voice, which can be moving and comforting.

[0758] Example prompt sentence:

[0759] "Generate an audiobook of this text read in the voice of the deceased: Complete Poems"

[0760] The above is a specific embodiment for carrying out the present invention. This system allows the user to re-create a heart-to-heart conversation with the deceased, and can bring about deep emotions and healing.

[0761] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0762] Step 1:

[0763] The user launches the dedicated application and clicks the "Upload audio data" button. The user selects the deceased's audio file (e.g., MP3, WAV) and starts uploading. The input is the audio file, and the output is the device sending the audio file to the server.

[0764] Step 2:

[0765] The terminal takes the selected audio file, checks the format and quality of the audio file, and if it matches, sends the audio file to the server. The input is the audio file, and the output is the audio data whose format and quality have been checked.

[0766] Step 3:

[0767] The server receives the audio file sent from the device. It analyzes the content of the audio file and checks that there are no problems with the format or quality. It stores the matching audio file in cloud storage and assigns a unique ID. The input is the audio file, and the output is the stored audio data and the unique ID.

[0768] Step 4:

[0769] The server analyzes the stored voice data and extracts feature data. Feature extraction algorithms are used to analyze the voice's pitch, speed, intonation, etc., and the data is used to train the generative AI model. The input is the stored voice data, and the output is the feature data.

[0770] Step 5:

[0771] The server uses the extracted feature data to train a generative AI model, which is optimized to reproduce the characteristics of the deceased person's voice. The input is the feature data, and the output is the trained generative AI model.

[0772] Step 6:

[0773] The user selects the text of a book or poem they want to read through the application and sends it to the server. The input is a text message, and the output is the sent text message.

[0774] Step 7:

[0775] The server analyzes the received text message and converts the text content into audio data using a trained generative AI model. The input is the text message and the output is the generated audio data.

[0776] Step 8:

[0777] The server stores the generated voice data in cloud storage and transmits it to the user's device. The input is the generated voice data, and the output is the transmitted voice data.

[0778] Step 9:

[0779] The device analyzes the audio data received from the server and verifies its integrity. The audio data is then played back within the application, allowing the user to hear the voice of the deceased. The input is the audio data received from the server, and the output is the played audio data.

[0780] The above are the specific processing steps of the system based on the application example, through which users can create audiobooks using the voices of deceased people and listen to them again.

[0781] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0782] This invention is a system that combines the voice data of the deceased with an emotion engine that recognizes the user's emotions, allowing bereaved family and friends to hear the deceased's voice again through AI and providing adapted responses based on the user's emotions. The program and specific processing flow of this system are shown below.

[0783] Program Overview

[0784] The program works through the following major steps:

[0785] 1. The user uploads the audio data of the deceased person.

[0786] 2. The device sends the audio data to the server.

[0787] 3. The server analyzes the voice data and extracts feature data.

[0788] 4. The server uses the feature data to train a generative AI model.

[0789] 5. The emotion engine recognizes the user's emotions.

[0790] 6. The user enters a text message, and the emotion engine sends the emotion analysis results to the server.

[0791] 7. The server uses the sentiment analysis results to generate audio data based on the text message.

[0792] 8. The device receives and plays the generated audio data.

[0793] Processing flow and specific examples

[0794] User voice data input

[0795] 1. The user launches the dedicated application and clicks the "Upload audio data" button.

[0796] 2. The user selects the audio file (e.g. MP3, WAV) of the deceased person and begins uploading.

[0797] Terminal handling

[0798] 3. The device will import the selected audio file and check the format and quality of the audio file.

[0799] Sending terminal

[0800] 4. If the format and quality match, the device sends the audio file to the server.

[0801] Validation and storage of voice data by the server

[0802] 1. The server receives the audio file sent from the device.

[0803] 2. The server analyzes the contents of the audio file to ensure there are no problems with the format or quality.

[0804] 3. If the audio file matches, the server stores the file in its database and assigns it a unique ID.

[0805] 4. The server notifies the device that the audio file has been saved.

[0806] Training generative AI models

[0807] 1. The server uses the stored audio files to characterize the deceased person's audio data.

[0808] 2. The server uses a feature extraction algorithm to analyze and extract features such as pitch, speed, and intonation of the speech.

[0809] 3. Train a generative AI model (e.g., a large-scale language model or a speech synthesis model) based on the extracted feature data.

[0810] 4. Once trained, the generative AI model is optimized to reproduce the vocal characteristics of the deceased person.

[0811] Emotion Engine Operation

[0812] 1. When a user uses an application, the emotion engine acquires the user's facial expression data and voice tone data through the camera and microphone.

[0813] 2. The emotion engine analyzes the acquired data and determines the user's emotional state (e.g., joy, sadness, surprise).

[0814] Enter and send text messages

[0815] 1. The user enters a text message through the application (e.g., "How have things been lately?").

[0816] 2. The emotion engine sends the user's text message along with the analyzed emotion results to the server.

[0817] Server-generated audio

[0818] 1. The server analyzes the received text message and the sentiment analysis results.

[0819] 2. The server inputs the analyzed text and emotion data into the generative AI model and generates corresponding voice data.

[0820] 3. Using the trained data, the generative AI model converts input text into audio data in the characteristic voice of the deceased based on the user's emotional state.

[0821] 4. The server sends the generated voice data to the terminal.

[0822] Receiving and playing audio on the device

[0823] 1. The device analyzes the voice data received from the server.

[0824] 2. The device checks the integrity of the audio data and prepares for playback.

[0825] 3. The device plays the audio data within the application, allowing the user to hear the voice of the deceased.

[0826] Specific examples

[0827] Consider a scenario where a user uses a smartphone app to upload audio data of their deceased grandmother. The audio data is uploaded to the app, and the smartphone sends the data to a server. The server analyzes the audio data, extracts the characteristics of the grandmother's voice, and trains an AI model. Next, when the user types a text message into the app saying, "Tell me about your favorite flower," the emotion engine analyzes the user's emotions and sends the emotion data to the server. Based on this information, the server uses the trained AI model to generate audio data in the grandmother's voice. The generated audio data is sent to the smartphone and played back by the app. The user can hear their grandmother's voice again, and experience deeper emotion and healing from the emotionally sensitive response.

[0828] The above is an example of the implementation of the present invention. This system recreates the heart-to-heart dialogue with the deceased and provides responses tailored to the user's emotions, allowing the bereaved to receive even deeper healing and empathy.

[0829] The processing flow will be explained below.

[0830] Step 1:

[0831] User

[0832] The user launches the dedicated application and clicks the "Upload audio data" button.

[0833] Step 2:

[0834] User

[0835] The user selects the deceased person's audio file (e.g. MP3, WAV) and begins uploading.

[0836] Step 3:

[0837] Terminal

[0838] The device captures the selected audio file and checks the format and quality of the audio file.

[0839] Step 4:

[0840] Terminal

[0841] If the format and quality match, the terminal sends the audio file to the server.

[0842] Step 5:

[0843] server

[0844] The server receives the audio file sent from the terminal.

[0845] Step 6:

[0846] server

[0847] The server analyzes the content of the audio file to ensure that there are no problems with the format or quality.

[0848] Step 7:

[0849] server

[0850] If the audio file matches, the server stores the file in a database and assigns it a unique ID.

[0851] Step 8:

[0852] server

[0853] The server notifies the terminal that the audio file has been saved.

[0854] Step 9:

[0855] server

[0856] The server uses the stored audio files to characterize the deceased person's audio data.

[0857] Step 10:

[0858] server

[0859] The server uses feature extraction algorithms to analyze and extract features such as pitch, speed, and intonation of the speech.

[0860] Step 11:

[0861] server

[0862] Based on the extracted feature data, generative AI models (e.g., large-scale language models or speech synthesis models) are trained.

[0863] Step 12:

[0864] server

[0865] Once trained, the generative AI model is optimized to reproduce the vocal characteristics of the deceased person.

[0866] Step 13:

[0867] Emotion Engine

[0868] When a user uses an application, the emotion engine acquires the user's facial expression data and voice tone data through a camera and microphone.

[0869] Step 14:

[0870] Emotion Engine

[0871] The emotion engine analyzes the acquired data and determines the user's emotional state (e.g., joy, sadness, surprise).

[0872] Step 15:

[0873] User

[0874] The user enters an arbitrary text message (e.g., "How have things been lately?") through the application.

[0875] Step 16:

[0876] User

[0877] The user clicks the "Send" button to send the text message to the server.

[0878] Step 17:

[0879] Emotion Engine

[0880] The emotion engine sends the user's text message along with the analyzed emotion results to the server.

[0881] Step 18:

[0882] server

[0883] The server analyzes the received text message and the sentiment analysis results.

[0884] Step 19:

[0885] server

[0886] The server inputs the analyzed text and emotion data into a generative AI model to generate corresponding voice data.

[0887] Step 20:

[0888] Generative AI Models

[0889] The generative AI model uses trained data to convert input text and emotional data into audio data in the characteristic voice of the deceased.

[0890] Step 21:

[0891] server

[0892] The server transmits the generated voice data to the terminal.

[0893] Step 22:

[0894] Terminal

[0895] The terminal analyzes the voice data received from the server.

[0896] Step 23:

[0897] Terminal

[0898] The device checks the integrity of the audio data and prepares it for playback.

[0899] Step 24:

[0900] Terminal

[0901] The device plays the audio data within the application, allowing the user to hear the voice of the deceased.

[0902] As a concrete example, consider the case where a user uses a smartphone app to upload voice data of their deceased grandmother and engages in emotionally appropriate conversations through an emotion engine. The voice file uploaded by the user is received by a server, which analyzes and extracts features. A generative AI model is trained based on this data. The user then performs emotion analysis through the emotion engine and enters and sends a text message. The data including the emotion analysis results is passed to the server, and emotion-appropriate voice data is generated in the grandmother's voice through the generative AI model. This voice data is received and played back on the device. The user can hear the grandmother's voice emotionally, and experience deep emotion and healing.

[0903] Example 2

[0904] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0905] For bereaved family members and friends who wish to reconnect with their loved ones, recreating their voice and receiving emotionally responsive responses is extremely valuable. However, simply playing back audio data does not provide a deep conversational experience with the deceased. Furthermore, current technology lacks a system that can recognize the user's emotions and respond accordingly.

[0906] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0907] In this invention, the server includes: [means for uploading the voice data of the deceased from the user terminal to the server; [means for analyzing the voice data received by the server and extracting feature data; [means for training a generative AI model using the extracted feature data; [means for an emotion engine to recognize the user's emotions and acquire facial expression data and voice tone data; [means for the server to receive a text message entered by the user and convert the text message into voice data using the emotion analysis results; and [means for the user terminal to receive the voice data sent from the server and play it back to the user. This enables the user to have a conversation with the deceased that is sensitive to their emotions.

[0908] "Voice data of a deceased person" refers to a digital file uploaded by a user that records the voice of a deceased person during their lifetime.

[0909] "User terminal" refers to electronic devices used by users, such as smartphones, tablets, and personal computers.

[0910] "Server" refers to a high-performance computer system for analyzing, storing, and generating voice data.

[0911] "Analysis" refers to the process of examining and examining the content of received voice data in detail to extract necessary feature data.

[0912] "Feature data" refers to information that quantifies the unique characteristics of voice data, such as voice pitch, speed, and intonation.

[0913] "Generative AI model" refers to an artificial intelligence model that generates voice based on feature data extracted from the voice data of a deceased person.

[0914] "Training" refers to the process of learning using data to enable a generative AI model to reproduce the vocal characteristics of the deceased person.

[0915] An "emotion engine" refers to an algorithm and system for recognizing and analyzing user emotions in real time.

[0916] "Facial expression data" refers to data that quantifies and classifies changes in a user's facial expressions captured through a camera.

[0917] "Voice tone data" refers to data obtained by quantifying and analyzing changes in the tone of a user's speech obtained through a microphone.

[0918] "Emotion analysis results" refers to data in which the emotion engine analyzes the user's emotional state and expresses the results in numerical values ​​and categories.

[0919] "Text message" refers to text information entered by a user through an application.

[0920] "Audio Data" refers to the digital audio files generated by the generative AI model that reproduce the voice of the deceased person.

[0921] This invention is a system that allows bereaved family and friends to hear the voice of the deceased again through AI by combining the voice data of the deceased with an emotion engine that recognizes the user's emotions. The system provides adaptive responses based on the emotions. Below, we will explain in detail how this system is implemented.

[0922] System configuration and processing overview

[0923] 1. User voice data input

[0924] The user launches the dedicated application and clicks the "Upload audio data" button.

[0925] The user selects the audio file of the deceased person (e.g., MP3, WAV format) and begins uploading.

[0926] 2. Terminal Processing

[0927] The device will then import the selected audio file and check the format and quality of the audio file using an audio analysis library (e.g., FFmpeg).

[0928] 3. Sending the device

[0929] If the format and quality match, the device sends the audio file to the server via an HTTP POST request.

[0930] 4. Verification and storage of audio data by the server

[0931] The server receives the audio file sent from the device, analyzes it again, and checks that there are no problems with the format or quality. If the audio file is suitable, the server stores it in a database (e.g., MongoDB) and assigns it a unique ID.

[0932] The server sends a storage completion notification to the terminal.

[0933] 5. Training the generative AI model

[0934] The server uses the stored audio files to perform a feature analysis of the deceased person's voice data. The feature analysis of the voice data is performed using a voice analysis tool (e.g., Praat, Kaldi).

[0935] The server uses feature extraction algorithms to extract features such as pitch, speed, and intonation, and then uses them to train a generative AI model (e.g., GPT-3, Tacotron 2).

[0936] 6. Operation of the Emotion Engine

[0937] The emotion engine captures facial expression data and voice tone data via a camera or microphone when a user uses an application, and analyzes the user's emotional state in real time using facial expression recognition algorithms (e.g., OpenCV, Dlib) and voice tone analysis tools (e.g., Praat).

[0938] 7. Enter and send text messages

[0939] The user types a text message through the application (e.g., "How's it been lately?").

[0940] The emotion engine sends a text message along with the analyzed emotion results to the server.

[0941] 8. Server-generated audio

[0942] The server analyzes the received text messages and emotion analysis results, inputs them into a generative AI model, and generates audio data in the voice of the deceased.

[0943] The server transmits the generated voice data to the terminal.

[0944] 9. Receiving and Playing Audio by the Device

[0945] The device analyzes the audio data received from the server, checks its integrity, and then plays the audio data within the application.

[0946] Specific use cases

[0947] For example, consider a case where a user uploads audio data of a deceased person using a smartphone app. The user uploads the audio data to the app, and the device sends the data to a server. The server analyzes the audio data, extracts the characteristics of the deceased's voice, and trains a generative AI model. Next, when the user types "Tell me about your favorite flower" into the app, the emotion engine analyzes the user's emotions and generates the most appropriate response based on those emotions. The generated audio data is sent to the smartphone, and the user can listen to the deceased's voice again in the app and enjoy a dialogue tailored to their emotions.

[0948] Prompt Sentence Examples

[0949] "Tell me about a flower you liked."

[0950] "How was your day today?"

[0951] "Tell me about a memory of going out together."

[0952] The above is a concrete embodiment for carrying out the present invention. Through this system, the user can hear the voice of the deceased again and receive emotional responses.

[0953] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0954] Step 1:

[0955] The user launches the dedicated application and clicks the "Upload audio data" button. The user selects the audio file (e.g., MP3, WAV) of the deceased person and begins uploading.

[0956] Input: User selection of an audio file.

[0957] Output: The selected audio file.

[0958] Step 2:

[0959] The device will then import the selected audio file and check the format and quality of the audio file, using an audio analysis library (e.g. FFmpeg) to analyze the bit rate, sampling rate, etc.

[0960] Input: The selected audio file.

[0961] Output: Data about the format and quality of the audio file.

[0962] Step 3:

[0963] If the audio file format and quality are compatible, the device sends the audio file to the server via an HTTP POST request.

[0964] Input: Quality checked audio files.

[0965] Output: The audio file sent to the server.

[0966] Step 4:

[0967] The server receives the audio file sent from the device and analyzes its content. It uses an audio analysis tool (e.g., Praat, Kaldi) to check for format and quality issues. If there are no problems, it stores the audio file in a database (e.g., MongoDB) and assigns a unique ID.

[0968] Input: The audio file sent from the device.

[0969] Output: The saved audio file and its unique ID.

[0970] Step 5:

[0971] The server uses the stored audio files to extract feature data. It uses speech analysis algorithms to analyze and extract features such as pitch, speed, and intonation. It then trains a generative AI model (e.g., GPT-3, Tacotron 2) based on the extracted feature data.

[0972] Input: Saved audio file.

[0973] Output: feature data and a trained generative AI model.

[0974] Step 6:

[0975] The emotion engine captures the user's facial expression data and voice tone data via the camera and microphone when the user uses the application, and performs real-time emotion analysis using facial expression recognition algorithms (e.g., OpenCV, Dlib) and voice tone analysis tools.

[0976] Input: User's facial expression data and voice tone data.

[0977] Output: User sentiment analysis results.

[0978] Step 7:

[0979] The user inputs a text message through the application (e.g., "How have things been lately?"). The emotion engine analyzes the text message and sends it to the server along with the resulting emotion.

[0980] Input: User text message and sentiment analysis results.

[0981] Output: The text message sent to the server and the sentiment result.

[0982] Step 8:

[0983] The server analyzes the received text message and the emotion analysis results, inputs them into a generative AI model, and generates audio data in the deceased's voice. The generated audio data is then sent from the server to the device.

[0984] Input: Text message, sentiment analysis results, pre-trained generative AI model.

[0985] Output: The generated audio data.

[0986] Step 9:

[0987] The device analyzes the audio data received from the server, checks its integrity, and then plays the audio data within the application.

[0988] Input: Audio data received from the server.

[0989] Output: The audio data that is played within the application.

[0990] (Application example 2)

[0991] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0992] In recent years, there has been an increasing need to cherish memories of deceased loved ones. However, conventional technologies have difficulty reproducing the voices of the deceased, and are unable to provide responses that are particularly sensitive to the user's emotions. Furthermore, more advanced technology is required to enable users to re-experience conversations with their deceased loved ones and achieve emotional healing. Furthermore, there is a need for a system that can provide a deeper experience in both real and virtual spaces by realizing this through specific devices.

[0993] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0994] In this invention, the server includes: [means for uploading the voice data of the deceased from the user terminal to the server; [means for analyzing the voice data received by the server and extracting feature data; [means for training a generative AI model using the extracted feature data; [means for the server to receive a text message entered by the user and convert the text message into voice data using the trained generative AI model; [means for the user terminal to receive the voice data sent from the server and play it back to the user; [means for an emotion engine to recognize the user's emotions and provide a response based on those emotions; and [means for playing back the response through a specific terminal (e.g., smartphone, smart glasses, etc.)]. This allows [the user to hear the voice of the deceased again and receive a response that is in tune with their emotions, thereby achieving deeper healing and empathy].

[0995] A "user terminal" is an electronic device operated by a user, and includes devices such as smartphones, tablets, and smart glasses.

[0996] A "server" is a central system that provides data and services over a network, and performs various processes such as analyzing, storing, and generating voice data.

[0997] "Audio data" refers to data in file format that records the voice of the deceased, and refers to information that includes the characteristics of the voice.

[0998] "Analysis" is the process in which the server analyzes the received voice data and extracts feature data.

[0999] "Feature data" refers to individual characteristic information such as pitch, speed, and intonation extracted from speech data.

[1000] A "generative AI model" is an artificial intelligence model trained based on extracted feature data, and is responsible for converting text messages into voice data.

[1001] A "text message" is a textual message entered by a user and sent to a server.

[1002] "Converting to voice data" refers to the process in which a generative AI model generates voice-format data based on the input text message.

[1003] An "emotion engine" is a system that recognizes a user's emotions and generates a response according to those emotions.

[1004] "Play" refers to the operation of the user terminal playing back the audio data received from the server as audio.

[1005] This invention is a system that reproduces the voice of a deceased person and provides responses that are in tune with the user's emotions. This system is realized mainly using a user terminal, a server, an emotion engine, and a generative AI model.

[1006] System configuration

[1007] User terminal

[1008] A user terminal is an electronic device operated by a user, including a smartphone, tablet, smart glasses, etc. This terminal is used by the user to upload voice data and perform subsequent interactions.

[1009] server

[1010] The server is a central system that provides data and services over the network. It is responsible for various processes such as analyzing, storing, and generating voice data. This server receives voice data sent from user devices, extracts feature data, and trains generative AI models. It is also responsible for the process of converting text messages entered by users into voice data.

[1011] Emotion Engine

[1012] The emotion engine is a system that recognizes the user's emotions and generates responses based on those emotions. The emotion engine acquires and analyzes the user's facial expressions and voice tone data via the camera and microphone on the user's device.

[1013] Generative AI Models

[1014] A generative AI model is an artificial intelligence model trained using extracted feature data. This model generates voice data that reproduces the characteristics of the deceased based on the input text message. Examples of such models include the Wav2Vec2 model and EmotionRecognizer.

[1015] Data processing and calculation

[1016] The system performs the following data processing and calculations:

[1017] 1. Analysis of speech data: The speech data of the deceased person uploaded from the user's device to the server is analyzed and feature data is extracted. This analysis is performed using models such as Wav2Vec2Processor and Wav2Vec2ForCTC.

[1018] 2. Emotion Recognition: The emotion engine analyzes voice tone and facial expression data to recognize the user's emotions. This analysis is performed using emotion recognition models such as EmotionRecognizer.

[1019] 3. Voice data generation: The server uses the generative AI model to generate voice data that reproduces the voice of the deceased person based on the text message entered by the user, matching the emotions analyzed by the emotion engine.

[1020] Specific examples

[1021] A specific usage scenario is when a user wears smart glasses at a memorial cafe and uploads their mother's voice data. The user enters a prompt such as, "Mom, tell me about your recent trip." In this case, the emotion engine analyzes the user's emotions, and the server generates an optimal response in the mother's voice. The user's device then receives the generated voice data and plays it back to the user, allowing the user to hear their mother's voice again. This allows for deep healing and empathy that is tailored to their emotions.

[1022] Prompt Sentence Examples

[1023] "Mom, tell me about your recent trip."

[1024] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1025] Step 1:

[1026] The user uploads the audio data of the deceased.

[1027] Description: The user launches the dedicated application and clicks the "Upload audio data" button. The user selects the audio file (e.g., MP3, WAV) of the deceased person and begins uploading.

[1028] Input: Audio file (e.g. MP3, WAV)

[1029] Output: Uploaded audio data

[1030] Step 2:

[1031] The terminal transmits the voice data to the server.

[1032] Description: The terminal captures the selected audio file, checks the format and quality of the audio file, and then sends the audio file to the server if it is suitable.

[1033] Input: Uploaded audio data

[1034] Output: Audio data sent to the server

[1035] Step 3:

[1036] The server analyzes the voice data and extracts feature data.

[1037] Description: The server analyzes the received audio data and extracts features such as pitch, speed, and intonation. The analysis is performed using Wav2Vec2Processor and Wav2Vec2ForCTC models.

[1038] Input: Audio data sent to the server

[1039] Output: feature data

[1040] Step 4:

[1041] The server uses the feature data to train a generative AI model.

[1042] Description: The server uses the extracted feature data to train a generative AI model, which is then optimized to reproduce the characteristics of the deceased person's voice.

[1043] Input: feature data

[1044] Output: A trained generative AI model

[1045] Step 5:

[1046] The emotion engine recognizes the user's emotions.

[1047] Description: When a user uses an application, the emotion engine captures and analyzes the user's facial expression data and voice tone data through the camera and microphone on the user's device. The analysis uses models such as EmotionRecognizer.

[1048] Input: User's facial expression data, voice tone data

[1049] Output: User emotion data

[1050] Step 6:

[1051] The user inputs a text message, and the emotion engine sends the emotion analysis results to the server.

[1052] Description: A user inputs any text message through the application. The emotion engine analyzes the user's text message and sends it to the server along with the emotion results.

[1053] Input: Text messages, emotion data

[1054] Output: Text message and emotion data sent to the server

[1055] Step 7:

[1056] The server generates voice data using the text message and the emotion data.

[1057] Description: The server inputs the received text message and emotional data into the generative AI model and generates corresponding voice data.

[1058] Input: Text messages, emotion data, pre-trained generative AI model

[1059] Output: Generated audio data

[1060] Step 8:

[1061] The terminal receives and plays the generated audio data.

[1062] Description: The device analyzes the audio data received from the server, checks its integrity, and then plays it back.

[1063] Input: Generated audio data

[1064] Output: Played audio

[1065] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1066] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1067] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[1068] [Third embodiment]

[1069] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[1070] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.

[1071] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1072] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[1073] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1074] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1075] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1076] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1077] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1078] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1079] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1080] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[1081] This invention uses the voice data of the deceased to allow their family and friends to hear their voice again through AI. The program and specific processing flow of this system are shown below.

[1082] Program Overview

[1083] The program works through the following major steps:

[1084] 1. The user uploads the audio data of the deceased person.

[1085] 2. The device sends the audio data to the server.

[1086] 3. The server analyzes the voice data and extracts feature data.

[1087] 4. The server uses the feature data to train a generative AI model.

[1088] 5. The server receives the user's text message and generates audio data.

[1089] 6. The device receives and plays the generated audio data.

[1090] Processing flow and specific examples

[1091] User voice data input

[1092] 1. The user launches the dedicated application and clicks the "Upload audio data" button.

[1093] 2. The user selects the audio file (e.g. MP3, WAV) of the deceased person and begins uploading.

[1094] Terminal handling

[1095] 3. The device will import the selected audio file and check the format and quality of the audio file.

[1096] 4. If the format and quality match, the device sends the audio file to the server.

[1097] Validation and storage of voice data by the server

[1098] 1. The server receives the audio file sent from the device.

[1099] 2. The server analyzes the contents of the audio file to ensure there are no problems with the format or quality.

[1100] 3. If the audio file matches, the server stores the file in its database and assigns it a unique ID.

[1101] 4. The server notifies the device that the audio file has been saved.

[1102] Training generative AI models

[1103] 1. After the audio file is saved, the server performs a characteristic analysis of the deceased person's audio data.

[1104] 2. The server uses a feature extraction algorithm to analyze and extract features such as pitch, speed, and intonation of the speech.

[1105] 3. Train a generative AI model (e.g., a large-scale language model or a speech synthesis model) based on the extracted feature data.

[1106] 4. Once trained, the generative AI model is optimized to reproduce the vocal characteristics of the deceased person.

[1107] Enter and send text messages

[1108] 1. The user enters an arbitrary text message through the application (e.g., "How are you feeling today?").

[1109] 2. The user clicks the "Send" button to send the text message to the server.

[1110] Server-generated audio

[1111] 1. The server parses the received text message.

[1112] 2. The server inputs the analyzed text into a generative AI model and generates corresponding audio data.

[1113] 3. The generative AI model uses the trained data to convert the input text into audio data in the characteristic voice of the deceased.

[1114] 4. The server sends the generated voice data to the terminal.

[1115] Receiving and playing audio on the device

[1116] 1. The device analyzes the voice data received from the server.

[1117] 2. The device checks the integrity of the audio data and prepares for playback.

[1118] 3. The device plays the audio data within the application, allowing the user to hear the voice of the deceased.

[1119] Specific examples

[1120] Consider a case where a user uses a smartphone app to upload audio data of their deceased grandmother. The audio data is uploaded to the app, and the smartphone sends the data to a server. The server analyzes the audio data, extracts the characteristics of the grandmother's voice, and trains an AI model. Next, when the user types a text message into the app saying, "Tell me about your favorite flower," the server receives this message and uses the trained AI model to generate audio data in the grandmother's voice. The generated audio data is sent to the smartphone and played back by the app. The user can hear their grandmother's voice again, which can be moving and comforting.

[1121] The above is an example of the implementation of the present invention. This system allows the bereaved family to re-create a heart-to-heart conversation with the deceased, and to gain healing and empathy.

[1122] The processing flow will be explained below.

[1123] Step 1:

[1124] User

[1125] The user launches the dedicated application and clicks the "Upload audio data" button.

[1126] Step 2:

[1127] User

[1128] The user selects the deceased person's audio file (e.g. MP3, WAV) and begins uploading.

[1129] Step 3:

[1130] Terminal

[1131] The device captures the selected audio file and checks the format and quality of the audio file.

[1132] Step 4:

[1133] Terminal

[1134] If the format and quality match, the terminal sends the audio file to the server.

[1135] Step 5:

[1136] server

[1137] The server receives the audio file sent from the terminal.

[1138] Step 6:

[1139] server

[1140] The server analyzes the content of the audio file to ensure that there are no problems with the format or quality.

[1141] Step 7:

[1142] server

[1143] If the audio file matches, the server stores the file in a database and assigns it a unique ID.

[1144] Step 8:

[1145] server

[1146] The server notifies the terminal that the audio file has been saved.

[1147] Step 9:

[1148] server

[1149] The server uses the stored audio files to characterize the deceased person's audio data.

[1150] Step 10:

[1151] server

[1152] The server uses feature extraction algorithms to analyze and extract features such as pitch, speed, and intonation of the speech.

[1153] Step 11:

[1154] server

[1155] Based on the extracted feature data, generative AI models (e.g., large-scale language models or speech synthesis models) are trained.

[1156] Step 12:

[1157] server

[1158] Once trained, the generative AI model is optimized to reproduce the vocal characteristics of the deceased person.

[1159] Step 13:

[1160] User

[1161] Through the application, the user enters an arbitrary text message (e.g., "How are you feeling today?").

[1162] Step 14:

[1163] User

[1164] The user clicks the "Send" button to send the text message to the server.

[1165] Step 15:

[1166] Terminal

[1167] The terminal transmits the input text message to the server and makes a request to the server.

[1168] Step 16:

[1169] server

[1170] The server parses the received text message.

[1171] Step 17:

[1172] server

[1173] The server inputs the analyzed text into a generative AI model and generates corresponding audio data.

[1174] Step 18:

[1175] Generative AI Models

[1176] The generative AI model uses trained data to convert input text into audio data in the characteristic voice of the deceased.

[1177] Step 19:

[1178] server

[1179] The server transmits the generated voice data to the terminal.

[1180] Step 20:

[1181] Terminal

[1182] The terminal analyzes the voice data received from the server.

[1183] Step 21:

[1184] Terminal

[1185] The device checks the integrity of the audio data and prepares it for playback.

[1186] Step 22:

[1187] Terminal

[1188] The device plays the audio data within the application, allowing the user to hear the voice of the deceased.

[1189] Example 1

[1190] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1191] Technology that allows bereaved families and friends to hear the voices of the deceased again can bring emotion and healing, but conventional systems require a lot of work, such as uploading the voice data, checking the quality, and training the AI ​​model, making it difficult to operate efficiently.In addition, because the processes of checking the voice data quality, extracting features, and training the AI ​​model are complex, making it easy for users to use is a challenge.

[1192] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1193] In this invention, the server includes a means for a user to upload voice data of the deceased via a terminal, a means for the terminal to check the format and quality of the voice data and send it to the server, and a means for the server to analyze the received voice data and extract feature data, thereby enabling a user to efficiently and easily use the voice data of the deceased and reproduce the deceased's voice using an AI model.

[1194] A "user" is a person who uses the system to upload and play back audio data of a deceased person.

[1195] A "terminal" is a device operated by a user that uploads audio data, checks the format and quality, transmits it to a server, and plays back the received audio data.

[1196] A "server" is a device that analyzes received voice data, extracts feature data, trains a generative AI model, generates voice data, and transmits it to a terminal.

[1197] "Audio data" refers to a digital data file (e.g., MP3, WAV) that records the voice of a deceased person.

[1198] "Uploading" is the act of a user sending audio data to a server via a terminal.

[1199] "Format" refers to the format of the audio data file (e.g., MP3, WAV).

[1200] "Quality" refers to characteristics related to the sound quality of data, such as the sampling rate and bit depth of the audio data.

[1201] "Analysis" is the process by which the server performs detailed analysis of the received audio data to extract specific information.

[1202] "Feature data" is information that indicates voice characteristics such as pitch, speed, and intonation extracted from voice data.

[1203] A "generative AI model" is an artificial intelligence model trained to convert text messages into audio data based on feature data.

[1204] "Training" is the process of learning from feature data so that the generative AI model can accurately reproduce the voice of the deceased.

[1205] A "text message" is text information that is entered by a user and sent to a server.

[1206] "Conversion" is the process of using a generative AI model to turn a text message into audio data.

[1207] "Reception" refers to the act of the server receiving data transmitted from the user terminal, or the act of the user terminal receiving data transmitted from the server.

[1208] "Playback" refers to the act of outputting audio data received by a user terminal as audio so that the user can hear it.

[1209] This invention is a system that uses the voice data of the deceased to allow bereaved family members and friends to hear the voice of the deceased again through AI. This system consists of three main elements: a user, a terminal, and a server. The specific configuration and embodiments of this system are described below.

[1210] First, the user launches a dedicated application on their device. This device can be a smartphone, tablet, or computer, which is capable of importing and playing audio data. When the user clicks the "Upload Audio Data" button, a screen appears allowing them to select the audio file (e.g., MP3, WAV) of the deceased person. The user selects the target audio file and begins uploading.

[1211] The device then retrieves the selected audio file and checks its format and quality. This check checks whether the audio file is in MP3 or WAV format and meets certain audio quality standards (e.g., 44.1kHz, 16-bit). If the quality is acceptable, the device sends the audio file to the server. During this process, the user sees an "Uploading" progress bar.

[1212] When the server receives the audio file sent from the device, it checks the format and quality of the audio data again. If there are no problems, the server saves the audio file in its database and assigns a unique identifier (ID) to the file. The server notifies the device that the file has been saved, allowing the user to confirm that the upload was successful.

[1213] The server extracts feature data from the stored audio files. Specifically, it uses a feature extraction algorithm to analyze and extract features such as pitch, speed, and intonation. This feature data is used to train a generative AI model (e.g., a large-scale language model or a speech synthesis model). This model is optimized to reproduce the characteristics of the deceased person's voice.

[1214] If a user wants to use the voice of a deceased person, they can enter any text message through the application, for example, "How are you feeling today?" or "Tell me about a flower you liked," and click the "Send" button. This message will be sent to the server.

[1215] The server analyzes the received text message and inputs it into a generative AI model to generate corresponding audio data. This AI model converts the input text into audio data in the deceased's characteristic voice based on the trained audio feature data. The generated audio data is then sent to the device.

[1216] The audio data received by the device is checked for integrity and then prepared for playback. Finally, the audio data is played back within the application, allowing the user to hear the voice of the deceased. This recreates an emotional conversation with the deceased, bringing emotion and healing to bereaved family and friends.

[1217] Specifically, when a user uses a smartphone app to upload audio data of their deceased grandmother, the audio data is uploaded to the app, and the smartphone sends the data to a server. The server analyzes the audio data, extracts the characteristics of the grandmother's voice, and trains an AI model. When the user types a text message into the app saying, "Tell me about your favorite flower," the server receives the message and uses the trained AI model to generate audio data in the grandmother's voice. The generated audio data is sent to the smartphone and played back in the app. The user can hear their grandmother's voice again, and feel moved and comforted.

[1218] An example of a prompt is as follows:

[1219] Prompt: "Tell me about a flower you loved."

[1220] Generated text: "Hello. I love talking about flowers you liked. My favorite was roses..."

[1221] This invention allows users to relive the heartfelt conversation by listening to the voice of the deceased again, and thus find healing.

[1222] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1223] Step 1:

[1224] The user launches the dedicated application on the device and clicks the "Upload audio data" button. When the user selects an audio file to upload, this file is imported into the device. The input is the audio file selected by the user, and the output is the imported audio file.

[1225] Step 2:

[1226] The device checks the format and quality of the selected audio file. Specifically, it checks that the file is in MP3 or WAV format and that the audio quality is 44.1 kHz and 16-bit. The input is the imported audio file, and the output is the result of matching the format and quality.

[1227] Step 3:

[1228] The device checks the format and quality and, if it is found to be compatible, sends the audio file to the server. A progress bar indicating "uploading" is displayed to the user during the upload. The input is the format and quality check result and the audio file, and the output is the audio file sent to the server.

[1229] Step 4:

[1230] The server receives the audio file sent from the terminal. The input is the audio file sent from the terminal, and the output is the received audio file.

[1231] Step 5:

[1232] The server rechecks the content of the received audio file and, if there are no problems, stores it in a database. Specifically, it analyzes the format and quality of the audio data again and assigns a unique identifier. The input is the received audio file, and the output is the audio file stored in the database and its unique identifier.

[1233] Step 6:

[1234] The server performs feature analysis on the audio files stored in the database. It uses a feature extraction algorithm to analyze and extract features such as pitch, speed, and intonation of the voice. Based on this analysis, feature data is generated. The input is the audio files stored in the database, and the output is the extracted feature data.

[1235] Step 7:

[1236] The server uses the extracted feature data to train a generative AI model. Specifically, the feature data is input into the AI ​​model and the AI ​​model is optimized through an appropriate learning process. The input is the feature data, and the output is a trained generative AI model.

[1237] Step 8:

[1238] A user enters a text message through an application. The user clicks a "Send" button to send the entered text message to a server. The input is the text message entered by the user, and the output is the text message sent to the server.

[1239] Step 9:

[1240] The server analyzes the received text message. The analyzed text message is input into the generative AI model to generate corresponding audio data. The input is the text message sent to the server, and the output is the generated audio data.

[1241] Step 10:

[1242] The server sends the generated voice data to the terminal. The input is the generated voice data, and the output is the voice data sent to the terminal.

[1243] Step 11:

[1244] The device analyzes the audio data received from the server. It checks the integrity of the audio data and prepares for playback. The input is the audio data received from the server, and the output is the audio data whose integrity has been checked.

[1245] Step 12:

[1246] The device plays back the audio data within the application, allowing the user to hear the voice of the deceased. The input is the audio data whose integrity has been verified, and the output is the played-back audio.

[1247] (Application example 1)

[1248] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1249] Technology that allows users to feel moved and comforted by re-listening to the voices of the deceased is extremely valuable. However, conventional systems have difficulty using the voice data of the deceased to read text aloud or generate audiobooks. This makes it difficult for users to experience the added emotional impact of listening to books or poems in the voice of the deceased.

[1250] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1251] In this invention, the server includes: [means for uploading the deceased's voice data from the user terminal to the server;] [means for analyzing the voice data received by the server and extracting feature data;] [means for training a generative AI model using the extracted feature data;] [means for the server to receive text selected by the user and convert the text into voice data using the trained generative AI model;] [means for the user terminal to receive voice data sent from the server and play it back to the user; and [means for sending the text selected by the user to the server and generating an audiobook in the deceased's voice. This allows the user to listen to books and poems in the deceased's voice.

[1252] "Voice data of the deceased" refers to information in the form of data that contains recorded or collected voices made by the deceased during their lifetime.

[1253] A "user terminal" is an information processing device such as a computer, smartphone, or tablet that is used by a user.

[1254] A "server" is a computer system that provides services over a network and sends and receives voice data from user terminals and analyzes it.

[1255] "Feature data" is data that indicates sound characteristics such as pitch, speed, and intonation extracted from voice data.

[1256] A "generative AI model" is an artificial intelligence model that learns the characteristics of voice data and is used to convert specified text into synthetic speech.

[1257] A "text message" is character string information entered by a user, which is received by the server and converted into voice data.

[1258] A "unique ID" is a unique identifier assigned to identify a particular piece of audio data within a database.

[1259] An "audiobook" is an audio version of the contents of a book or poem, and is audio data generated using the voice of a deceased person.

[1260] MODE FOR CARRYING OUT THE INVENTION

[1261] This invention is a system that allows users to feel touched or comforted by using the voice data of the deceased. The specific configuration and processing flow are described below.

[1262] System program and processing description

[1263] This system is constructed from a user device, a server, and a generative AI model. Each component plays the following role:

[1264] Hardware and software used:

[1265] User device: A smartphone, tablet, or computer is used to run applications, upload audio data, and play audiobooks.

[1266] Server: Located on the cloud, it is responsible for analyzing voice data, extracting features, and training and generating generative AI models.

[1267] Generative AI models: Convert text to audio using large-scale language models (e.g., GPT-3) and speech synthesis models (e.g., TTS models).

[1268] Cloud storage: Amazon Web Services (AWS) S3 and other cloud storage services are used to store audio data and generated audiobooks.

[1269] Data processing and data calculations:

[1270] 1. Uploading and sending audio data

[1271] The user uploads the deceased's audio data to the server from their device. First, the user selects an audio file (e.g., MP3, WAV) via a dedicated app and sends the data to the server. The server checks the format and quality of the received audio data and stores the appropriate audio data in cloud storage.

[1272] 2. Extracting feature data and training the AI ​​model

[1273] The server analyzes the stored audio data and extracts feature data. The feature extraction algorithm analyzes characteristics such as pitch, speed, and intonation of the voice, and uses this data to train a generative AI model. The trained model is optimized to reproduce the characteristics of the deceased person's voice.

[1274] 3. Entering text messages and generating voice data

[1275] The user selects the text of a book or poem they want to read through the application and sends it to the server, which analyzes the text and uses a generative AI model to convert the text into audio data in the deceased person's voice. The converted audiobook data is then stored in cloud storage.

[1276] 4. Receiving and playing audio data

[1277] The user device receives the generated audiobook from the server, verifies its integrity, and then plays it within the application, allowing the user to listen to books and poems in the voice of the deceased.

[1278] Examples:

[1279] For example, if a user uses a smartphone app to upload audio data of a deceased person, the audio data is uploaded to the app, and the smartphone sends the data to a server. The server analyzes the audio data, extracts the voice characteristics of the deceased person (e.g., grandmother), and trains an AI model. Next, when the user selects a poetry collection and enters it into the app, the server receives the poetry collection and uses the trained AI model to generate audio data in the deceased person's (grandmother's) voice. The generated audio data is sent to the smartphone and played back in the app. The user can relisten to the poetry collection in the deceased's voice, which can be moving and comforting.

[1280] Example prompt sentence:

[1281] "Generate an audiobook of this text read in the voice of the deceased: Complete Poems"

[1282] The above is a specific embodiment for carrying out the present invention. This system allows the user to re-create a heart-to-heart conversation with the deceased, and can bring about deep emotions and healing.

[1283] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1284] Step 1:

[1285] The user launches the dedicated application and clicks the "Upload audio data" button. The user selects the deceased's audio file (e.g., MP3, WAV) and starts uploading. The input is the audio file, and the output is the device sending the audio file to the server.

[1286] Step 2:

[1287] The terminal takes the selected audio file, checks the format and quality of the audio file, and if it matches, sends the audio file to the server. The input is the audio file, and the output is the audio data whose format and quality have been checked.

[1288] Step 3:

[1289] The server receives the audio file sent from the device. It analyzes the content of the audio file and checks that there are no problems with the format or quality. It stores the matching audio file in cloud storage and assigns a unique ID. The input is the audio file, and the output is the stored audio data and the unique ID.

[1290] Step 4:

[1291] The server analyzes the stored voice data and extracts feature data. Feature extraction algorithms are used to analyze the voice's pitch, speed, intonation, etc., and the data is used to train the generative AI model. The input is the stored voice data, and the output is the feature data.

[1292] Step 5:

[1293] The server uses the extracted feature data to train a generative AI model, which is optimized to reproduce the characteristics of the deceased person's voice. The input is the feature data, and the output is the trained generative AI model.

[1294] Step 6:

[1295] The user selects the text of a book or poem they want to read through the application and sends it to the server. The input is a text message, and the output is the sent text message.

[1296] Step 7:

[1297] The server analyzes the received text message and converts the text content into audio data using a trained generative AI model. The input is the text message and the output is the generated audio data.

[1298] Step 8:

[1299] The server stores the generated voice data in cloud storage and transmits it to the user's device. The input is the generated voice data, and the output is the transmitted voice data.

[1300] Step 9:

[1301] The device analyzes the audio data received from the server and verifies its integrity. The audio data is then played back within the application, allowing the user to hear the voice of the deceased. The input is the audio data received from the server, and the output is the played audio data.

[1302] The above are the specific processing steps of the system based on the application example, through which users can create audiobooks using the voices of deceased people and listen to them again.

[1303] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1304] This invention is a system that combines the voice data of the deceased with an emotion engine that recognizes the user's emotions, allowing bereaved family and friends to hear the deceased's voice again through AI and providing adapted responses based on the user's emotions. The program and specific processing flow of this system are shown below.

[1305] Program Overview

[1306] The program works through the following major steps:

[1307] 1. The user uploads the audio data of the deceased person.

[1308] 2. The device sends the audio data to the server.

[1309] 3. The server analyzes the voice data and extracts feature data.

[1310] 4. The server uses the feature data to train a generative AI model.

[1311] 5. The emotion engine recognizes the user's emotions.

[1312] 6. The user enters a text message, and the emotion engine sends the emotion analysis results to the server.

[1313] 7. The server uses the sentiment analysis results to generate audio data based on the text message.

[1314] 8. The device receives and plays the generated audio data.

[1315] Processing flow and specific examples

[1316] User voice data input

[1317] 1. The user launches the dedicated application and clicks the "Upload audio data" button.

[1318] 2. The user selects the audio file (e.g. MP3, WAV) of the deceased person and begins uploading.

[1319] Terminal handling

[1320] 3. The device will import the selected audio file and check the format and quality of the audio file.

[1321] Sending terminal

[1322] 4. If the format and quality match, the device sends the audio file to the server.

[1323] Validation and storage of voice data by the server

[1324] 1. The server receives the audio file sent from the device.

[1325] 2. The server analyzes the contents of the audio file to ensure there are no problems with the format or quality.

[1326] 3. If the audio file matches, the server stores the file in its database and assigns it a unique ID.

[1327] 4. The server notifies the device that the audio file has been saved.

[1328] Training generative AI models

[1329] 1. The server uses the stored audio files to characterize the deceased person's audio data.

[1330] 2. The server uses a feature extraction algorithm to analyze and extract features such as pitch, speed, and intonation of the speech.

[1331] 3. Train a generative AI model (e.g., a large-scale language model or a speech synthesis model) based on the extracted feature data.

[1332] 4. Once trained, the generative AI model is optimized to reproduce the vocal characteristics of the deceased person.

[1333] Emotion Engine Operation

[1334] 1. When a user uses an application, the emotion engine acquires the user's facial expression data and voice tone data through the camera and microphone.

[1335] 2. The emotion engine analyzes the acquired data and determines the user's emotional state (e.g., joy, sadness, surprise).

[1336] Enter and send text messages

[1337] 1. The user enters a text message through the application (e.g., "How have things been lately?").

[1338] 2. The emotion engine sends the user's text message along with the analyzed emotion results to the server.

[1339] Server-generated audio

[1340] 1. The server analyzes the received text message and the sentiment analysis results.

[1341] 2. The server inputs the analyzed text and emotion data into the generative AI model and generates corresponding voice data.

[1342] 3. Using the trained data, the generative AI model converts input text into audio data in the characteristic voice of the deceased based on the user's emotional state.

[1343] 4. The server sends the generated voice data to the terminal.

[1344] Receiving and playing audio on the device

[1345] 1. The device analyzes the voice data received from the server.

[1346] 2. The device checks the integrity of the audio data and prepares for playback.

[1347] 3. The device plays the audio data within the application, allowing the user to hear the voice of the deceased.

[1348] Specific examples

[1349] Consider a scenario where a user uses a smartphone app to upload audio data of their deceased grandmother. The audio data is uploaded to the app, and the smartphone sends the data to a server. The server analyzes the audio data, extracts the characteristics of the grandmother's voice, and trains an AI model. Next, when the user types a text message into the app saying, "Tell me about your favorite flower," the emotion engine analyzes the user's emotions and sends the emotion data to the server. Based on this information, the server uses the trained AI model to generate audio data in the grandmother's voice. The generated audio data is sent to the smartphone and played back by the app. The user can hear their grandmother's voice again, and experience deeper emotion and healing from the emotionally sensitive response.

[1350] The above is an example of the implementation of the present invention. This system recreates the heart-to-heart dialogue with the deceased and provides responses tailored to the user's emotions, allowing the bereaved to receive even deeper healing and empathy.

[1351] The processing flow will be explained below.

[1352] Step 1:

[1353] User

[1354] The user launches the dedicated application and clicks the "Upload audio data" button.

[1355] Step 2:

[1356] User

[1357] The user selects the deceased person's audio file (e.g. MP3, WAV) and begins uploading.

[1358] Step 3:

[1359] Terminal

[1360] The device captures the selected audio file and checks the format and quality of the audio file.

[1361] Step 4:

[1362] Terminal

[1363] If the format and quality match, the terminal sends the audio file to the server.

[1364] Step 5:

[1365] server

[1366] The server receives the audio file sent from the terminal.

[1367] Step 6:

[1368] server

[1369] The server analyzes the content of the audio file to ensure that there are no problems with the format or quality.

[1370] Step 7:

[1371] server

[1372] If the audio file matches, the server stores the file in a database and assigns it a unique ID.

[1373] Step 8:

[1374] server

[1375] The server notifies the terminal that the audio file has been saved.

[1376] Step 9:

[1377] server

[1378] The server uses the stored audio files to characterize the deceased person's audio data.

[1379] Step 10:

[1380] server

[1381] The server uses feature extraction algorithms to analyze and extract features such as pitch, speed, and intonation of the speech.

[1382] Step 11:

[1383] server

[1384] Based on the extracted feature data, generative AI models (e.g., large-scale language models or speech synthesis models) are trained.

[1385] Step 12:

[1386] server

[1387] Once trained, the generative AI model is optimized to reproduce the vocal characteristics of the deceased person.

[1388] Step 13:

[1389] Emotion Engine

[1390] When a user uses an application, the emotion engine acquires the user's facial expression data and voice tone data through a camera and microphone.

[1391] Step 14:

[1392] Emotion Engine

[1393] The emotion engine analyzes the acquired data and determines the user's emotional state (e.g., joy, sadness, surprise).

[1394] Step 15:

[1395] User

[1396] The user enters an arbitrary text message (e.g., "How have things been lately?") through the application.

[1397] Step 16:

[1398] User

[1399] The user clicks the "Send" button to send the text message to the server.

[1400] Step 17:

[1401] Emotion Engine

[1402] The emotion engine sends the user's text message along with the analyzed emotion results to the server.

[1403] Step 18:

[1404] server

[1405] The server analyzes the received text message and the sentiment analysis results.

[1406] Step 19:

[1407] server

[1408] The server inputs the analyzed text and emotion data into a generative AI model to generate corresponding voice data.

[1409] Step 20:

[1410] Generative AI Models

[1411] The generative AI model uses trained data to convert input text and emotional data into audio data in the characteristic voice of the deceased.

[1412] Step 21:

[1413] server

[1414] The server transmits the generated voice data to the terminal.

[1415] Step 22:

[1416] Terminal

[1417] The terminal analyzes the voice data received from the server.

[1418] Step 23:

[1419] Terminal

[1420] The device checks the integrity of the audio data and prepares it for playback.

[1421] Step 24:

[1422] Terminal

[1423] The device plays the audio data within the application, allowing the user to hear the voice of the deceased.

[1424] As a concrete example, consider the case where a user uses a smartphone app to upload voice data of their deceased grandmother and engages in emotionally appropriate conversations through an emotion engine. The voice file uploaded by the user is received by a server, which analyzes and extracts features. A generative AI model is trained based on this data. The user then performs emotion analysis through the emotion engine and enters and sends a text message. The data including the emotion analysis results is passed to the server, and emotion-appropriate voice data is generated in the grandmother's voice through the generative AI model. This voice data is received and played back on the device. The user can hear the grandmother's voice emotionally, and experience deep emotion and healing.

[1425] Example 2

[1426] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1427] For bereaved family members and friends who wish to reconnect with their loved ones, recreating their voice and receiving emotionally responsive responses is extremely valuable. However, simply playing back audio data does not provide a deep conversational experience with the deceased. Furthermore, current technology lacks a system that can recognize the user's emotions and respond accordingly.

[1428] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1429] In this invention, the server includes: [means for uploading the voice data of the deceased from the user terminal to the server; [means for analyzing the voice data received by the server and extracting feature data; [means for training a generative AI model using the extracted feature data; [means for an emotion engine to recognize the user's emotions and acquire facial expression data and voice tone data; [means for the server to receive a text message entered by the user and convert the text message into voice data using the emotion analysis results; and [means for the user terminal to receive the voice data sent from the server and play it back to the user. This enables the user to have a conversation with the deceased that is sensitive to their emotions.

[1430] "Voice data of a deceased person" refers to a digital file uploaded by a user that records the voice of a deceased person during their lifetime.

[1431] "User terminal" refers to electronic devices used by users, such as smartphones, tablets, and personal computers.

[1432] "Server" refers to a high-performance computer system for analyzing, storing, and generating voice data.

[1433] "Analysis" refers to the process of examining and examining the content of received voice data in detail to extract necessary feature data.

[1434] "Feature data" refers to information that quantifies the unique characteristics of voice data, such as voice pitch, speed, and intonation.

[1435] "Generative AI model" refers to an artificial intelligence model that generates voice based on feature data extracted from the voice data of a deceased person.

[1436] "Training" refers to the process of learning using data to enable a generative AI model to reproduce the vocal characteristics of the deceased person.

[1437] An "emotion engine" refers to an algorithm and system for recognizing and analyzing user emotions in real time.

[1438] "Facial expression data" refers to data that quantifies and classifies changes in a user's facial expressions captured through a camera.

[1439] "Voice tone data" refers to data obtained by quantifying and analyzing changes in the tone of a user's speech obtained through a microphone.

[1440] "Emotion analysis results" refers to data in which the emotion engine analyzes the user's emotional state and expresses the results in numerical values ​​and categories.

[1441] "Text message" refers to text information entered by a user through an application.

[1442] "Audio Data" refers to the digital audio files generated by the generative AI model that reproduce the voice of the deceased person.

[1443] This invention is a system that allows bereaved family and friends to hear the voice of the deceased again through AI by combining the voice data of the deceased with an emotion engine that recognizes the user's emotions. The system provides adaptive responses based on the emotions. Below, we will explain in detail how this system is implemented.

[1444] System configuration and processing overview

[1445] 1. User voice data input

[1446] The user launches the dedicated application and clicks the "Upload audio data" button.

[1447] The user selects the audio file of the deceased person (e.g., MP3, WAV format) and begins uploading.

[1448] 2. Terminal Processing

[1449] The device will then import the selected audio file and check the format and quality of the audio file using an audio analysis library (e.g., FFmpeg).

[1450] 3. Sending the device

[1451] If the format and quality match, the device sends the audio file to the server via an HTTP POST request.

[1452] 4. Verification and storage of audio data by the server

[1453] The server receives the audio file sent from the device, analyzes it again, and checks that there are no problems with the format or quality. If the audio file is suitable, the server stores it in a database (e.g., MongoDB) and assigns it a unique ID.

[1454] The server sends a storage completion notification to the terminal.

[1455] 5. Training the generative AI model

[1456] The server uses the stored audio files to perform a feature analysis of the deceased person's voice data. The feature analysis of the voice data is performed using a voice analysis tool (e.g., Praat, Kaldi).

[1457] The server uses feature extraction algorithms to extract features such as pitch, speed, and intonation, and then uses them to train a generative AI model (e.g., GPT-3, Tacotron 2).

[1458] 6. Operation of the Emotion Engine

[1459] The emotion engine captures facial expression data and voice tone data via a camera or microphone when a user uses an application, and analyzes the user's emotional state in real time using facial expression recognition algorithms (e.g., OpenCV, Dlib) and voice tone analysis tools (e.g., Praat).

[1460] 7. Enter and send text messages

[1461] The user types a text message through the application (e.g., "How's it been lately?").

[1462] The emotion engine sends a text message along with the analyzed emotion results to the server.

[1463] 8. Server-generated audio

[1464] The server analyzes the received text messages and emotion analysis results, inputs them into a generative AI model, and generates audio data in the voice of the deceased.

[1465] The server transmits the generated voice data to the terminal.

[1466] 9. Receiving and Playing Audio by the Device

[1467] The device analyzes the audio data received from the server, checks its integrity, and then plays the audio data within the application.

[1468] Specific use cases

[1469] For example, consider a case where a user uploads audio data of a deceased person using a smartphone app. The user uploads the audio data to the app, and the device sends the data to a server. The server analyzes the audio data, extracts the characteristics of the deceased's voice, and trains a generative AI model. Next, when the user types "Tell me about your favorite flower" into the app, the emotion engine analyzes the user's emotions and generates the most appropriate response based on those emotions. The generated audio data is sent to the smartphone, and the user can listen to the deceased's voice again in the app and enjoy a dialogue tailored to their emotions.

[1470] Prompt Sentence Examples

[1471] "Tell me about a flower you liked."

[1472] "How was your day today?"

[1473] "Tell me about a memory of going out together."

[1474] The above is a concrete embodiment for carrying out the present invention. Through this system, the user can hear the voice of the deceased again and receive emotional responses.

[1475] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1476] Step 1:

[1477] The user launches the dedicated application and clicks the "Upload audio data" button. The user selects the audio file (e.g., MP3, WAV) of the deceased person and begins uploading.

[1478] Input: User selection of an audio file.

[1479] Output: The selected audio file.

[1480] Step 2:

[1481] The device will then import the selected audio file and check the format and quality of the audio file, using an audio analysis library (e.g. FFmpeg) to analyze the bit rate, sampling rate, etc.

[1482] Input: The selected audio file.

[1483] Output: Data about the format and quality of the audio file.

[1484] Step 3:

[1485] If the audio file format and quality are compatible, the device sends the audio file to the server via an HTTP POST request.

[1486] Input: Quality checked audio files.

[1487] Output: The audio file sent to the server.

[1488] Step 4:

[1489] The server receives the audio file sent from the device and analyzes its content. It uses an audio analysis tool (e.g., Praat, Kaldi) to check for format and quality issues. If there are no problems, it stores the audio file in a database (e.g., MongoDB) and assigns a unique ID.

[1490] Input: The audio file sent from the device.

[1491] Output: The saved audio file and its unique ID.

[1492] Step 5:

[1493] The server uses the stored audio files to extract feature data. It uses speech analysis algorithms to analyze and extract features such as pitch, speed, and intonation. It then trains a generative AI model (e.g., GPT-3, Tacotron 2) based on the extracted feature data.

[1494] Input: Saved audio file.

[1495] Output: feature data and a trained generative AI model.

[1496] Step 6:

[1497] The emotion engine captures the user's facial expression data and voice tone data via the camera and microphone when the user uses the application, and performs real-time emotion analysis using facial expression recognition algorithms (e.g., OpenCV, Dlib) and voice tone analysis tools.

[1498] Input: User's facial expression data and voice tone data.

[1499] Output: User sentiment analysis results.

[1500] Step 7:

[1501] The user inputs a text message through the application (e.g., "How have things been lately?"). The emotion engine analyzes the text message and sends it to the server along with the resulting emotion.

[1502] Input: User text message and sentiment analysis results.

[1503] Output: The text message sent to the server and the sentiment result.

[1504] Step 8:

[1505] The server analyzes the received text message and the emotion analysis results, inputs them into a generative AI model, and generates audio data in the deceased's voice. The generated audio data is then sent from the server to the device.

[1506] Input: Text message, sentiment analysis results, pre-trained generative AI model.

[1507] Output: The generated audio data.

[1508] Step 9:

[1509] The device analyzes the audio data received from the server, checks its integrity, and then plays the audio data within the application.

[1510] Input: Audio data received from the server.

[1511] Output: The audio data that is played within the application.

[1512] (Application example 2)

[1513] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1514] In recent years, there has been an increasing need to cherish memories of deceased loved ones. However, conventional technologies have difficulty reproducing the voices of the deceased, and are unable to provide responses that are particularly sensitive to the user's emotions. Furthermore, more advanced technology is required to enable users to re-experience conversations with their deceased loved ones and achieve emotional healing. Furthermore, there is a need for a system that can provide a deeper experience in both real and virtual spaces by realizing this through specific devices.

[1515] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1516] In this invention, the server includes: [means for uploading the voice data of the deceased from the user terminal to the server; [means for analyzing the voice data received by the server and extracting feature data; [means for training a generative AI model using the extracted feature data; [means for the server to receive a text message entered by the user and convert the text message into voice data using the trained generative AI model; [means for the user terminal to receive the voice data sent from the server and play it back to the user; [means for an emotion engine to recognize the user's emotions and provide a response based on those emotions; and [means for playing back the response through a specific terminal (e.g., smartphone, smart glasses, etc.)]. This allows [the user to hear the voice of the deceased again and receive a response that is in tune with their emotions, thereby achieving deeper healing and empathy].

[1517] A "user terminal" is an electronic device operated by a user, and includes devices such as smartphones, tablets, and smart glasses.

[1518] A "server" is a central system that provides data and services over a network, and performs various processes such as analyzing, storing, and generating voice data.

[1519] "Audio data" refers to data in file format that records the voice of the deceased, and refers to information that includes the characteristics of the voice.

[1520] "Analysis" is the process in which the server analyzes the received voice data and extracts feature data.

[1521] "Feature data" refers to individual characteristic information such as pitch, speed, and intonation extracted from speech data.

[1522] A "generative AI model" is an artificial intelligence model trained based on extracted feature data, and is responsible for converting text messages into voice data.

[1523] A "text message" is a textual message entered by a user and sent to a server.

[1524] "Converting to voice data" refers to the process in which a generative AI model generates voice-format data based on the input text message.

[1525] An "emotion engine" is a system that recognizes a user's emotions and generates a response according to those emotions.

[1526] "Play" refers to the operation of the user terminal playing back the audio data received from the server as audio.

[1527] This invention is a system that reproduces the voice of a deceased person and provides responses that are in tune with the user's emotions. This system is realized mainly using a user terminal, a server, an emotion engine, and a generative AI model.

[1528] System configuration

[1529] User terminal

[1530] A user terminal is an electronic device operated by a user, including a smartphone, tablet, smart glasses, etc. This terminal is used by the user to upload voice data and perform subsequent interactions.

[1531] server

[1532] The server is a central system that provides data and services over the network. It is responsible for various processes such as analyzing, storing, and generating voice data. This server receives voice data sent from user devices, extracts feature data, and trains generative AI models. It is also responsible for the process of converting text messages entered by users into voice data.

[1533] Emotion Engine

[1534] The emotion engine is a system that recognizes the user's emotions and generates responses based on those emotions. The emotion engine acquires and analyzes the user's facial expressions and voice tone data via the camera and microphone on the user's device.

[1535] Generative AI Models

[1536] A generative AI model is an artificial intelligence model trained using extracted feature data. This model generates voice data that reproduces the characteristics of the deceased based on the input text message. Examples of such models include the Wav2Vec2 model and EmotionRecognizer.

[1537] Data processing and calculation

[1538] The system performs the following data processing and calculations:

[1539] 1. Analysis of speech data: The speech data of the deceased person uploaded from the user's device to the server is analyzed and feature data is extracted. This analysis is performed using models such as Wav2Vec2Processor and Wav2Vec2ForCTC.

[1540] 2. Emotion Recognition: The emotion engine analyzes voice tone and facial expression data to recognize the user's emotions. This analysis is performed using emotion recognition models such as EmotionRecognizer.

[1541] 3. Voice data generation: The server uses the generative AI model to generate voice data that reproduces the voice of the deceased person based on the text message entered by the user, matching the emotions analyzed by the emotion engine.

[1542] Specific examples

[1543] A specific usage scenario is when a user wears smart glasses at a memorial cafe and uploads their mother's voice data. The user enters a prompt such as, "Mom, tell me about your recent trip." In this case, the emotion engine analyzes the user's emotions, and the server generates an optimal response in the mother's voice. The user's device then receives the generated voice data and plays it back to the user, allowing the user to hear their mother's voice again. This allows for deep healing and empathy that is tailored to their emotions.

[1544] Prompt Sentence Examples

[1545] "Mom, tell me about your recent trip."

[1546] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1547] Step 1:

[1548] The user uploads the audio data of the deceased.

[1549] Description: The user launches the dedicated application and clicks the "Upload audio data" button. The user selects the audio file (e.g., MP3, WAV) of the deceased person and begins uploading.

[1550] Input: Audio file (e.g. MP3, WAV)

[1551] Output: Uploaded audio data

[1552] Step 2:

[1553] The terminal transmits the voice data to the server.

[1554] Description: The terminal captures the selected audio file, checks the format and quality of the audio file, and then sends the audio file to the server if it is suitable.

[1555] Input: Uploaded audio data

[1556] Output: Audio data sent to the server

[1557] Step 3:

[1558] The server analyzes the voice data and extracts feature data.

[1559] Description: The server analyzes the received audio data and extracts features such as pitch, speed, and intonation. The analysis is performed using Wav2Vec2Processor and Wav2Vec2ForCTC models.

[1560] Input: Audio data sent to the server

[1561] Output: feature data

[1562] Step 4:

[1563] The server uses the feature data to train a generative AI model.

[1564] Description: The server uses the extracted feature data to train a generative AI model, which is then optimized to reproduce the characteristics of the deceased person's voice.

[1565] Input: feature data

[1566] Output: A trained generative AI model

[1567] Step 5:

[1568] The emotion engine recognizes the user's emotions.

[1569] Description: When a user uses an application, the emotion engine captures and analyzes the user's facial expression data and voice tone data through the camera and microphone on the user's device. The analysis uses models such as EmotionRecognizer.

[1570] Input: User's facial expression data, voice tone data

[1571] Output: User emotion data

[1572] Step 6:

[1573] The user inputs a text message, and the emotion engine sends the emotion analysis results to the server.

[1574] Description: A user inputs any text message through the application. The emotion engine analyzes the user's text message and sends it to the server along with the emotion results.

[1575] Input: Text messages, emotion data

[1576] Output: Text message and emotion data sent to the server

[1577] Step 7:

[1578] The server generates voice data using the text message and the emotion data.

[1579] Description: The server inputs the received text message and emotional data into the generative AI model and generates corresponding voice data.

[1580] Input: Text messages, emotion data, pre-trained generative AI model

[1581] Output: Generated audio data

[1582] Step 8:

[1583] The terminal receives and plays the generated audio data.

[1584] Description: The device analyzes the audio data received from the server, checks its integrity, and then plays it back.

[1585] Input: Generated audio data

[1586] Output: Played audio

[1587] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1588] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1589] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1590] [Fourth embodiment]

[1591] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1592] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1593] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1594] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1595] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1596] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1597] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1598] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1599] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1600] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1601] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1602] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1603] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1604] This invention uses the voice data of the deceased to allow their family and friends to hear their voice again through AI. The program and specific processing flow of this system are shown below.

[1605] Program Overview

[1606] The program works through the following major steps:

[1607] 1. The user uploads the audio data of the deceased person.

[1608] 2. The device sends the audio data to the server.

[1609] 3. The server analyzes the voice data and extracts feature data.

[1610] 4. The server uses the feature data to train a generative AI model.

[1611] 5. The server receives the user's text message and generates audio data.

[1612] 6. The device receives and plays the generated audio data.

[1613] Processing flow and specific examples

[1614] User voice data input

[1615] 1. The user launches the dedicated application and clicks the "Upload audio data" button.

[1616] 2. The user selects the audio file (e.g. MP3, WAV) of the deceased person and begins uploading.

[1617] Terminal handling

[1618] 3. The device will import the selected audio file and check the format and quality of the audio file.

[1619] 4. If the format and quality match, the device sends the audio file to the server.

[1620] Validation and storage of voice data by the server

[1621] 1. The server receives the audio file sent from the device.

[1622] 2. The server analyzes the contents of the audio file to ensure there are no problems with the format or quality.

[1623] 3. If the audio file matches, the server stores the file in its database and assigns it a unique ID.

[1624] 4. The server notifies the device that the audio file has been saved.

[1625] Training generative AI models

[1626] 1. After the audio file is saved, the server performs a characteristic analysis of the deceased person's audio data.

[1627] 2. The server uses a feature extraction algorithm to analyze and extract features such as pitch, speed, and intonation of the speech.

[1628] 3. Train a generative AI model (e.g., a large-scale language model or a speech synthesis model) based on the extracted feature data.

[1629] 4. Once trained, the generative AI model is optimized to reproduce the vocal characteristics of the deceased person.

[1630] Enter and send text messages

[1631] 1. The user enters an arbitrary text message through the application (e.g., "How are you feeling today?").

[1632] 2. The user clicks the "Send" button to send the text message to the server.

[1633] Server-generated audio

[1634] 1. The server parses the received text message.

[1635] 2. The server inputs the analyzed text into a generative AI model and generates corresponding audio data.

[1636] 3. The generative AI model uses the trained data to convert the input text into audio data in the characteristic voice of the deceased.

[1637] 4. The server sends the generated voice data to the terminal.

[1638] Receiving and playing audio on the device

[1639] 1. The device analyzes the voice data received from the server.

[1640] 2. The device checks the integrity of the audio data and prepares for playback.

[1641] 3. The device plays the audio data within the application, allowing the user to hear the voice of the deceased.

[1642] Specific examples

[1643] Consider a case where a user uses a smartphone app to upload audio data of their deceased grandmother. The audio data is uploaded to the app, and the smartphone sends the data to a server. The server analyzes the audio data, extracts the characteristics of the grandmother's voice, and trains an AI model. Next, when the user types a text message into the app saying, "Tell me about your favorite flower," the server receives this message and uses the trained AI model to generate audio data in the grandmother's voice. The generated audio data is sent to the smartphone and played back by the app. The user can hear their grandmother's voice again, which can be moving and comforting.

[1644] The above is an example of the implementation of the present invention. This system allows the bereaved family to re-create a heart-to-heart conversation with the deceased, and to gain healing and empathy.

[1645] The processing flow will be explained below.

[1646] Step 1:

[1647] User

[1648] The user launches the dedicated application and clicks the "Upload audio data" button.

[1649] Step 2:

[1650] User

[1651] The user selects the deceased person's audio file (e.g. MP3, WAV) and begins uploading.

[1652] Step 3:

[1653] Terminal

[1654] The device captures the selected audio file and checks the format and quality of the audio file.

[1655] Step 4:

[1656] Terminal

[1657] If the format and quality match, the terminal sends the audio file to the server.

[1658] Step 5:

[1659] server

[1660] The server receives the audio file sent from the terminal.

[1661] Step 6:

[1662] server

[1663] The server analyzes the content of the audio file to ensure that there are no problems with the format or quality.

[1664] Step 7:

[1665] server

[1666] If the audio file matches, the server stores the file in a database and assigns it a unique ID.

[1667] Step 8:

[1668] server

[1669] The server notifies the terminal that the audio file has been saved.

[1670] Step 9:

[1671] server

[1672] The server uses the stored audio files to characterize the deceased person's audio data.

[1673] Step 10:

[1674] server

[1675] The server uses feature extraction algorithms to analyze and extract features such as pitch, speed, and intonation of the speech.

[1676] Step 11:

[1677] server

[1678] Based on the extracted feature data, generative AI models (e.g., large-scale language models or speech synthesis models) are trained.

[1679] Step 12:

[1680] server

[1681] Once trained, the generative AI model is optimized to reproduce the vocal characteristics of the deceased person.

[1682] Step 13:

[1683] User

[1684] Through the application, the user enters an arbitrary text message (e.g., "How are you feeling today?").

[1685] Step 14:

[1686] User

[1687] The user clicks the "Send" button to send the text message to the server.

[1688] Step 15:

[1689] Terminal

[1690] The terminal transmits the input text message to the server and makes a request to the server.

[1691] Step 16:

[1692] server

[1693] The server parses the received text message.

[1694] Step 17:

[1695] server

[1696] The server inputs the analyzed text into a generative AI model and generates corresponding audio data.

[1697] Step 18:

[1698] Generative AI Models

[1699] The generative AI model uses trained data to convert input text into audio data in the characteristic voice of the deceased.

[1700] Step 19:

[1701] server

[1702] The server transmits the generated voice data to the terminal.

[1703] Step 20:

[1704] Terminal

[1705] The terminal analyzes the voice data received from the server.

[1706] Step 21:

[1707] Terminal

[1708] The device checks the integrity of the audio data and prepares it for playback.

[1709] Step 22:

[1710] Terminal

[1711] The device plays the audio data within the application, allowing the user to hear the voice of the deceased.

[1712] Example 1

[1713] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1714] Technology that allows bereaved families and friends to hear the voices of the deceased again can bring emotion and healing, but conventional systems require a lot of work, such as uploading the voice data, checking the quality, and training the AI ​​model, making it difficult to operate efficiently.In addition, because the processes of checking the voice data quality, extracting features, and training the AI ​​model are complex, making it easy for users to use is a challenge.

[1715] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1716] In this invention, the server includes a means for a user to upload voice data of the deceased via a terminal, a means for the terminal to check the format and quality of the voice data and send it to the server, and a means for the server to analyze the received voice data and extract feature data, thereby enabling a user to efficiently and easily use the voice data of the deceased and reproduce the deceased's voice using an AI model.

[1717] A "user" is a person who uses the system to upload and play back audio data of a deceased person.

[1718] A "terminal" is a device operated by a user that uploads audio data, checks the format and quality, transmits it to a server, and plays back the received audio data.

[1719] A "server" is a device that analyzes received voice data, extracts feature data, trains a generative AI model, generates voice data, and transmits it to a terminal.

[1720] "Audio data" refers to a digital data file (e.g., MP3, WAV) that records the voice of a deceased person.

[1721] "Uploading" is the act of a user sending audio data to a server via a terminal.

[1722] "Format" refers to the format of the audio data file (e.g., MP3, WAV).

[1723] "Quality" refers to characteristics related to the sound quality of data, such as the sampling rate and bit depth of the audio data.

[1724] "Analysis" is the process by which the server performs detailed analysis of the received audio data to extract specific information.

[1725] "Feature data" is information that indicates voice characteristics such as pitch, speed, and intonation extracted from voice data.

[1726] A "generative AI model" is an artificial intelligence model trained to convert text messages into audio data based on feature data.

[1727] "Training" is the process of learning from feature data so that the generative AI model can accurately reproduce the voice of the deceased.

[1728] A "text message" is text information that is entered by a user and sent to a server.

[1729] "Conversion" is the process of using a generative AI model to turn a text message into audio data.

[1730] "Reception" refers to the act of the server receiving data transmitted from the user terminal, or the act of the user terminal receiving data transmitted from the server.

[1731] "Playback" refers to the act of outputting audio data received by a user terminal as audio so that the user can hear it.

[1732] This invention is a system that uses the voice data of the deceased to allow bereaved family members and friends to hear the voice of the deceased again through AI. This system consists of three main elements: a user, a terminal, and a server. The specific configuration and embodiments of this system are described below.

[1733] First, the user launches a dedicated application on their device. This device can be a smartphone, tablet, or computer, which is capable of importing and playing audio data. When the user clicks the "Upload Audio Data" button, a screen appears allowing them to select the audio file (e.g., MP3, WAV) of the deceased person. The user selects the target audio file and begins uploading.

[1734] The device then retrieves the selected audio file and checks its format and quality. This check checks whether the audio file is in MP3 or WAV format and meets certain audio quality standards (e.g., 44.1kHz, 16-bit). If the quality is acceptable, the device sends the audio file to the server. During this process, the user sees an "Uploading" progress bar.

[1735] When the server receives the audio file sent from the device, it checks the format and quality of the audio data again. If there are no problems, the server saves the audio file in its database and assigns a unique identifier (ID) to the file. The server notifies the device that the file has been saved, allowing the user to confirm that the upload was successful.

[1736] The server extracts feature data from the stored audio files. Specifically, it uses a feature extraction algorithm to analyze and extract features such as pitch, speed, and intonation. This feature data is used to train a generative AI model (e.g., a large-scale language model or a speech synthesis model). This model is optimized to reproduce the characteristics of the deceased person's voice.

[1737] If a user wants to use the voice of a deceased person, they can enter any text message through the application, for example, "How are you feeling today?" or "Tell me about a flower you liked," and click the "Send" button. This message will be sent to the server.

[1738] The server analyzes the received text message and inputs it into a generative AI model to generate corresponding audio data. This AI model converts the input text into audio data in the deceased's characteristic voice based on the trained audio feature data. The generated audio data is then sent to the device.

[1739] The audio data received by the device is checked for integrity and then prepared for playback. Finally, the audio data is played back within the application, allowing the user to hear the voice of the deceased. This recreates an emotional conversation with the deceased, bringing emotion and healing to bereaved family and friends.

[1740] Specifically, when a user uses a smartphone app to upload audio data of their deceased grandmother, the audio data is uploaded to the app, and the smartphone sends the data to a server. The server analyzes the audio data, extracts the characteristics of the grandmother's voice, and trains an AI model. When the user types a text message into the app saying, "Tell me about your favorite flower," the server receives the message and uses the trained AI model to generate audio data in the grandmother's voice. The generated audio data is sent to the smartphone and played back in the app. The user can hear their grandmother's voice again, and feel moved and comforted.

[1741] An example of a prompt is as follows:

[1742] Prompt: "Tell me about a flower you loved."

[1743] Generated text: "Hello. I love talking about flowers you liked. My favorite was roses..."

[1744] This invention allows users to relive the heartfelt conversation by listening to the voice of the deceased again, and thus find healing.

[1745] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1746] Step 1:

[1747] The user launches the dedicated application on the device and clicks the "Upload audio data" button. When the user selects an audio file to upload, this file is imported into the device. The input is the audio file selected by the user, and the output is the imported audio file.

[1748] Step 2:

[1749] The device checks the format and quality of the selected audio file. Specifically, it checks that the file is in MP3 or WAV format and that the audio quality is 44.1 kHz and 16-bit. The input is the imported audio file, and the output is the result of matching the format and quality.

[1750] Step 3:

[1751] The device checks the format and quality and, if it is found to be compatible, sends the audio file to the server. A progress bar indicating "uploading" is displayed to the user during the upload. The input is the format and quality check result and the audio file, and the output is the audio file sent to the server.

[1752] Step 4:

[1753] The server receives the audio file sent from the terminal. The input is the audio file sent from the terminal, and the output is the received audio file.

[1754] Step 5:

[1755] The server rechecks the content of the received audio file and, if there are no problems, stores it in a database. Specifically, it analyzes the format and quality of the audio data again and assigns a unique identifier. The input is the received audio file, and the output is the audio file stored in the database and its unique identifier.

[1756] Step 6:

[1757] The server performs feature analysis on the audio files stored in the database. It uses a feature extraction algorithm to analyze and extract features such as pitch, speed, and intonation of the voice. Based on this analysis, feature data is generated. The input is the audio files stored in the database, and the output is the extracted feature data.

[1758] Step 7:

[1759] The server uses the extracted feature data to train a generative AI model. Specifically, the feature data is input into the AI ​​model and the AI ​​model is optimized through an appropriate learning process. The input is the feature data, and the output is a trained generative AI model.

[1760] Step 8:

[1761] A user enters a text message through an application. The user clicks a "Send" button to send the entered text message to a server. The input is the text message entered by the user, and the output is the text message sent to the server.

[1762] Step 9:

[1763] The server analyzes the received text message. The analyzed text message is input into the generative AI model to generate corresponding audio data. The input is the text message sent to the server, and the output is the generated audio data.

[1764] Step 10:

[1765] The server sends the generated voice data to the terminal. The input is the generated voice data, and the output is the voice data sent to the terminal.

[1766] Step 11:

[1767] The device analyzes the audio data received from the server. It checks the integrity of the audio data and prepares for playback. The input is the audio data received from the server, and the output is the audio data whose integrity has been checked.

[1768] Step 12:

[1769] The device plays back the audio data within the application, allowing the user to hear the voice of the deceased. The input is the audio data whose integrity has been verified, and the output is the played-back audio.

[1770] (Application example 1)

[1771] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1772] Technology that allows users to feel moved and comforted by re-listening to the voices of the deceased is extremely valuable. However, conventional systems have difficulty using the voice data of the deceased to read text aloud or generate audiobooks. This makes it difficult for users to experience the added emotional impact of listening to books or poems in the voice of the deceased.

[1773] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1774] In this invention, the server includes: [means for uploading the deceased's voice data from the user terminal to the server;] [means for analyzing the voice data received by the server and extracting feature data;] [means for training a generative AI model using the extracted feature data;] [means for the server to receive text selected by the user and convert the text into voice data using the trained generative AI model;] [means for the user terminal to receive voice data sent from the server and play it back to the user; and [means for sending the text selected by the user to the server and generating an audiobook in the deceased's voice. This allows the user to listen to books and poems in the deceased's voice.

[1775] "Voice data of the deceased" refers to information in the form of data that contains recorded or collected voices made by the deceased during their lifetime.

[1776] A "user terminal" is an information processing device such as a computer, smartphone, or tablet that is used by a user.

[1777] A "server" is a computer system that provides services over a network and sends and receives voice data from user terminals and analyzes it.

[1778] "Feature data" is data that indicates sound characteristics such as pitch, speed, and intonation extracted from voice data.

[1779] A "generative AI model" is an artificial intelligence model that learns the characteristics of voice data and is used to convert specified text into synthetic speech.

[1780] A "text message" is character string information entered by a user, which is received by the server and converted into voice data.

[1781] A "unique ID" is a unique identifier assigned to identify a particular piece of audio data within a database.

[1782] An "audiobook" is an audio version of the contents of a book or poem, and is audio data generated using the voice of a deceased person.

[1783] MODE FOR CARRYING OUT THE INVENTION

[1784] This invention is a system that allows users to feel touched or comforted by using the voice data of the deceased. The specific configuration and processing flow are described below.

[1785] System program and processing description

[1786] This system is constructed from a user device, a server, and a generative AI model. Each component plays the following role:

[1787] Hardware and software used:

[1788] User device: A smartphone, tablet, or computer is used to run applications, upload audio data, and play audiobooks.

[1789] Server: Located on the cloud, it is responsible for analyzing voice data, extracting features, and training and generating generative AI models.

[1790] Generative AI models: Convert text to audio using large-scale language models (e.g., GPT-3) and speech synthesis models (e.g., TTS models).

[1791] Cloud storage: Amazon Web Services (AWS) S3 and other cloud storage services are used to store audio data and generated audiobooks.

[1792] Data processing and data calculations:

[1793] 1. Uploading and sending audio data

[1794] The user uploads the deceased's audio data to the server from their device. First, the user selects an audio file (e.g., MP3, WAV) via a dedicated app and sends the data to the server. The server checks the format and quality of the received audio data and stores the appropriate audio data in cloud storage.

[1795] 2. Extracting feature data and training the AI ​​model

[1796] The server analyzes the stored audio data and extracts feature data. The feature extraction algorithm analyzes characteristics such as pitch, speed, and intonation of the voice, and uses this data to train a generative AI model. The trained model is optimized to reproduce the characteristics of the deceased person's voice.

[1797] 3. Entering text messages and generating voice data

[1798] The user selects the text of a book or poem they want to read through the application and sends it to the server, which analyzes the text and uses a generative AI model to convert the text into audio data in the deceased person's voice. The converted audiobook data is then stored in cloud storage.

[1799] 4. Receiving and playing audio data

[1800] The user device receives the generated audiobook from the server, verifies its integrity, and then plays it within the application, allowing the user to listen to books and poems in the voice of the deceased.

[1801] Examples:

[1802] For example, if a user uses a smartphone app to upload audio data of a deceased person, the audio data is uploaded to the app, and the smartphone sends the data to a server. The server analyzes the audio data, extracts the voice characteristics of the deceased person (e.g., grandmother), and trains an AI model. Next, when the user selects a poetry collection and enters it into the app, the server receives the poetry collection and uses the trained AI model to generate audio data in the deceased person's (grandmother's) voice. The generated audio data is sent to the smartphone and played back in the app. The user can relisten to the poetry collection in the deceased's voice, which can be moving and comforting.

[1803] Example prompt sentence:

[1804] "Generate an audiobook of this text read in the voice of the deceased: Complete Poems"

[1805] The above is a specific embodiment for carrying out the present invention. This system allows the user to re-create a heart-to-heart conversation with the deceased, and can bring about deep emotions and healing.

[1806] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1807] Step 1:

[1808] The user launches the dedicated application and clicks the "Upload audio data" button. The user selects the deceased's audio file (e.g., MP3, WAV) and starts uploading. The input is the audio file, and the output is the device sending the audio file to the server.

[1809] Step 2:

[1810] The terminal takes the selected audio file, checks the format and quality of the audio file, and if it matches, sends the audio file to the server. The input is the audio file, and the output is the audio data whose format and quality have been checked.

[1811] Step 3:

[1812] The server receives the audio file sent from the device. It analyzes the content of the audio file and checks that there are no problems with the format or quality. It stores the matching audio file in cloud storage and assigns a unique ID. The input is the audio file, and the output is the stored audio data and the unique ID.

[1813] Step 4:

[1814] The server analyzes the stored voice data and extracts feature data. Feature extraction algorithms are used to analyze the voice's pitch, speed, intonation, etc., and the data is used to train the generative AI model. The input is the stored voice data, and the output is the feature data.

[1815] Step 5:

[1816] The server uses the extracted feature data to train a generative AI model, which is optimized to reproduce the characteristics of the deceased person's voice. The input is the feature data, and the output is the trained generative AI model.

[1817] Step 6:

[1818] The user selects the text of a book or poem they want to read through the application and sends it to the server. The input is a text message, and the output is the sent text message.

[1819] Step 7:

[1820] The server analyzes the received text message and converts the text content into audio data using a trained generative AI model. The input is the text message and the output is the generated audio data.

[1821] Step 8:

[1822] The server stores the generated voice data in cloud storage and transmits it to the user's device. The input is the generated voice data, and the output is the transmitted voice data.

[1823] Step 9:

[1824] The device analyzes the audio data received from the server and verifies its integrity. The audio data is then played back within the application, allowing the user to hear the voice of the deceased. The input is the audio data received from the server, and the output is the played audio data.

[1825] The above are the specific processing steps of the system based on the application example, through which users can create audiobooks using the voices of deceased people and listen to them again.

[1826] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1827] This invention is a system that combines the voice data of the deceased with an emotion engine that recognizes the user's emotions, allowing bereaved family and friends to hear the deceased's voice again through AI and providing adapted responses based on the user's emotions. The program and specific processing flow of this system are shown below.

[1828] Program Overview

[1829] The program works through the following major steps:

[1830] 1. The user uploads the audio data of the deceased person.

[1831] 2. The device sends the audio data to the server.

[1832] 3. The server analyzes the voice data and extracts feature data.

[1833] 4. The server uses the feature data to train a generative AI model.

[1834] 5. The emotion engine recognizes the user's emotions.

[1835] 6. The user enters a text message, and the emotion engine sends the emotion analysis results to the server.

[1836] 7. The server uses the sentiment analysis results to generate audio data based on the text message.

[1837] 8. The device receives and plays the generated audio data.

[1838] Processing flow and specific examples

[1839] User voice data input

[1840] 1. The user launches the dedicated application and clicks the "Upload audio data" button.

[1841] 2. The user selects the audio file (e.g. MP3, WAV) of the deceased person and begins uploading.

[1842] Terminal handling

[1843] 3. The device will import the selected audio file and check the format and quality of the audio file.

[1844] Sending terminal

[1845] 4. If the format and quality match, the device sends the audio file to the server.

[1846] Validation and storage of voice data by the server

[1847] 1. The server receives the audio file sent from the device.

[1848] 2. The server analyzes the contents of the audio file to ensure there are no problems with the format or quality.

[1849] 3. If the audio file matches, the server stores the file in its database and assigns it a unique ID.

[1850] 4. The server notifies the device that the audio file has been saved.

[1851] Training generative AI models

[1852] 1. The server uses the stored audio files to characterize the deceased person's audio data.

[1853] 2. The server uses a feature extraction algorithm to analyze and extract features such as pitch, speed, and intonation of the speech.

[1854] 3. Train a generative AI model (e.g., a large-scale language model or a speech synthesis model) based on the extracted feature data.

[1855] 4. Once trained, the generative AI model is optimized to reproduce the vocal characteristics of the deceased person.

[1856] Emotion Engine Operation

[1857] 1. When a user uses an application, the emotion engine acquires the user's facial expression data and voice tone data through the camera and microphone.

[1858] 2. The emotion engine analyzes the acquired data and determines the user's emotional state (e.g., joy, sadness, surprise).

[1859] Enter and send text messages

[1860] 1. The user enters a text message through the application (e.g., "How have things been lately?").

[1861] 2. The emotion engine sends the user's text message along with the analyzed emotion results to the server.

[1862] Server-generated audio

[1863] 1. The server analyzes the received text message and the sentiment analysis results.

[1864] 2. The server inputs the analyzed text and emotion data into the generative AI model and generates corresponding voice data.

[1865] 3. Using the trained data, the generative AI model converts input text into audio data in the characteristic voice of the deceased based on the user's emotional state.

[1866] 4. The server sends the generated voice data to the terminal.

[1867] Receiving and playing audio on the device

[1868] 1. The device analyzes the voice data received from the server.

[1869] 2. The device checks the integrity of the audio data and prepares for playback.

[1870] 3. The device plays the audio data within the application, allowing the user to hear the voice of the deceased.

[1871] Specific examples

[1872] Consider a scenario where a user uses a smartphone app to upload audio data of their deceased grandmother. The audio data is uploaded to the app, and the smartphone sends the data to a server. The server analyzes the audio data, extracts the characteristics of the grandmother's voice, and trains an AI model. Next, when the user types a text message into the app saying, "Tell me about your favorite flower," the emotion engine analyzes the user's emotions and sends the emotion data to the server. Based on this information, the server uses the trained AI model to generate audio data in the grandmother's voice. The generated audio data is sent to the smartphone and played back by the app. The user can hear their grandmother's voice again, and experience deeper emotion and healing from the emotionally sensitive response.

[1873] The above is an example of the implementation of the present invention. This system recreates the heart-to-heart dialogue with the deceased and provides responses tailored to the user's emotions, allowing the bereaved to receive even deeper healing and empathy.

[1874] The processing flow will be explained below.

[1875] Step 1:

[1876] User

[1877] The user launches the dedicated application and clicks the "Upload audio data" button.

[1878] Step 2:

[1879] User

[1880] The user selects the deceased person's audio file (e.g. MP3, WAV) and begins uploading.

[1881] Step 3:

[1882] Terminal

[1883] The device captures the selected audio file and checks the format and quality of the audio file.

[1884] Step 4:

[1885] Terminal

[1886] If the format and quality match, the terminal sends the audio file to the server.

[1887] Step 5:

[1888] server

[1889] The server receives the audio file sent from the terminal.

[1890] Step 6:

[1891] server

[1892] The server analyzes the content of the audio file to ensure that there are no problems with the format or quality.

[1893] Step 7:

[1894] server

[1895] If the audio file matches, the server stores the file in a database and assigns it a unique ID.

[1896] Step 8:

[1897] server

[1898] The server notifies the terminal that the audio file has been saved.

[1899] Step 9:

[1900] server

[1901] The server uses the stored audio files to characterize the deceased person's audio data.

[1902] Step 10:

[1903] server

[1904] The server uses feature extraction algorithms to analyze and extract features such as pitch, speed, and intonation of the speech.

[1905] Step 11:

[1906] server

[1907] Based on the extracted feature data, generative AI models (e.g., large-scale language models or speech synthesis models) are trained.

[1908] Step 12:

[1909] server

[1910] Once trained, the generative AI model is optimized to reproduce the vocal characteristics of the deceased person.

[1911] Step 13:

[1912] Emotion Engine

[1913] When a user uses an application, the emotion engine acquires the user's facial expression data and voice tone data through a camera and microphone.

[1914] Step 14:

[1915] Emotion Engine

[1916] The emotion engine analyzes the acquired data and determines the user's emotional state (e.g., joy, sadness, surprise).

[1917] Step 15:

[1918] User

[1919] The user enters an arbitrary text message (e.g., "How have things been lately?") through the application.

[1920] Step 16:

[1921] User

[1922] The user clicks the "Send" button to send the text message to the server.

[1923] Step 17:

[1924] Emotion Engine

[1925] The emotion engine sends the user's text message along with the analyzed emotion results to the server.

[1926] Step 18:

[1927] server

[1928] The server analyzes the received text message and the sentiment analysis results.

[1929] Step 19:

[1930] server

[1931] The server inputs the analyzed text and emotion data into a generative AI model to generate corresponding voice data.

[1932] Step 20:

[1933] Generative AI Models

[1934] The generative AI model uses trained data to convert input text and emotional data into audio data in the characteristic voice of the deceased.

[1935] Step 21:

[1936] server

[1937] The server transmits the generated voice data to the terminal.

[1938] Step 22:

[1939] Terminal

[1940] The terminal analyzes the voice data received from the server.

[1941] Step 23:

[1942] Terminal

[1943] The device checks the integrity of the audio data and prepares it for playback.

[1944] Step 24:

[1945] Terminal

[1946] The device plays the audio data within the application, allowing the user to hear the voice of the deceased.

[1947] As a concrete example, consider the case where a user uses a smartphone app to upload voice data of their deceased grandmother and engages in emotionally appropriate conversations through an emotion engine. The voice file uploaded by the user is received by a server, which analyzes and extracts features. A generative AI model is trained based on this data. The user then performs emotion analysis through the emotion engine and enters and sends a text message. The data including the emotion analysis results is passed to the server, and emotion-appropriate voice data is generated in the grandmother's voice through the generative AI model. This voice data is received and played back on the device. The user can hear the grandmother's voice emotionally, and experience deep emotion and healing.

[1948] Example 2

[1949] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1950] For bereaved family members and friends who wish to reconnect with their loved ones, recreating their voice and receiving emotionally responsive responses is extremely valuable. However, simply playing back audio data does not provide a deep conversational experience with the deceased. Furthermore, current technology lacks a system that can recognize the user's emotions and respond accordingly.

[1951] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1952] In this invention, the server includes: [means for uploading the voice data of the deceased from the user terminal to the server; [means for analyzing the voice data received by the server and extracting feature data; [means for training a generative AI model using the extracted feature data; [means for an emotion engine to recognize the user's emotions and acquire facial expression data and voice tone data; [means for the server to receive a text message entered by the user and convert the text message into voice data using the emotion analysis results; and [means for the user terminal to receive the voice data sent from the server and play it back to the user. This enables the user to have a conversation with the deceased that is sensitive to their emotions.

[1953] "Voice data of a deceased person" refers to a digital file uploaded by a user that records the voice of a deceased person during their lifetime.

[1954] "User terminal" refers to electronic devices used by users, such as smartphones, tablets, and personal computers.

[1955] "Server" refers to a high-performance computer system for analyzing, storing, and generating voice data.

[1956] "Analysis" refers to the process of examining and examining the content of received voice data in detail to extract necessary feature data.

[1957] "Feature data" refers to information that quantifies the unique characteristics of voice data, such as voice pitch, speed, and intonation.

[1958] "Generative AI model" refers to an artificial intelligence model that generates voice based on feature data extracted from the voice data of a deceased person.

[1959] "Training" refers to the process of learning using data to enable a generative AI model to reproduce the vocal characteristics of the deceased person.

[1960] An "emotion engine" refers to an algorithm and system for recognizing and analyzing user emotions in real time.

[1961] "Facial expression data" refers to data that quantifies and classifies changes in a user's facial expressions captured through a camera.

[1962] "Voice tone data" refers to data obtained by quantifying and analyzing changes in the tone of a user's speech obtained through a microphone.

[1963] "Emotion analysis results" refers to data in which the emotion engine analyzes the user's emotional state and expresses the results in numerical values ​​and categories.

[1964] "Text message" refers to text information entered by a user through an application.

[1965] "Audio Data" refers to the digital audio files generated by the generative AI model that reproduce the voice of the deceased person.

[1966] This invention is a system that allows bereaved family and friends to hear the voice of the deceased again through AI by combining the voice data of the deceased with an emotion engine that recognizes the user's emotions. The system provides adaptive responses based on the emotions. Below, we will explain in detail how this system is implemented.

[1967] System configuration and processing overview

[1968] 1. User voice data input

[1969] The user launches the dedicated application and clicks the "Upload audio data" button.

[1970] The user selects the audio file of the deceased person (e.g., MP3, WAV format) and begins uploading.

[1971] 2. Terminal Processing

[1972] The device will then import the selected audio file and check the format and quality of the audio file using an audio analysis library (e.g., FFmpeg).

[1973] 3. Sending the device

[1974] If the format and quality match, the device sends the audio file to the server via an HTTP POST request.

[1975] 4. Verification and storage of audio data by the server

[1976] The server receives the audio file sent from the device, analyzes it again, and checks that there are no problems with the format or quality. If the audio file is suitable, the server stores it in a database (e.g., MongoDB) and assigns it a unique ID.

[1977] The server sends a storage completion notification to the terminal.

[1978] 5. Training the generative AI model

[1979] The server uses the stored audio files to perform a feature analysis of the deceased person's voice data. The feature analysis of the voice data is performed using a voice analysis tool (e.g., Praat, Kaldi).

[1980] The server uses feature extraction algorithms to extract features such as pitch, speed, and intonation, and then uses them to train a generative AI model (e.g., GPT-3, Tacotron 2).

[1981] 6. Operation of the Emotion Engine

[1982] The emotion engine captures facial expression data and voice tone data via a camera or microphone when a user uses an application, and analyzes the user's emotional state in real time using facial expression recognition algorithms (e.g., OpenCV, Dlib) and voice tone analysis tools (e.g., Praat).

[1983] 7. Enter and send text messages

[1984] The user types a text message through the application (e.g., "How's it been lately?").

[1985] The emotion engine sends a text message along with the analyzed emotion results to the server.

[1986] 8. Server-generated audio

[1987] The server analyzes the received text messages and emotion analysis results, inputs them into a generative AI model, and generates audio data in the voice of the deceased.

[1988] The server transmits the generated voice data to the terminal.

[1989] 9. Receiving and Playing Audio by the Device

[1990] The device analyzes the audio data received from the server, checks its integrity, and then plays the audio data within the application.

[1991] Specific use cases

[1992] For example, consider a case where a user uploads audio data of a deceased person using a smartphone app. The user uploads the audio data to the app, and the device sends the data to a server. The server analyzes the audio data, extracts the characteristics of the deceased's voice, and trains a generative AI model. Next, when the user types "Tell me about your favorite flower" into the app, the emotion engine analyzes the user's emotions and generates the most appropriate response based on those emotions. The generated audio data is sent to the smartphone, and the user can listen to the deceased's voice again in the app and enjoy a dialogue tailored to their emotions.

[1993] Prompt Sentence Examples

[1994] "Tell me about a flower you liked."

[1995] "How was your day today?"

[1996] "Tell me about a memory of going out together."

[1997] The above is a concrete embodiment for carrying out the present invention. Through this system, the user can hear the voice of the deceased again and receive emotional responses.

[1998] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1999] Step 1:

[2000] The user launches the dedicated application and clicks the "Upload audio data" button. The user selects the audio file (e.g., MP3, WAV) of the deceased person and begins uploading.

[2001] Input: User selection of an audio file.

[2002] Output: The selected audio file.

[2003] Step 2:

[2004] The device will then import the selected audio file and check the format and quality of the audio file, using an audio analysis library (e.g. FFmpeg) to analyze the bit rate, sampling rate, etc.

[2005] Input: The selected audio file.

[2006] Output: Data about the format and quality of the audio file.

[2007] Step 3:

[2008] If the audio file format and quality are compatible, the device sends the audio file to the server via an HTTP POST request.

[2009] Input: Quality checked audio files.

[2010] Output: The audio file sent to the server.

[2011] Step 4:

[2012] The server receives the audio file sent from the device and analyzes its content. It uses an audio analysis tool (e.g., Praat, Kaldi) to check for format and quality issues. If there are no problems, it stores the audio file in a database (e.g., MongoDB) and assigns a unique ID.

[2013] Input: The audio file sent from the device.

[2014] Output: The saved audio file and its unique ID.

[2015] Step 5:

[2016] The server uses the stored audio files to extract feature data. It uses speech analysis algorithms to analyze and extract features such as pitch, speed, and intonation. It then trains a generative AI model (e.g., GPT-3, Tacotron 2) based on the extracted feature data.

[2017] Input: Saved audio file.

[2018] Output: feature data and a trained generative AI model.

[2019] Step 6:

[2020] The emotion engine captures the user's facial expression data and voice tone data via the camera and microphone when the user uses the application, and performs real-time emotion analysis using facial expression recognition algorithms (e.g., OpenCV, Dlib) and voice tone analysis tools.

[2021] Input: User's facial expression data and voice tone data.

[2022] Output: User sentiment analysis results.

[2023] Step 7:

[2024] The user inputs a text message through the application (e.g., "How have things been lately?"). The emotion engine analyzes the text message and sends it to the server along with the resulting emotion.

[2025] Input: User text message and sentiment analysis results.

[2026] Output: The text message sent to the server and the sentiment result.

[2027] Step 8:

[2028] The server analyzes the received text message and the emotion analysis results, inputs them into a generative AI model, and generates audio data in the deceased's voice. The generated audio data is then sent from the server to the device.

[2029] Input: Text message, sentiment analysis results, pre-trained generative AI model.

[2030] Output: The generated audio data.

[2031] Step 9:

[2032] The device analyzes the audio data received from the server, checks its integrity, and then plays the audio data within the application.

[2033] Input: Audio data received from the server.

[2034] Output: The audio data that is played within the application.

[2035] (Application example 2)

[2036] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[2037] In recent years, there has been an increasing need to cherish memories of deceased loved ones. However, conventional technologies have difficulty reproducing the voices of the deceased, and are unable to provide responses that are particularly sensitive to the user's emotions. Furthermore, more advanced technology is required to enable users to re-experience conversations with their deceased loved ones and achieve emotional healing. Furthermore, there is a need for a system that can provide a deeper experience in both real and virtual spaces by realizing this through specific devices.

[2038] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[2039] In this invention, the server includes: [means for uploading the voice data of the deceased from the user terminal to the server; [means for analyzing the voice data received by the server and extracting feature data; [means for training a generative AI model using the extracted feature data; [means for the server to receive a text message entered by the user and convert the text message into voice data using the trained generative AI model; [means for the user terminal to receive the voice data sent from the server and play it back to the user; [means for an emotion engine to recognize the user's emotions and provide a response based on those emotions; and [means for playing back the response through a specific terminal (e.g., smartphone, smart glasses, etc.)]. This allows [the user to hear the voice of the deceased again and receive a response that is in tune with their emotions, thereby achieving deeper healing and empathy].

[2040] A "user terminal" is an electronic device operated by a user, and includes devices such as smartphones, tablets, and smart glasses.

[2041] A "server" is a central system that provides data and services over a network, and performs various processes such as analyzing, storing, and generating voice data.

[2042] "Audio data" refers to data in file format that records the voice of the deceased, and refers to information that includes the characteristics of the voice.

[2043] "Analysis" is the process in which the server analyzes the received voice data and extracts feature data.

[2044] "Feature data" refers to individual characteristic information such as pitch, speed, and intonation extracted from speech data.

[2045] A "generative AI model" is an artificial intelligence model trained based on extracted feature data, and is responsible for converting text messages into voice data.

[2046] A "text message" is a textual message entered by a user and sent to a server.

[2047] "Converting to voice data" refers to the process in which a generative AI model generates voice-format data based on the input text message.

[2048] An "emotion engine" is a system that recognizes a user's emotions and generates a response according to those emotions.

[2049] "Play" refers to the operation of the user terminal playing back the audio data received from the server as audio.

[2050] This invention is a system that reproduces the voice of a deceased person and provides responses that are in tune with the user's emotions. This system is realized mainly using a user terminal, a server, an emotion engine, and a generative AI model.

[2051] System configuration

[2052] User terminal

[2053] A user terminal is an electronic device operated by a user, including a smartphone, tablet, smart glasses, etc. This terminal is used by the user to upload voice data and perform subsequent interactions.

[2054] server

[2055] The server is a central system that provides data and services over the network. It is responsible for various processes such as analyzing, storing, and generating voice data. This server receives voice data sent from user devices, extracts feature data, and trains generative AI models. It is also responsible for the process of converting text messages entered by users into voice data.

[2056] Emotion Engine

[2057] The emotion engine is a system that recognizes the user's emotions and generates responses based on those emotions. The emotion engine acquires and analyzes the user's facial expressions and voice tone data via the camera and microphone on the user's device.

[2058] Generative AI Models

[2059] A generative AI model is an artificial intelligence model trained using extracted feature data. This model generates voice data that reproduces the characteristics of the deceased based on the input text message. Examples of such models include the Wav2Vec2 model and EmotionRecognizer.

[2060] Data processing and calculation

[2061] The system performs the following data processing and calculations:

[2062] 1. Analysis of speech data: The speech data of the deceased person uploaded from the user's device to the server is analyzed and feature data is extracted. This analysis is performed using models such as Wav2Vec2Processor and Wav2Vec2ForCTC.

[2063] 2. Emotion Recognition: The emotion engine analyzes voice tone and facial expression data to recognize the user's emotions. This analysis is performed using emotion recognition models such as EmotionRecognizer.

[2064] 3. Voice data generation: The server uses the generative AI model to generate voice data that reproduces the voice of the deceased person based on the text message entered by the user, matching the emotions analyzed by the emotion engine.

[2065] Specific examples

[2066] A specific usage scenario is when a user wears smart glasses at a memorial cafe and uploads their mother's voice data. The user enters a prompt such as, "Mom, tell me about your recent trip." In this case, the emotion engine analyzes the user's emotions, and the server generates an optimal response in the mother's voice. The user's device then receives the generated voice data and plays it back to the user, allowing the user to hear their mother's voice again. This allows for deep healing and empathy that is tailored to their emotions.

[2067] Prompt Sentence Examples

[2068] "Mom, tell me about your recent trip."

[2069] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[2070] Step 1:

[2071] The user uploads the audio data of the deceased.

[2072] Description: The user launches the dedicated application and clicks the "Upload audio data" button. The user selects the audio file (e.g., MP3, WAV) of the deceased person and begins uploading.

[2073] Input: Audio file (e.g. MP3, WAV)

[2074] Output: Uploaded audio data

[2075] Step 2:

[2076] The terminal transmits the voice data to the server.

[2077] Description: The terminal captures the selected audio file, checks the format and quality of the audio file, and then sends the audio file to the server if it is suitable.

[2078] Input: Uploaded audio data

[2079] Output: Audio data sent to the server

[2080] Step 3:

[2081] The server analyzes the voice data and extracts feature data.

[2082] Description: The server analyzes the received audio data and extracts features such as pitch, speed, and intonation. The analysis is performed using Wav2Vec2Processor and Wav2Vec2ForCTC models.

[2083] Input: Audio data sent to the server

[2084] Output: feature data

[2085] Step 4:

[2086] The server uses the feature data to train a generative AI model.

[2087] Description: The server uses the extracted feature data to train a generative AI model, which is then optimized to reproduce the characteristics of the deceased person's voice.

[2088] Input: feature data

[2089] Output: A trained generative AI model

[2090] Step 5:

[2091] The emotion engine recognizes the user's emotions.

[2092] Description: When a user uses an application, the emotion engine captures and analyzes the user's facial expression data and voice tone data through the camera and microphone on the user's device. The analysis uses models such as EmotionRecognizer.

[2093] Input: User's facial expression data, voice tone data

[2094] Output: User emotion data

[2095] Step 6:

[2096] The user inputs a text message, and the emotion engine sends the emotion analysis results to the server.

[2097] Description: A user inputs any text message through the application. The emotion engine analyzes the user's text message and sends it to the server along with the emotion results.

[2098] Input: Text messages, emotion data

[2099] Output: Text message and emotion data sent to the server

[2100] Step 7:

[2101] The server generates voice data using the text message and the emotion data.

[2102] Description: The server inputs the received text message and emotional data into the generative AI model and generates corresponding voice data.

[2103] Input: Text messages, emotion data, pre-trained generative AI model

[2104] Output: Generated audio data

[2105] Step 8:

[2106] The terminal receives and plays the generated audio data.

[2107] Description: The device analyzes the audio data received from the server, checks its integrity, and then plays it back.

[2108] Input: Generated audio data

[2109] Output: Played audio

[2110] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[2111] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[2112] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[2113] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[2114] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[2115] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[2116] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[2117] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[2118] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[2119] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[2120] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[2121] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[2122] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[2123] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[2124] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[2125] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[2126] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[2127] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[2128] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[2129] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[2130] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[2131] The following is further disclosed regarding the above embodiment.

[2132] (Claim 1)

[2133] [Means for uploading the voice data of the deceased person from the user terminal to the server;

[2134] [Means for analyzing the voice data received by the server and extracting feature data;

[2135] [means for training a generative AI model using the extracted feature data; and

[2136] [Means for a server to receive a text message entered by a user and convert the text message into voice data using a trained generative AI model;

[2137] [Means for the user terminal to receive the audio data transmitted from the server and play it back to the user;

[2138] A system including:

[2139] (Claim 2)

[2140] The system of claim 1, further comprising: verifying the format and quality of the audio data.

[2141] (Claim 3)

[2142] [The system of claim 1, wherein the server stores the voice data in a database and assigns a unique ID.

[2143] "Example 1"

[2144] (Claim 1)

[2145] [Means for a user to upload audio data of the deceased via a terminal;

[2146] [Means for the terminal to check the format and quality of the audio data and send it to the server;

[2147] [Means for analyzing the voice data received by the server and extracting feature data;

[2148] [means for training a generative AI model using the extracted feature data; and

[2149] [Means for a server to receive a text message entered by a user and convert the text message into voice data using a trained generative AI model;

[2150] [Means for the user terminal to receive the audio data transmitted from the server and play it back to the user;

[2151] A system including:

[2152] (Claim 2)

[2153] The system of claim 1, further comprising: verifying the format and quality of the audio data.

[2154] (Claim 3)

[2155] [The system of claim 1, wherein the server stores the voice data in a database and assigns a unique identifier.

[2156] "Application Example 1"

[2157] (Claim 1)

[2158] [Means for uploading the voice data of the deceased person from the user terminal to the server;

[2159] [Means for analyzing the voice data received by the server and extracting feature data;

[2160] [means for training a generative AI model using the extracted feature data; and

[2161] [Means for the server to receive the text selected by the user and convert the text into speech data using a trained generative AI model;

[2162] [Means for the user terminal to receive the audio data transmitted from the server and play it back to the user;

[2163] A system including:

[2164] (Claim 2)

[2165] The system of claim 1, further comprising: verifying the format and quality of the audio data.

[2166] (Claim 3)

[2167] [The system of claim 1, wherein the server stores the voice data in a database and assigns a unique ID.

[2168] (Claim 4)

[2169] The system of claim 1, wherein the user sends selected text to a server, which generates an audiobook in the deceased person's voice.

[2170] "Example 2: Combining Emotion Engines"

[2171] (Claim 1)

[2172] [Means for uploading the voice data of the deceased person from the user terminal to the server;

[2173] [Means for analyzing the voice data received by the server and extracting feature data;

[2174] [means for training a generative AI model using the extracted feature data; and

[2175] [Means for the emotion engine to recognize the user's emotion and acquire facial expression data and voice tone data;

[2176] [Means for the server to receive the text message entered by the user and convert the text message into voice data using the emotion analysis results;

[2177] [Means for the user terminal to receive the audio data transmitted from the server and play it back to the user;

[2178] A system including:

[2179] (Claim 2)

[2180] The system of claim 1, further comprising: verifying the format and quality of the audio data.

[2181] (Claim 3)

[2182] [The system of claim 1, wherein the server stores the voice data in a database and assigns a unique ID.

[2183] "Application example 2 when combining emotion engines"

[2184] (Claim 1)

[2185] [Means for uploading the voice data of the deceased person from the user terminal to the server;

[2186] [Means for analyzing the voice data received by the server and extracting feature data;

[2187] [means for training a generative AI model using the extracted feature data; and

[2188] [Means for a server to receive a text message entered by a user and convert the text message into voice data using a trained generative AI model;

[2189] [Means for the user terminal to receive the audio data transmitted from the server and play it back to the user;

[2190] [means for the emotion engine to recognize the user's emotion and provide a response based on the emotion; and

[2191] [Means for playing back responses through a specific device (e.g., a smartphone, smart glasses, etc.);

[2192] A system including:

[2193] (Claim 2)

[2194] The system of claim 1, further comprising: verifying the format and quality of the audio data.

[2195] (Claim 3)

[2196] [The system of claim 1, wherein the server stores the voice data in a database and assigns a unique ID. [Explanation of symbols]

[2197] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. A means for uploading voice data of the deceased person from a user terminal to a server; A means for analyzing the voice data received by the server and extracting feature data; means for training a generative AI model using the extracted feature data; A means for a server to receive a text message entered by a user and convert the text message into voice data using a trained generative AI model; a means for receiving the audio data transmitted from the server by the user terminal and playing it back to the user; A system including:

2. 10. The system of claim 1, wherein the system verifies the format and quality of the audio data.

3. 2. The system of claim 1, wherein the server stores the voice data in a database and assigns a unique ID to the voice data.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A