System
The system addresses the challenge of recreating deceased loved ones' voices and faces by analyzing photographic and audio data to generate realistic video messages, ensuring secure storage and playback, thereby enhancing emotional connections.
Patent Information
- Application Number
- JP2024128497
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-08-02
- Publication Date
- 2026-02-16
AI Technical Summary
Existing systems struggle to realistically recreate the voices and faces of deceased loved ones, making it difficult for users to vividly recall their memories, and lack sufficient data protection and security measures.
A system that uploads photographic and audio data from a family member's lifetime, analyzes facial and vocal characteristics, generates text messages, converts them into audio messages, and creates videos using facial features, while ensuring secure storage and playback through cloud servers and security tokens.
Enables the realistic recreation of fond memories by generating high-quality, emotionally rich video messages that can be safely accessed and played in streaming format, enhancing emotional connections with loved ones.
Smart Images

Figure 2026025685000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] In many modern homes, especially in our digitalized society, there are limited ways to preserve fond memories of family members who have passed away. Photographs and audio recordings alone make it difficult to recreate memories in one's mind. Furthermore, as we grow older, it becomes increasingly difficult to vividly recall the voices and expressions of loved ones. Given this background, there is a demand for technology that can realistically recreate the voices and faces of cherished family members and leave a heartwarming message behind. [Means for solving the problem]
[0005] The present invention relates to a system that includes a means for uploading photographic data and audio data from a family member's lifetime and a means for analyzing the data to extract facial and vocal characteristics of the family member. It also includes a means for generating a text message based on the extracted facial and vocal characteristics and converting the text message into an audio message. It also includes a means for generating a video using the generated audio message and facial characteristics, saving it in cloud storage, and enabling it to be played in streaming format. It also includes a means for issuing a security token to verify user access rights for streaming playback, and a means for using a database of dialects and catchphrases to reflect specific patterns in the analysis and generation of audio data. In this way, a system is provided that can more realistically recreate fond memories and provide them to users.
[0006] "Photographic data taken before death" refers to image data of a person taken while they were alive.
[0007] "Early life voice data" refers to data containing voice information of a person that was recorded while the person was alive.
[0008] "Means for uploading" refers to the ability to send data from a client device to a cloud service.
[0009] "Means for analyzing" refers to the ability to perform a process to extract specific features or information from the received data.
[0010] "Facial features" are data that indicate the unique shape and pattern of a person's face.
[0011] "Voice characteristics" are data that indicate the unique tone, accent, and catchphrases of a person's voice.
[0012] "Means for generating text messages" refers to a function that generates naturally-phrased sentences using input information.
[0013] The "means for converting into a voice message" refers to a function for synthesizing voice based on the generated text message and outputting it as voice data.
[0014] "Means for generating video" refers to the function of outputting moving images using analyzed facial features and generated audio messages.
[0015] "Cloud storage" refers to a storage service for storing data on remote online servers.
[0016] "Streaming format" refers to a technology that plays media such as video and audio in real time.
[0017] "Security Token" refers to authentication information issued to ensure secure access to data.
[0018] "Means for verifying access rights" refers to a function that verifies whether a user has the appropriate rights.
[0019] A "dialect database" refers to a collection of information that stores expressions and pronunciations unique to the language of a particular region.
[0020] A "catchphrase database" refers to a collection of information that stores words and expressions that a particular person frequently uses.
[0021] "Analysis of audio data" refers to the process of extracting and analyzing characteristic information from an audio signal.
[0022] "Means to reflect specific patterns" refers to the function of reproducing specific dialects or catchphrases in speech based on analyzed data. [Brief explanation of the drawings]
[0023] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2]1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0024] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0025] First, the terms used in the following description will be explained.
[0026] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0027] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0028] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0029] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0030] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0031] [First embodiment]
[0032] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0033] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0034] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0035] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0036] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0037] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0038] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0039] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0040] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0041] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0042] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0043] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0044] This invention provides a system for generating video messages based on photographic and audio data of loved ones, and describes an embodiment of the system. By implementing this system, users can recreate the voices and faces of loved ones and receive them as heartwarming messages.
[0045] System configuration
[0046] This system consists of a user terminal, a cloud server, and a storage system. Each component is explained below.
[0047] User's device
[0048] The user's device has the function of uploading photo data and audio data to the cloud server, and also has the function of playing streaming video provided by the cloud server.
[0049] Cloud Server
[0050] The cloud server has the functions of receiving and analyzing photo data and audio data, generating text messages based on the analysis results and converting them into audio messages, and generating videos based on the generated audio messages and facial features, and storing them in cloud storage after ensuring appropriate security.
[0051] Cloud Storage
[0052] Cloud storage is a database that securely stores the generated videos and provides them to users in streaming format. By using security tokens, user access rights are confirmed and secure data management is achieved.
[0053] Program processing
[0054] The program in this system plays three main roles: user, terminal, and server. The process flow and specific operations are explained below.
[0055] First, the user accesses the cloud service and logs in. Next, the user uploads photos and recorded voices of their deceased family members from their device. The device receives this data, converts it into an appropriate format (e.g., JPEG and WAV), and sends it to the cloud server.
[0056] The server then receives the uploaded data and performs image analysis. The image analysis module extracts facial features and generates a high-resolution image. The audio analysis module then analyzes the audio data and extracts voice characteristics, including dialect and speech patterns.
[0057] Next, based on the text message entered by the user, the server uses a generative AI model to generate a naturally-phrased text message, which is then converted into a voice message using a speech synthesis engine, reflecting the analyzed voice characteristics.
[0058] The server then uses the generated voice message and facial features to generate a video that appears to be moving, which is then stored in cloud storage and a security token is issued to protect the data.
[0059] Finally, the user accesses the cloud service again and plays the generated video in streaming format. The server verifies the user's access privileges and then plays the video securely using a security token.
[0060] Specific examples
[0061] 1. The user logs in to the cloud service and uploads a photo of their deceased mother and her last recorded message (e.g., a birthday message) from their device to the cloud server.
[0062] 2. The server analyzes the uploaded photo, extracts the mother's facial features, and generates a high-resolution image. At the same time, it performs audio analysis to extract the mother's tone of voice, accent, and speaking habits.
[0063] 3. The user requests a video of their mother wishing them a happy birthday. Based on this request, the server uses a generative AI model to generate a text message and converts it into a voice message.
[0064] 4. The server generates a video of the mother speaking using the generated voice message and the analyzed high-resolution facial features.
[0065] 5. The server stores the generated video in cloud storage and issues a security token to protect the data.
[0066] 6. The user accesses the cloud service to play the generated video and watch the heartwarming message from their deceased mother.
[0067] The above is a specific description of the embodiment of the "Animated My Album" system of the present invention.
[0068] The processing flow will be explained below.
[0069] Step 1: Log in and upload data
[0070] 1. The user accesses the cloud service and logs in by entering their user ID and password.
[0071] 2. The device sends the login information to the cloud server for authentication.
[0072] 3. The server verifies the authentication information and, if successful, grants access to the user.
[0073] 4. The user selects the photo and audio data of the deceased family member on the device and uploads it to the cloud server.
[0074] 5. The device converts the selected data into the appropriate format (e.g., JPEG and WAV) and sends it to the cloud server.
[0075] 6. The server stores the received photo and audio data in cloud storage.
[0076] Step 2: Data analysis
[0077] 1. The server passes the uploaded photo data to the image analysis module.
[0078] 2. The server uses an image analysis module to extract facial features (such as the position of the eyes, nose, and mouth) of family members from the photo.
[0079] 3. The server uses super-resolution technology to generate a high-resolution facial image based on the extracted feature points.
[0080] 4. The server passes the uploaded voice data to the voice analysis module.
[0081] 5. The server uses a speech analysis module to extract tone, rate, accent, and dialect and diction patterns from the speech data.
[0082] Step 3: Message Generation
[0083] 1. The user enters the content of the message they want to generate (e.g., "Happy Birthday") through the cloud service interface.
[0084] 2. The server passes the input text message to a generative AI model (e.g., GPT-4) to generate a naturally phrased text message.
[0085] 3. The server uses a speech synthesis engine to generate a voice message that sounds like the family member's voice based on the generated text message and the extracted voice characteristics.
[0086] Step 4: Video Generation
[0087] 1. The server uses a high-resolution facial image and the generated voice message to generate a video that matches the facial movements and voice using ravioli generation technology (e.g., Facial Animation Technology).
[0088] 2. The server stores the generated video in cloud storage and applies appropriate encryption processes.
[0089] Step 5: Streaming
[0090] 1. The user clicks the play button to request playback of the video generated from the cloud service.
[0091] 2. The server verifies the user's access rights, issues a security token, and sends it to the user's device.
[0092] 3. The device receives the security token and sends a request to the server again to start streaming the video.
[0093] 4. The server verifies the security token and sends the video data to the device in streaming format.
[0094] 5. The device plays the received video data, allowing the user to view the video.
[0095] The above is a description of the specific operations divided into each processing step.
[0096] Example 1
[0097] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0098] Conventional commemorative video generation systems that use photographs and audio data have difficulty adequately reproducing the face and voice of the deceased as desired by users, and there are also issues with the quality and realism of the generated videos. As a result, it is difficult for users to receive the heartwarming messages of their loved ones in a realistic way. Furthermore, there are insufficient aspects of data protection and security, which increases the risk of personal information being leaked.
[0099] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0100] In this invention, the server includes: means for a user to access digital data storage and upload photo data and audio data; means for converting the uploaded photo data and audio data into an appropriate format; means for transmitting the converted photo data and audio data to a cloud server; means for the cloud server to analyze the received photo data and audio data and extract facial and vocal characteristics of a person; means for generating a text message using a generative AI model based on the extracted facial and vocal characteristics and converting the text message into an audio message; means for generating a video using the generated audio message and the analyzed facial characteristics; and means for saving the generated video in cloud storage and making it playable in streaming format. This allows users to safely and easily generate high-quality, realistic commemorative videos and receive warm messages from their loved ones in real time.
[0101] "User" refers to an individual who uses the system and uploads photo data and audio data.
[0102] "Digital data storage" refers to a system that stores and makes accessible information electronically.
[0103] "Photo data" refers to still images saved in a format such as JPEG.
[0104] "Audio Data" refers to recorded audio information stored in a format such as WAV.
[0105] "Format" refers to a standard or method for organizing data into a specific form.
[0106] "Cloud server" refers to a remote server that stores, processes, and manages data over the Internet.
[0107] "Analysis" refers to the process of examining received data in detail and extracting features and patterns.
[0108] "Facial features of a person" refers to information about the position and shape of the eyes, nose, mouth, etc. extracted from photographic data.
[0109] "Voice characteristics" refers to information such as voice tone, pronunciation habits, etc. extracted from audio data.
[0110] A "generative AI model" refers to an artificial intelligence model that uses machine learning algorithms to generate new data and information.
[0111] "Text Message" means a message expressed as written words or sentences.
[0112] "Voice message" refers to a message that is an audio representation of a text message.
[0113] "Video" refers to a media format that reproduces movement and sound through the combination of multiple still images and audio.
[0114] "Cloud storage" refers to remote storage that stores and manages data via the Internet.
[0115] "Streaming format" refers to a format in which data is played in real time without downloading.
[0116] "Security token" refers to identification information issued to verify access rights and prevent unauthorized access.
[0117] This invention provides a system for generating video messages based on photographic and audio data of loved ones, and describes an embodiment of the system. By implementing this system, users can recreate the voices and faces of loved ones and receive them as heartwarming messages.
[0118] System configuration
[0119] This system consists of a user terminal, a cloud server, and a cloud storage system. Each component is explained below.
[0120] User's device
[0121] The user's device has the function of uploading photo data and audio data to the cloud server, and also has the function of playing streaming video provided by the cloud server.
[0122] Cloud Server
[0123] The cloud server receives and analyzes photo and audio data. It also generates text messages based on the analysis results and converts them into audio messages. It also generates videos based on the generated audio messages and facial features and saves them in cloud storage. A generative AI model is used to generate text messages with natural wording.
[0124] Cloud Storage
[0125] Cloud storage is a database that securely stores the generated videos and provides them to users in streaming format. By using security tokens, user access rights are confirmed and secure data management is achieved.
[0126] What the program does
[0127] The program in this system plays three main roles: user, terminal, and server. The specific operation is explained below.
[0128] First, users access the cloud service and log in. Then, they upload photos and audio recordings of their deceased family members from their devices. The devices receive these data, convert them into appropriate formats (e.g., JPEG and WAV), and send them to the cloud server.
[0129] The server then receives the uploaded data and first performs image analysis. The image analysis module extracts facial features and generates high-resolution images. The voice analysis module then analyzes the audio data and extracts voice characteristics, including dialect and speech patterns.
[0130] Next, based on the text message entered by the user, the server uses a generative AI model to generate a naturally-phrased text message. The text message is then converted into a voice message using a speech synthesis engine, reflecting the analyzed voice characteristics. An example of a prompt is a request such as, "Please generate a message for my mother celebrating her birthday using her voice and characteristics."
[0131] The server then uses the generated voice message and facial features to generate a video that appears to be moving, which is then stored in cloud storage and a security token is issued to protect the data.
[0132] Finally, the user accesses the cloud service again and plays the generated video in streaming format. The server verifies the user's access privileges and then plays the video securely using a security token.
[0133] Specific examples
[0134] The user logs in to the cloud service and uploads a photo of their deceased mother and her last recorded message (e.g., a birthday message) from their device to the cloud server.
[0135] The server analyzes the uploaded photo, extracts the mother's facial features, and generates a high-resolution image, while also performing audio analysis to extract the mother's tone of voice, accent, and speaking habits.
[0136] A user requests a video of their mother wishing them a happy birthday. Based on this request, the server uses a generative AI model to generate a text message and convert it into an audio message.
[0137] The server uses the generated audio message and analyzed high-resolution facial features to generate a video of the mother speaking.
[0138] The server stores the generated video in cloud storage and issues a security token to protect the data.
[0139] The user accesses the cloud service to play the generated video and watch the heartwarming message from their deceased mother.
[0140] The above is a specific description of the embodiment of the "Animated My Album" system of the present invention.
[0141] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0142] Step 1:
[0143] A user logs in to the cloud service and uploads photos and audio data. The inputs include the user's account information, the JPEG photo data to be uploaded, and the WAV audio data. The output is then prepared on the user's device.
[0144] Step 2:
[0145] The device sends the photo data and audio data received from the user to the server. The inputs are uploaded JPEG photo data and WAV audio data. Specifically, the device sends this data to the cloud server using an HTTP POST request, and the cloud server receives the data as output.
[0146] Step 3:
[0147] The server receives the photo data and audio data and begins analyzing it. The input is JPEG photo data and WAV audio data. Specifically, the server uses an image analysis module to extract facial features from the photo data. A face detection algorithm is used for this analysis, and facial features are obtained as the output.
[0148] Step 4:
[0149] The server analyzes the audio data and extracts voice characteristics. The input is audio data in WAV format. Specifically, the server uses acoustic analysis tools to extract voice characteristics such as tone, accent, and catchphrases from the audio data. The voice characteristics are obtained as output.
[0150] Step 5:
[0151] The server converts the text message entered by the user into a natural-sounding text message based on the generative AI model. The inputs include a text message provided by the user and a prompt: "Generate a message to celebrate my mother's birthday using her voice and characteristics." The output is a natural-sounding text message generated by the generative AI model.
[0152] Step 6:
[0153] The server converts the text message into a voice message using a speech synthesis engine. The inputs are the generated text message and the analyzed voice characteristics. In concrete terms, the server converts the text message into voice using the speech synthesis engine, and the output is the voice message.
[0154] Step 7:
[0155] The server generates a video based on the generated voice message and facial features. The inputs are the voice message and facial features. Specifically, it starts the video generation engine, synchronizes the facial movements with the voice, and generates a video that appears to be speaking as output.
[0156] Step 8:
[0157] The server stores the generated video in cloud storage and issues a security token. The input is the generated video data. Specifically, the server encrypts the video data, stores it in cloud storage, and issues a security token to verify the user's access rights. The output is the securely stored video data and a security token.
[0158] Step 9:
[0159] The user accesses the cloud service again and plays the generated video in streaming format. The inputs are the user's account information and a security token. Specifically, the server verifies the security token and grants the user permission to play the video. The output is the video played in streaming format and provided to the user.
[0160] (Application example 1)
[0161] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0162] Current memorial videos and photo albums are limited to still images or simple slideshows, making it difficult for bereaved families to visually and powerfully reproduce the voices and facial movements of their loved ones. Furthermore, privacy and security concerns pose a demand for a reliable system for safely creating, storing, and playing video content. Therefore, a video message generation system that can convey more personal emotions and deepen emotional connections is needed.
[0163] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0164] In this invention, the server is installed on a smartphone or tablet and includes means for generating, saving, and playing back video messages of memories based on the voices and photos of loved ones, means for analyzing uploaded photo and audio data to extract facial and vocal characteristics of people, and means for saving the generated videos in cloud storage and making them playable in streaming format. This makes it possible to generate memorial videos of loved ones and strengthen emotional ties in a safe and reliable manner.
[0165] "Photographic data from before death" refers to digital data of photographs taken of the deceased person before death.
[0166] "Voice data from before death" refers to digital data of the voice of a deceased person recorded before death.
[0167] "Uploading means" refers to a method or device for transmitting data owned by a user to a cloud server via the Internet.
[0168] "Facial features" are characteristics or specific points of a face extracted by facial recognition technology and are used to identify an individual.
[0169] "Voice features" are characteristics of a voice extracted by voice recognition technology and are used to identify an individual's voice.
[0170] A "text message" is character data in a natural language, and is generated as a message from a loved one.
[0171] A "voice message" is data obtained by converting a generated text message into voice.
[0172] The "means for generating moving images" refers to a method or device that converts still images into dynamic moving images based on facial features and voice messages.
[0173] "Cloud storage" is a data storage area on a remote server for storing files on the Internet.
[0174] "Streaming format" is a technology that allows data to be played in real time while being downloaded.
[0175] A "smartphone or tablet" is a type of mobile device that has advanced computing power and internet connectivity.
[0176] "Memory Video Message" is an emotional video message created using the face and voice of the deceased.
[0177] A "security token" is temporary authentication information used to verify access privileges.
[0178] A "dialect and speech database" is a collection of data that records the speech characteristics of a particular region or particular individual.
[0179] This invention provides a system for generating video messages based on photographic and audio data of loved ones, allowing users to receive emotionally rich messages by recreating the voices and faces of loved ones.
[0180] System configuration
[0181] This system consists of a user device, a cloud server, and cloud storage. Each component is explained below.
[0182] User's device
[0183] The user's device is a smartphone or tablet, which has the function of uploading photo data and audio data to a cloud server.
[0184] It also has the ability to play streaming videos provided by cloud servers.
[0185] Cloud Server
[0186] The cloud server has the function of receiving and analyzing the photo data and audio data.
[0187] It also generates text messages based on the analysis results and converts them into voice messages, and generates videos based on the generated voice messages and facial features, saving them to cloud storage.
[0188] The main software used includes Amazon Rekognition for image analysis, Pyttsx3 for speech synthesis, and OpenCV and MoviePy for video generation.
[0189] Cloud Storage
[0190] Cloud storage is a database that safely stores the generated videos and provides them to users in streaming format.
[0191] By using security tokens, user access rights can be confirmed and data management can be ensured safely.
[0192] Program processing
[0193] On the user's device:
[0194] Users access the cloud service, log in, and upload photos and audio data of their precious family members from their devices. At this time, the photo data is converted to JPEG format, and the audio data is converted to WAV format.
[0195] Cloud Server:
[0196] After receiving the uploaded data, an image analysis module extracts facial features and generates a high-resolution image.
[0197] A speech analysis module analyzes the audio data and extracts vocal characteristics, in this case reflecting specific patterns using a database of dialects and accents.
[0198] Based on the text message entered by the user, a generative AI model is used to generate a naturally-phrased text message, which is then converted into a voice message using a speech synthesis engine, incorporating the analyzed voice characteristics.
[0199] Finally, the generated voice message and facial features are used to generate a video that makes the face appear to be moving.
[0200] Cloud Storage:
[0201] The generated video is stored in cloud storage and a security token is issued to protect the data.
[0202] When a user plays a video, the cloud server checks the user's access rights and then provides the video securely in streaming format.
[0203] Specific examples
[0204] Example 1:
[0205] A user logs in to a cloud service and uploads a photo of their deceased mother and her last recorded message (e.g., a birthday message). The cloud server analyzes the received image data using Amazon Rekognition to extract the mother's facial features. Next, the audio data is analyzed using Pyttsx3 to extract the mother's tone and accent. When the user requests a message saying "Happy Birthday," the generative AI model generates a naturally phrased text message, which is then converted into an audio message using Pyttsx3. Finally, a video of the mother speaking is generated using OpenCV and MoviePy and saved to cloud storage. The user then accesses the cloud service again to play the generated video.
[0206] Prompt for the generative AI model:
[0207] prompt:
[0208] "Generate a text message when a user requests a message like 'Happy Birthday.' The generated text should be natural-sounding and emotive."
[0209] The above is a specific description for carrying out the invention.
[0210] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0211] Step 1:
[0212] A user accesses a cloud service using a smartphone or tablet and logs in. The login information is sent to an authentication server, which returns an authentication result. This identifies the user and allows them to use the system.
[0213] Step 2:
[0214] Users upload photos and audio data of their deceased family members. The photo files (JPEG format) and audio files (WAV format) selected by the user are sent via the Internet to a cloud server, which receives the data and formats it appropriately.
[0215] Step 3:
[0216] The cloud server analyzes the received photo data. First, it uses Amazon Rekognition to analyze the image and extract facial features. This includes information such as the position of the eyes, nose, and mouth, as well as facial contours and skin tone. The extracted facial feature data is used in the next processing step.
[0217] Step 4:
[0218] The cloud server analyzes the received voice data using Pyttsx3 to extract voice features, including patterns such as tone, accent, and speaking habits. The extracted voice feature data is then used to generate the voice message.
[0219] Step 5:
[0220] The text message entered by the user is sent to a cloud server, which uses a generative AI model to generate a text message with natural wording. In this process, a more emotionally charged message is generated based on the user's requested phrases.
[0221] Step 6:
[0222] The cloud server converts the generated text message into a voice message using the Pyttsx3 speech synthesis engine, which generates a voice message that reflects the analyzed voice characteristics, ensuring that the voice message sounds as close as possible to the voice of the deceased person that the user is familiar with.
[0223] Step 7:
[0224] The cloud server generates a video using the generated voice message and the extracted facial features. Using tools such as OpenCV and MoviePy, the video is created so that the face appears to be moving. The video is designed to play back the message entered by the user emotionally.
[0225] Step 8:
[0226] The cloud server saves the generated video in cloud storage and issues a security token to protect the data. This token is used to verify the user's access rights when playing the video. Storing the video in cloud storage ensures secure data management.
[0227] Step 9:
[0228] The user then accesses the cloud service again and plays the generated video in streaming format. The video data is provided from the cloud storage and played in real time on the user's device. At this time, the user's access authority is confirmed using a security token.
[0229] This series of processing steps generates an emotional video message that reproduces the voices and faces of loved ones, allowing users to watch it safely.
[0230] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0231] This invention relates to a system that generates video messages based on photographic and audio data from the deceased's life, and also describes an embodiment that combines an emotion engine. This system allows users to recreate the voice and face of a loved one and receive a heartfelt message that reflects their emotions.
[0232] System configuration
[0233] This system consists of a user's device, a cloud server, cloud storage, and an emotion engine. Each component will be explained below.
[0234] User's device
[0235] The user's device has the function of uploading photo data and audio data to the cloud server, the function of playing streaming videos provided by the cloud server, and the function of collecting the user's emotional data and sending it to the cloud server.
[0236] Cloud Server
[0237] The cloud server receives and analyzes photo and audio data. It also generates text messages based on the analysis results and converts them into audio messages. It also uses an emotion engine to recognize the user's emotions, adjusts the message content and tone as necessary, and generates videos.
[0238] Cloud Storage
[0239] Cloud storage is a database that securely stores the generated videos and provides them to users in streaming format. By using security tokens, user access rights are confirmed and secure data management is achieved.
[0240] Emotion Engine
[0241] The emotion engine analyzes emotional data sent from the user's device and recognizes the user's emotional state in real time. Based on this information, it adjusts the tone and content of generated messages, audio, and video.
[0242] Program processing
[0243] The program in this system plays four main roles: user, terminal, server, and emotion engine. The processing flow and specific operations are explained below.
[0244] First, the user accesses the cloud service and logs in. Next, the user uploads photos and recorded voices of their deceased family members from their device. The device converts these data into appropriate formats (e.g., JPEG and WAV) and sends them to the cloud server.
[0245] Once the server receives the uploaded data, it first performs image analysis. The image analysis module extracts facial features and generates a high-resolution image. Then, the voice analysis module analyzes the audio data and extracts voice characteristics, including dialect and speech patterns.
[0246] Next, based on the text message entered by the user, the server uses a generative AI model to generate a naturally-phrased text message, which is then converted into a voice message using a speech synthesis engine, reflecting the analyzed voice characteristics.
[0247] Furthermore, the emotion engine receives emotion data sent from the user's device and analyzes the user's emotional state. Based on the results, the generated message content and voice tone are adjusted. Specifically, different messages and tones are adopted depending on changes in emotion.
[0248] The server then uses the generated voice message and facial features to generate a video that appears to be moving, which is then stored in cloud storage and a security token is issued to protect the data.
[0249] Finally, the user accesses the cloud service again and plays the generated video in streaming format. The server verifies the user's access privileges and then plays the video securely using a security token.
[0250] Specific examples
[0251] 1. The user logs in to the cloud service and uploads a photo of their deceased mother and her last recorded message (e.g., a birthday message) from their device to the cloud server.
[0252] 2. The server analyzes the uploaded photo, extracts the mother's facial features, and generates a high-resolution image. At the same time, it performs audio analysis to extract the mother's tone of voice, accent, and speaking habits.
[0253] 3. The user requests a video of their mother wishing them a happy birthday. Based on this request, the server uses a generative AI model to generate a text message and converts it into a voice message.
[0254] 4. At this time, the emotion engine analyzes the user's emotions and, for example, if it recognizes that the user is emotional, it adjusts the voice tone to a gentler one.
[0255] 5. The server generates a video of the mother speaking using the generated voice message and the analyzed high-resolution facial features.
[0256] 6. The server stores the generated video in cloud storage and issues a security token to protect the data.
[0257] 7. The user accesses the cloud service to play the generated video and watch the heartwarming message from their deceased mother.
[0258] The above is a specific description of an embodiment of the "My Album in Action" system that combines the emotion engine of the present invention.
[0259] The processing flow will be explained below.
[0260] Step 1: Log in and upload data
[0261] 1. The user accesses the cloud service and logs in by entering their user ID and password.
[0262] 2. The device sends the login information to the cloud server for authentication.
[0263] 3. The server verifies the authentication information and, if successful, grants access to the user.
[0264] 4. The user selects the photo and audio data of the deceased family member on the device and uploads it to the cloud server.
[0265] 5. The device converts the selected data into the appropriate format (e.g., JPEG and WAV) and sends it to the cloud server.
[0266] 6. The server stores the received photo and audio data in cloud storage.
[0267] Step 2: Data analysis
[0268] 1. The server passes the uploaded photo data to the image analysis module.
[0269] 2. The server uses an image analysis module to extract facial features (such as the position of the eyes, nose, and mouth) of family members from the photo.
[0270] 3. The server uses super-resolution technology to generate a high-resolution facial image based on the extracted feature points.
[0271] 4. The server passes the uploaded voice data to the voice analysis module.
[0272] 5. The server uses a speech analysis module to extract tone, rate, accent, and dialect and diction patterns from the speech data.
[0273] Step 3: Collect and analyze emotion data
[0274] 1. The user interacts with the cloud service interface and gives permission to collect emotion data.
[0275] 2. The device collects emotional data such as the user's facial expressions and tone of voice and sends it to a cloud server.
[0276] 3. The server passes the received emotion data to the emotion engine and analyzes the user's emotional state.
[0277] Step 4: Message Generation
[0278] 1. The user enters the content of the message they want to generate (e.g., "Happy Birthday") through the cloud service interface.
[0279] 2. The server passes the input text message to a generative AI model (e.g., GPT-4) to generate a naturally phrased text message.
[0280] 3. The server uses a speech synthesis engine to generate a voice message that sounds like the family member's voice based on the generated text message and the extracted voice characteristics.
[0281] 4. The server adjusts the tone of the voice and the content of the message based on the user's emotional data provided by the emotion engine.
[0282] Step 5: Video Generation
[0283] 1. The server uses a high-resolution facial image and the generated voice message to generate a video that matches the facial movements and voice using ravioli generation technology (e.g., Facial Animation Technology).
[0284] 2. The server stores the generated video in cloud storage and applies appropriate encryption processes.
[0285] Step 6: Streaming
[0286] 1. The user clicks the play button to request playback of the video generated from the cloud service.
[0287] 2. The server verifies the user's access rights, issues a security token, and sends it to the user's device.
[0288] 3. The device receives the security token and sends a request to the server again to start streaming the video.
[0289] 4. The server verifies the security token and sends the video data to the device in streaming format.
[0290] 5. The device plays the received video data, allowing the user to watch the video.
[0291] The above is a description of the specific operations divided into processing steps of the "Moving My Album" system.
[0292] Example 2
[0293] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0294] In recent years, there has been a demand for preserving memories of deceased loved ones and receiving them as emotionally warm messages. However, conventional technologies have struggled to reproduce the voice and appearance of the deceased in a natural way. It has also been difficult to generate messages and videos that correspond to the user's emotional state. Furthermore, ensuring the security of the generated data and managing access rights are also important issues.
[0295] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0296] In this invention, the server includes means for acquiring image data of the deceased person, means for acquiring audio data of the deceased person, means for analyzing the acquired image data and audio data and extracting facial and vocal characteristics of the person, means for generating a natural language message based on the extracted facial and vocal characteristics and converting the message into synthetic speech, means for analyzing the user's emotional data and reflecting the emotional state in the generation process, means for generating a video using the generated synthetic speech and the analyzed facial characteristics, and means for saving the generated video in a storage device and making it playable in streaming format. This makes it possible to reproduce the voice and appearance of the deceased person with high accuracy, and to easily generate and safely manage warm video messages that correspond to the user's emotional state.
[0297] "Image data from before death" refers to photographs and image files taken of a person before their death.
[0298] "Pre-death acoustic data" refers to voices and audio files recorded during a person's lifetime.
[0299] "Means of acquisition" refers to the functions and interfaces that allow users to select and upload data.
[0300] "Means for analysis" refers to the algorithms and software used to analyze the acquired data and extract the necessary information.
[0301] "Facial features" refer to identifiable parts or attributes of a person's face that are extracted from an image.
[0302] "Voice features" refer to identifiable parts or attributes such as tone, accent, and catchphrases extracted from voice data.
[0303] A "natural language message" refers to a text message that is written in a way that is natural for humans to understand.
[0304] "Synthetic voice" refers to data generated as voice based on a text message.
[0305] "Emotional Data" refers to information collected to describe a user's emotional state.
[0306] The "generation process" refers to a series of steps to generate messages and videos based on analyzed data.
[0307] "Storage device" refers to a storage system for safely storing the generated video.
[0308] "Streaming format" refers to a format that allows users to watch videos in real time over the Internet.
[0309] This invention relates to a system that generates video messages based on image and audio data from a person's life, and by combining it with an emotion engine, it allows users to receive the messages as warm messages. This system consists of a user's device, a cloud server, cloud storage, and an emotion engine.
[0310] System configuration
[0311] User's device
[0312] The user's device has the function of uploading image data and audio data to the cloud server, the function of playing streaming videos provided by the cloud server, and the function of collecting the user's emotional data and sending it to the cloud server.
[0313] Cloud Server
[0314] The cloud server is equipped with software and hardware for receiving and analyzing image and acoustic data. The analysis uses OpenCV, an image processing software, and NVIDIA NeMo, a voice analysis software. Based on the text message entered by the user, a generative AI model (e.g., a generative AI model) is used to generate a text message with natural wording.
[0315] The server converts the generated text message into a voice message using a speech synthesis engine, incorporating the analyzed voice characteristics. This process uses the speech synthesis technology of Azure Cognitive Services.
[0316] Furthermore, the Emotion Analysis Engine analyzes the user's emotional data and adjusts the generated message content and voice tone based on the user's emotional state. For example, if the user is emotional, the voice tone will be made gentler.
[0317] The server uses the generated voice message and high-resolution facial features to generate a video that appears to show a moving face using DeepFaceLab. The video is encoded using FFmpeg. The video is then stored in cloud storage, where data is protected by a security token.
[0318] Cloud Storage
[0319] Cloud storage is a database that securely stores the generated videos and provides them to users in streaming format. By using security tokens, user access rights are confirmed and secure data management is achieved.
[0320] Emotion Engine
[0321] The emotion engine analyzes emotional data sent from the user's device and recognizes the user's emotional state in real time. Based on this information, it adjusts the tone and content of generated messages, audio, and video.
[0322] Specific examples
[0323] A user logs in to the cloud service and uploads a photo (JPEG format) of a deceased family member and audio data (WAV format) from their device to the cloud server. The server analyzes the uploaded data using OpenCV and NVIDIA NeMo to extract facial and vocal characteristics. Next, the generative AI model converts the user's requested message (e.g., "Happy Birthday") into natural-sounding text, which is then converted into a voice message using Azure Cognitive Services' speech synthesis engine. The emotion engine also analyzes the user's emotional data and adjusts the message content and voice tone. Finally, the server generates a video using DeepFaceLab and saves it in cloud storage. The user can then access the cloud service again to play the generated video in streaming format.
[0324] Example prompt: "Generate a sweet message from my late mother wishing me a happy birthday."
[0325] The above is a specific description of an embodiment of the "My Album in Action" system that combines the emotion engine of the present invention.
[0326] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0327] Step 1:
[0328] A user logs in
[0329] A user accesses a cloud service using a browser or application and enters a user ID and password, which authenticates the user and allows them to log in to the cloud service.
[0330] Input: User ID, Password
[0331] Output: Authentication token, indicating successful login
[0332] Step 2:
[0333] The user uploads photos and audio data.
[0334] The user selects a photo (JPEG format) and audio data (WAV format) of the deceased family member from the device and clicks the upload button. The device converts the data into the appropriate format (JPEG and WAV) and sends it to the cloud server.
[0335] Input: Photo data (JPEG), audio data (WAV)
[0336] Output: Photo and audio data uploaded to the cloud server
[0337] Step 3:
[0338] The server receives the photo and audio data.
[0339] The server receives the photo and audio data sent from the device and temporarily stores it in the specified storage. At this stage, the server checks the integrity of the data to ensure there is no unauthorized data.
[0340] Input: Photo data and audio data uploaded from the device
[0341] Output: Temporarily saved photo and audio data
[0342] Step 4:
[0343] The server performs image analysis
[0344] The server analyzes the received photo data using OpenCV, specifically applying a facial recognition algorithm to extract facial features.
[0345] Input: Temporarily saved photo data
[0346] Output: Facial features (position, shape, landmark points, etc.)
[0347] Step 5:
[0348] The server performs voice analysis
[0349] The server uses NVIDIA NeMo to analyze the received voice data, extracting voice tone, accent, and speaking habits.
[0350] Input: Temporarily saved audio data
[0351] Output: Voice characteristics (tone, accent, speech patterns)
[0352] Step 6:
[0353] The server generates a text message
[0354] Based on a prompt phrase entered by the user (e.g., "Say happy birthday"), the server uses a generative AI model (e.g., GPT-3) to generate a natural-sounding text message.
[0355] Input: The prompt text entered by the user
[0356] Output: The generated text message
[0357] Step 7:
[0358] The server converts it into a voice message
[0359] Based on the generated text message, the server uses Azure Cognitive Services' speech synthesis engine to convert it into a voice message, which reflects the analyzed voice characteristics.
[0360] Input: Generated text message, voice features
[0361] Output: Synthesized voice message
[0362] Step 8:
[0363] The emotion engine analyzes the emotion data
[0364] The emotion engine analyzes emotion data sent from the user's device and recognizes the user's emotional state in real time.
[0365] Input: User emotion data (facial expressions, tone of voice, etc.)
[0366] Output: Emotional state analysis result
[0367] Step 9:
[0368] The server generates a video based on the generated data.
[0369] The server uses DeepFaceLab to generate a video that appears to show a moving face based on the generated voice message and high-resolution facial features, and encodes the video using FFmpeg.
[0370] Input: Synthesized voice message, facial features
[0371] Output: Generated video file
[0372] Step 10:
[0373] The server saves the video to cloud storage.
[0374] The generated video is stored in cloud storage. A security token is issued to verify the user's access rights and protect the data.
[0375] Input: Generated video file
[0376] Output: Video stored in cloud storage, security token
[0377] Step 11:
[0378] Playing user-generated videos
[0379] The user then accesses the cloud service again and plays the video. The server then verifies the user's access privileges and uses the security token to play the video securely in streaming format.
[0380] Input: Security token, saved video file
[0381] Output: Streamed video
[0382] The above are the detailed processing steps of this system.
[0383] (Application example 2)
[0384] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0385] In order to recreate the memory of the deceased and move the bereaved family, it is desirable to provide more realistic and emotionally relevant video messages. However, current technology has problems in that it is difficult to adjust the tone and content of the message according to the user's emotional state, and the generated videos lack a sense of realism. In addition, there is a lack of appropriate display methods for viewing moving videos in real time. There is a need to solve these problems.
[0386] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0387] In this invention, the server includes means for uploading photo data of the person before death, means for uploading audio data of the person before death, means for analyzing the uploaded photo data and audio data and extracting facial and vocal characteristics of the person, means for generating a text message based on the extracted facial and vocal characteristics and converting the text message into an audio message, means for generating a video using the generated audio message and the analyzed facial characteristics, means for storing the generated video in cloud storage and making it playable in streaming format, means for collecting emotional data and analyzing the data to adjust the tone of the text message or audio message, and means for displaying the generated video using a head-mounted display, thereby making it possible to provide a realistic video message that matches the emotional state of the user.
[0388] "Photo data from before death" refers to image files of a person taken before death.
[0389] "Voice data from before death" refers to an audio file of a person's voice recorded before death.
[0390] The "means for uploading" refers to a method and apparatus for transmitting data from a terminal to a cloud server.
[0391] "Means for analyzing" refers to methods and devices that process uploaded data to extract information.
[0392] "Facial features" are characteristic parts of a person's face that are extracted from the analyzed image data.
[0393] "Voice features" are characteristics of a person's voice extracted from analyzed voice data.
[0394] A "means for generating a text message" is a method and device that creates a sentence based on input information.
[0395] A "means for converting to voice message" is a method and apparatus for converting a text message to voice data.
[0396] The "means for generating video" refers to a method and device for creating a video file based on image data and audio data.
[0397] "Cloud storage means" refers to a method and device for storing data on a remote server.
[0398] "Means for enabling playback in streaming format" refers to a method and apparatus for playing back stored video data in real time via the Internet.
[0399] A "means for collecting emotional data" is a method and apparatus for collecting a user's emotional state.
[0400] A "means for analyzing emotion data" is a method and apparatus for processing collected emotion data to make sense of the information.
[0401] "Means for adjusting the tone of text and voice messages" refers to a method and apparatus for changing the tone of message content based on analyzed emotional data.
[0402] A "head-mounted display" is a display device worn on the head, and is used to display videos and images.
[0403] This invention relates to a system for generating moving video messages using photographic and audio data from a person's lifetime. Furthermore, by collecting and analyzing the user's emotional data, the system adjusts the message content and tone to generate more personalized and emotional videos. A specific embodiment of this system is described below.
[0404] System configuration
[0405] This system consists of a user's device, a cloud server, cloud storage, an emotion engine, and a head-mounted display.
[0406] User's device
[0407] The user's device has the function of uploading photo data and audio data to the cloud server, the function of playing streaming videos provided by the cloud server, and the function of collecting the user's emotional data and sending it to the cloud server.
[0408] Cloud Server
[0409] The cloud server receives and analyzes photo and audio data. It also generates text messages based on the analysis results and converts them into audio messages. It also uses an emotion engine to recognize the user's emotions, adjusts the message content and tone as necessary, and generates videos.
[0410] Cloud Storage
[0411] Cloud storage is a database that securely stores the generated videos and provides them to users in streaming format. By using security tokens, user access rights are confirmed and secure data management is achieved.
[0412] Emotion Engine
[0413] The emotion engine analyzes emotional data sent from the user's device and recognizes the user's emotional state in real time. Based on this information, it adjusts the tone and content of generated messages, audio, and video.
[0414] head-mounted display
[0415] A head-mounted display is a device worn on the head that displays images to the user. It displays video received from a cloud server in real time, providing the user with a realistic viewing experience.
[0416] Program processing
[0417] The program in this system mainly plays four main roles: user, terminal, server, and emotion engine.
[0418] First, the user accesses the cloud service and logs in. The user uploads photos and audio data from their device to the cloud server. The device converts the data into an appropriate format (e.g., JPEG and WAV) and sends it to the cloud server.
[0419] Once the cloud server receives the uploaded data, it first performs image analysis. The image analysis module extracts facial features and generates a high-resolution image. The voice analysis module then analyzes the audio data and extracts voice characteristics, including dialect and speech patterns.
[0420] Next, based on the text message entered by the user, the cloud server uses a generative AI model to generate a naturally-phrased text message, which is then converted into a voice message using a speech synthesis engine, reflecting the analyzed voice characteristics.
[0421] Furthermore, the emotion engine receives emotion data sent from the user's device and analyzes the user's emotional state. Based on the results, it adjusts the generated message content and voice tone. Specifically, it adopts different messages and tones depending on changes in the user's emotions.
[0422] The cloud server then uses the generated voice message and facial features to generate a video that appears to be moving, which is then stored in cloud storage and a security token is issued to protect the data.
[0423] Finally, the user wears a head-mounted display and watches the generated video. The cloud server verifies the user's access rights and then plays the video securely using a security token.
[0424] Specific examples
[0425] A user visits a funeral home and uploads a photo of their deceased father and a recorded voice message (e.g., a birthday message). The cloud server analyzes the uploaded photo, extracts the father's facial features, and generates a high-resolution image. At the same time, it performs voice analysis to extract the father's tone of voice, accent, and catchphrases.
[0426] A user can request a generative AI model to generate a video of their father saying, "Take care." The emotion engine then analyzes the user's emotions and adjusts the tone of the voice to be gentler if it detects that the user is emotional, for example.
[0427] The cloud server then uses the generated voice message and analyzed high-resolution facial features to generate a video of the father speaking, which is then stored in cloud storage and a security token is issued to protect the data.
[0428] Finally, the user can wear a head-mounted display to watch the generated video and receive a heartwarming message from their deceased father.
[0429] Example prompts for generative AI models
[0430] Below is an example of a prompt sentence for the generative AI model (generative AI model).
[0431] "Dear visitors, wipe away your tears and upload your favorite photos and audio recordings here. We will recreate the voice and face of your loved one and deliver a heartwarming video message. Let's share this special moment together. We will recreate phrases such as 'Thank you' and 'Take care' in the video."
[0432] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0433] Step 1:
[0434] The user accesses the cloud service and logs in.
[0435] Input: User authentication information (user ID, password)
[0436] Output: Authentication result (success / failure)
[0437] Specific behavior:
[0438] The user's device sends authentication information to the cloud server, which checks it against the database and, if authentication is successful, proceeds to the next step.
[0439] Step 2:
[0440] Users upload photo data and audio data from their pre-death lives from their devices to a cloud server.
[0441] Input: Photo data from before death (JPEG), audio data from before death (WAV)
[0442] Output: Data stored on the cloud server
[0443] Specific behavior:
[0444] Users simply select photos and audio data from their device and click the upload button. The device then sends the data to the cloud server, which stores it securely.
[0445] Step 3:
[0446] The cloud server analyzes the uploaded photo data.
[0447] Input: Photo data from before death (JPEG)
[0448] Output: Facial feature data (facial feature points, facial expressions)
[0449] Specific behavior:
[0450] The server uses an image analysis module (e.g., OpenCV) to extract facial features from the photo data and generate high-resolution facial images.
[0451] Step 4:
[0452] The cloud server analyzes the voice data from the person's life.
[0453] Input: Audio data from before death (WAV)
[0454] Output: Voice characteristics data (tone of voice, accent, catchphrases)
[0455] Specific behavior:
[0456] The server uses a voice analysis module (e.g., librosa) to extract voice features from the audio data.
[0457] Step 5:
[0458] The cloud server generates a message using a generative AI model based on the text message entered by the user.
[0459] Input: User-entered text messages, facial feature data, and voice feature data
[0460] Output: Naturally-phrased text message
[0461] Specific behavior:
[0462] The server sends a prompt to a generative AI model (e.g., GPT-3) to generate a naturally phrased text message.
[0463] Step 6:
[0464] The cloud server converts the generated text message into a voice message.
[0465] Input: Naturally-phrased text messages, voice feature data
[0466] Output: Voice message (audio file)
[0467] Specific behavior:
[0468] The server uses a speech synthesis engine (e.g., Watson Text to Speech) to convert the text message into speech, and the generated voice message reflects the voice characteristics.
[0469] Step 7:
[0470] The emotion engine analyzes the user's emotion data and transmits the emotional state to the cloud server.
[0471] Input: User emotion data (real-time emotional state)
[0472] Output: Emotion analysis result (emotional state)
[0473] Specific behavior:
[0474] The emotion engine analyzes the emotion data collected from the user's device and sends the results to a cloud server.
[0475] Step 8:
[0476] The cloud server adjusts the message content and voice tone based on the results of emotion analysis.
[0477] Input: Emotion analysis results, voice message, facial feature data
[0478] Output: Adjusted message content, audio tone
[0479] Specific behavior:
[0480] The results of the emotion engine are reflected and the tone of the generated text and voice messages is adjusted to match the user's emotions.
[0481] Step 9:
[0482] The cloud server generates a video using the generated voice message and facial features.
[0483] Input: Facial feature data, voice message
[0484] Output: Video file
[0485] Specific behavior:
[0486] The server uses a video generation module (for example, Adobe After Effects API) to generate a video in which a person's face moves based on the voice message and facial feature data.
[0487] Step 10:
[0488] The cloud server stores the generated video in cloud storage and issues a security token to protect the data.
[0489] Input: Video file
[0490] Output: Saved video file, security token
[0491] Specific behavior:
[0492] The server stores the video file in cloud storage (e.g., AWS S3) and issues a security token for access control.
[0493] Step 11:
[0494] The user wears a head-mounted display and watches videos generated from a cloud server.
[0495] Input: Security Token
[0496] Output: Video Streaming
[0497] Specific behavior:
[0498] The user accesses the cloud service again and watches the video in real time using a head-mounted display (e.g., HoloLens, Oculus Rift). The cloud server uses a security token to verify the user's access privileges and plays the video securely.
[0499] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0500] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0501] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0502] [Second embodiment]
[0503] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0504] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0505] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0506] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0507] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0508] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0509] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0510] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0511] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0512] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0513] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0514] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0515] This invention provides a system for generating video messages based on photographic and audio data of loved ones, and describes an embodiment of the system. By implementing this system, users can recreate the voices and faces of loved ones and receive them as heartwarming messages.
[0516] System configuration
[0517] This system consists of a user terminal, a cloud server, and a storage system. Each component is explained below.
[0518] User's device
[0519] The user's device has the function of uploading photo data and audio data to the cloud server, and also has the function of playing streaming video provided by the cloud server.
[0520] Cloud Server
[0521] The cloud server has the functions of receiving and analyzing photo data and audio data, generating text messages based on the analysis results and converting them into audio messages, and generating videos based on the generated audio messages and facial features, and storing them in cloud storage after ensuring appropriate security.
[0522] Cloud Storage
[0523] Cloud storage is a database that securely stores the generated videos and provides them to users in streaming format. By using security tokens, user access rights are confirmed and secure data management is achieved.
[0524] Program processing
[0525] The program in this system plays three main roles: user, terminal, and server. The process flow and specific operations are explained below.
[0526] First, the user accesses the cloud service and logs in. Next, the user uploads photos and recorded voices of their deceased family members from their device. The device receives this data, converts it into an appropriate format (e.g., JPEG and WAV), and sends it to the cloud server.
[0527] The server then receives the uploaded data and performs image analysis. The image analysis module extracts facial features and generates a high-resolution image. The audio analysis module then analyzes the audio data and extracts voice characteristics, including dialect and speech patterns.
[0528] Next, based on the text message entered by the user, the server uses a generative AI model to generate a naturally-phrased text message, which is then converted into a voice message using a speech synthesis engine, reflecting the analyzed voice characteristics.
[0529] The server then uses the generated voice message and facial features to generate a video that appears to be moving, which is then stored in cloud storage and a security token is issued to protect the data.
[0530] Finally, the user accesses the cloud service again and plays the generated video in streaming format. The server verifies the user's access privileges and then plays the video securely using a security token.
[0531] Specific examples
[0532] 1. The user logs in to the cloud service and uploads a photo of their deceased mother and her last recorded message (e.g., a birthday message) from their device to the cloud server.
[0533] 2. The server analyzes the uploaded photo, extracts the mother's facial features, and generates a high-resolution image. At the same time, it performs audio analysis to extract the mother's tone of voice, accent, and speaking habits.
[0534] 3. The user requests a video of their mother wishing them a happy birthday. Based on this request, the server uses a generative AI model to generate a text message and converts it into a voice message.
[0535] 4. The server generates a video of the mother speaking using the generated voice message and the analyzed high-resolution facial features.
[0536] 5. The server stores the generated video in cloud storage and issues a security token to protect the data.
[0537] 6. The user accesses the cloud service to play the generated video and watch the heartwarming message from their deceased mother.
[0538] The above is a specific description of the embodiment of the "Animated My Album" system of the present invention.
[0539] The processing flow will be explained below.
[0540] Step 1: Log in and upload data
[0541] 1. The user accesses the cloud service and logs in by entering their user ID and password.
[0542] 2. The device sends the login information to the cloud server for authentication.
[0543] 3. The server verifies the authentication information and, if successful, grants access to the user.
[0544] 4. The user selects the photo and audio data of the deceased family member on the device and uploads it to the cloud server.
[0545] 5. The device converts the selected data into the appropriate format (e.g., JPEG and WAV) and sends it to the cloud server.
[0546] 6. The server stores the received photo and audio data in cloud storage.
[0547] Step 2: Data analysis
[0548] 1. The server passes the uploaded photo data to the image analysis module.
[0549] 2. The server uses an image analysis module to extract facial features (such as the position of the eyes, nose, and mouth) of family members from the photo.
[0550] 3. The server uses super-resolution technology to generate a high-resolution facial image based on the extracted feature points.
[0551] 4. The server passes the uploaded voice data to the voice analysis module.
[0552] 5. The server uses a speech analysis module to extract tone, rate, accent, and dialect and diction patterns from the speech data.
[0553] Step 3: Message Generation
[0554] 1. The user enters the content of the message they want to generate (e.g., "Happy Birthday") through the cloud service interface.
[0555] 2. The server passes the input text message to a generative AI model (e.g., GPT-4) to generate a naturally phrased text message.
[0556] 3. The server uses a speech synthesis engine to generate a voice message that sounds like the family member's voice based on the generated text message and the extracted voice characteristics.
[0557] Step 4: Video Generation
[0558] 1. The server uses a high-resolution facial image and the generated voice message to generate a video that matches the facial movements and voice using ravioli generation technology (e.g., Facial Animation Technology).
[0559] 2. The server stores the generated video in cloud storage and applies appropriate encryption processes.
[0560] Step 5: Streaming
[0561] 1. The user clicks the play button to request playback of the video generated from the cloud service.
[0562] 2. The server verifies the user's access rights, issues a security token, and sends it to the user's device.
[0563] 3. The device receives the security token and sends a request to the server again to start streaming the video.
[0564] 4. The server verifies the security token and sends the video data to the device in streaming format.
[0565] 5. The device plays the received video data, allowing the user to view the video.
[0566] The above is a description of the specific operations divided into each processing step.
[0567] Example 1
[0568] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0569] Conventional commemorative video generation systems that use photographs and audio data have difficulty adequately reproducing the face and voice of the deceased as desired by users, and there are also issues with the quality and realism of the generated videos. As a result, it is difficult for users to receive the heartwarming messages of their loved ones in a realistic way. Furthermore, there are insufficient aspects of data protection and security, which increases the risk of personal information being leaked.
[0570] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0571] In this invention, the server includes: means for a user to access digital data storage and upload photo data and audio data; means for converting the uploaded photo data and audio data into an appropriate format; means for transmitting the converted photo data and audio data to a cloud server; means for the cloud server to analyze the received photo data and audio data and extract facial and vocal characteristics of a person; means for generating a text message using a generative AI model based on the extracted facial and vocal characteristics and converting the text message into an audio message; means for generating a video using the generated audio message and the analyzed facial characteristics; and means for saving the generated video in cloud storage and making it playable in streaming format. This allows users to safely and easily generate high-quality, realistic commemorative videos and receive warm messages from their loved ones in real time.
[0572] "User" refers to an individual who uses the system and uploads photo data and audio data.
[0573] "Digital data storage" refers to a system that stores and makes accessible information electronically.
[0574] "Photo data" refers to still images saved in a format such as JPEG.
[0575] "Audio Data" refers to recorded audio information stored in a format such as WAV.
[0576] "Format" refers to a standard or method for organizing data into a specific form.
[0577] "Cloud server" refers to a remote server that stores, processes, and manages data over the Internet.
[0578] "Analysis" refers to the process of examining received data in detail and extracting features and patterns.
[0579] "Facial features of a person" refers to information about the position and shape of the eyes, nose, mouth, etc. extracted from photographic data.
[0580] "Voice characteristics" refers to information such as voice tone, pronunciation habits, etc. extracted from audio data.
[0581] A "generative AI model" refers to an artificial intelligence model that uses machine learning algorithms to generate new data and information.
[0582] "Text Message" means a message expressed as written words or sentences.
[0583] "Voice message" refers to a message that is an audio representation of a text message.
[0584] "Video" refers to a media format that reproduces movement and sound through the combination of multiple still images and audio.
[0585] "Cloud storage" refers to remote storage that stores and manages data via the Internet.
[0586] "Streaming format" refers to a format in which data is played in real time without downloading.
[0587] "Security token" refers to identification information issued to verify access rights and prevent unauthorized access.
[0588] This invention provides a system for generating video messages based on photographic and audio data of loved ones, and describes an embodiment of the system. By implementing this system, users can recreate the voices and faces of loved ones and receive them as heartwarming messages.
[0589] System configuration
[0590] This system consists of a user terminal, a cloud server, and a cloud storage system. Each component is explained below.
[0591] User's device
[0592] The user's device has the function of uploading photo data and audio data to the cloud server, and also has the function of playing streaming video provided by the cloud server.
[0593] Cloud Server
[0594] The cloud server receives and analyzes photo and audio data. It also generates text messages based on the analysis results and converts them into audio messages. It also generates videos based on the generated audio messages and facial features and saves them in cloud storage. A generative AI model is used to generate text messages with natural wording.
[0595] Cloud Storage
[0596] Cloud storage is a database that securely stores the generated videos and provides them to users in streaming format. By using security tokens, user access rights are confirmed and secure data management is achieved.
[0597] What the program does
[0598] The program in this system plays three main roles: user, terminal, and server. The specific operation is explained below.
[0599] First, users access the cloud service and log in. Then, they upload photos and audio recordings of their deceased family members from their devices. The devices receive these data, convert them into appropriate formats (e.g., JPEG and WAV), and send them to the cloud server.
[0600] The server then receives the uploaded data and first performs image analysis. The image analysis module extracts facial features and generates high-resolution images. The voice analysis module then analyzes the audio data and extracts voice characteristics, including dialect and speech patterns.
[0601] Next, based on the text message entered by the user, the server uses a generative AI model to generate a naturally-phrased text message. The text message is then converted into a voice message using a speech synthesis engine, reflecting the analyzed voice characteristics. An example of a prompt is a request such as, "Please generate a message for my mother celebrating her birthday using her voice and characteristics."
[0602] The server then uses the generated voice message and facial features to generate a video that appears to be moving, which is then stored in cloud storage and a security token is issued to protect the data.
[0603] Finally, the user accesses the cloud service again and plays the generated video in streaming format. The server verifies the user's access privileges and then plays the video securely using a security token.
[0604] Specific examples
[0605] The user logs in to the cloud service and uploads a photo of their deceased mother and her last recorded message (e.g., a birthday message) from their device to the cloud server.
[0606] The server analyzes the uploaded photo, extracts the mother's facial features, and generates a high-resolution image, while also performing audio analysis to extract the mother's tone of voice, accent, and speaking habits.
[0607] A user requests a video of their mother wishing them a happy birthday. Based on this request, the server uses a generative AI model to generate a text message and convert it into an audio message.
[0608] The server uses the generated audio message and analyzed high-resolution facial features to generate a video of the mother speaking.
[0609] The server stores the generated video in cloud storage and issues a security token to protect the data.
[0610] The user accesses the cloud service to play the generated video and watch the heartwarming message from their deceased mother.
[0611] The above is a specific description of the embodiment of the "Animated My Album" system of the present invention.
[0612] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0613] Step 1:
[0614] A user logs in to the cloud service and uploads photos and audio data. The inputs include the user's account information, the JPEG photo data to be uploaded, and the WAV audio data. The output is then prepared on the user's device.
[0615] Step 2:
[0616] The device sends the photo data and audio data received from the user to the server. The inputs are uploaded JPEG photo data and WAV audio data. Specifically, the device sends this data to the cloud server using an HTTP POST request, and the cloud server receives the data as output.
[0617] Step 3:
[0618] The server receives the photo data and audio data and begins analyzing it. The input is JPEG photo data and WAV audio data. Specifically, the server uses an image analysis module to extract facial features from the photo data. A face detection algorithm is used for this analysis, and facial features are obtained as the output.
[0619] Step 4:
[0620] The server analyzes the audio data and extracts voice characteristics. The input is audio data in WAV format. Specifically, the server uses acoustic analysis tools to extract voice characteristics such as tone, accent, and catchphrases from the audio data. The voice characteristics are obtained as output.
[0621] Step 5:
[0622] The server converts the text message entered by the user into a natural-sounding text message based on the generative AI model. The inputs include a text message provided by the user and a prompt: "Generate a message to celebrate my mother's birthday using her voice and characteristics." The output is a natural-sounding text message generated by the generative AI model.
[0623] Step 6:
[0624] The server converts the text message into a voice message using a speech synthesis engine. The inputs are the generated text message and the analyzed voice characteristics. In concrete terms, the server converts the text message into voice using the speech synthesis engine, and the output is the voice message.
[0625] Step 7:
[0626] The server generates a video based on the generated voice message and facial features. The inputs are the voice message and facial features. Specifically, it starts the video generation engine, synchronizes the facial movements with the voice, and generates a video that appears to be speaking as output.
[0627] Step 8:
[0628] The server stores the generated video in cloud storage and issues a security token. The input is the generated video data. Specifically, the server encrypts the video data, stores it in cloud storage, and issues a security token to verify the user's access rights. The output is the securely stored video data and a security token.
[0629] Step 9:
[0630] The user accesses the cloud service again and plays the generated video in streaming format. The inputs are the user's account information and a security token. Specifically, the server verifies the security token and grants the user permission to play the video. The output is the video played in streaming format and provided to the user.
[0631] (Application example 1)
[0632] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0633] Current memorial videos and photo albums are limited to still images or simple slideshows, making it difficult for bereaved families to visually and powerfully reproduce the voices and facial movements of their loved ones. Furthermore, privacy and security concerns pose a demand for a reliable system for safely creating, storing, and playing video content. Therefore, a video message generation system that can convey more personal emotions and deepen emotional connections is needed.
[0634] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0635] In this invention, the server is installed on a smartphone or tablet and includes means for generating, saving, and playing back video messages of memories based on the voices and photos of loved ones, means for analyzing uploaded photo and audio data to extract facial and vocal characteristics of people, and means for saving the generated videos in cloud storage and making them playable in streaming format. This makes it possible to generate memorial videos of loved ones and strengthen emotional ties in a safe and reliable manner.
[0636] "Photographic data from before death" refers to digital data of photographs taken of the deceased person before death.
[0637] "Voice data from before death" refers to digital data of the voice of a deceased person recorded before death.
[0638] "Uploading means" refers to a method or device for transmitting data owned by a user to a cloud server via the Internet.
[0639] "Facial features" are characteristics or specific points of a face extracted by facial recognition technology and are used to identify an individual.
[0640] "Voice features" are characteristics of a voice extracted by voice recognition technology and are used to identify an individual's voice.
[0641] A "text message" is character data in a natural language, and is generated as a message from a loved one.
[0642] A "voice message" is data obtained by converting a generated text message into voice.
[0643] The "means for generating moving images" refers to a method or device that converts still images into dynamic moving images based on facial features and voice messages.
[0644] "Cloud storage" is a data storage area on a remote server for storing files on the Internet.
[0645] "Streaming format" is a technology that allows data to be played in real time while being downloaded.
[0646] A "smartphone or tablet" is a type of mobile device that has advanced computing power and internet connectivity.
[0647] "Memory Video Message" is an emotional video message created using the face and voice of the deceased.
[0648] A "security token" is temporary authentication information used to verify access privileges.
[0649] A "dialect and speech database" is a collection of data that records the speech characteristics of a particular region or particular individual.
[0650] This invention provides a system for generating video messages based on photographic and audio data of loved ones, allowing users to receive emotionally rich messages by recreating the voices and faces of loved ones.
[0651] System configuration
[0652] This system consists of a user device, a cloud server, and cloud storage. Each component is explained below.
[0653] User's device
[0654] The user's device is a smartphone or tablet, which has the function of uploading photo data and audio data to a cloud server.
[0655] It also has the ability to play streaming videos provided by cloud servers.
[0656] Cloud Server
[0657] The cloud server has the function of receiving and analyzing the photo data and audio data.
[0658] It also generates text messages based on the analysis results and converts them into voice messages, and generates videos based on the generated voice messages and facial features, saving them to cloud storage.
[0659] The main software used includes Amazon Rekognition for image analysis, Pyttsx3 for speech synthesis, and OpenCV and MoviePy for video generation.
[0660] Cloud Storage
[0661] Cloud storage is a database that safely stores the generated videos and provides them to users in streaming format.
[0662] By using security tokens, user access rights can be confirmed and data management can be ensured safely.
[0663] Program processing
[0664] On the user's device:
[0665] Users access the cloud service, log in, and upload photos and audio data of their precious family members from their devices. At this time, the photo data is converted to JPEG format, and the audio data is converted to WAV format.
[0666] Cloud Server:
[0667] After receiving the uploaded data, an image analysis module extracts facial features and generates a high-resolution image.
[0668] A speech analysis module analyzes the audio data and extracts vocal characteristics, in this case reflecting specific patterns using a database of dialects and accents.
[0669] Based on the text message entered by the user, a generative AI model is used to generate a naturally-phrased text message, which is then converted into a voice message using a speech synthesis engine, incorporating the analyzed voice characteristics.
[0670] Finally, the generated voice message and facial features are used to generate a video that makes the face appear to be moving.
[0671] Cloud Storage:
[0672] The generated video is stored in cloud storage and a security token is issued to protect the data.
[0673] When a user plays a video, the cloud server checks the user's access rights and then provides the video securely in streaming format.
[0674] Specific examples
[0675] Example 1:
[0676] A user logs in to a cloud service and uploads a photo of their deceased mother and her last recorded message (e.g., a birthday message). The cloud server analyzes the received image data using Amazon Rekognition to extract the mother's facial features. Next, the audio data is analyzed using Pyttsx3 to extract the mother's tone and accent. When the user requests a message saying "Happy Birthday," the generative AI model generates a naturally phrased text message, which is then converted into an audio message using Pyttsx3. Finally, a video of the mother speaking is generated using OpenCV and MoviePy and saved to cloud storage. The user then accesses the cloud service again to play the generated video.
[0677] Prompt for the generative AI model:
[0678] prompt:
[0679] "Generate a text message when a user requests a message like 'Happy Birthday.' The generated text should be natural-sounding and emotive."
[0680] The above is a specific description for carrying out the invention.
[0681] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0682] Step 1:
[0683] A user accesses a cloud service using a smartphone or tablet and logs in. The login information is sent to an authentication server, which returns an authentication result. This identifies the user and allows them to use the system.
[0684] Step 2:
[0685] Users upload photos and audio data of their deceased family members. The photo files (JPEG format) and audio files (WAV format) selected by the user are sent via the Internet to a cloud server, which receives the data and formats it appropriately.
[0686] Step 3:
[0687] The cloud server analyzes the received photo data. First, it uses Amazon Rekognition to analyze the image and extract facial features. This includes information such as the position of the eyes, nose, and mouth, as well as facial contours and skin tone. The extracted facial feature data is used in the next processing step.
[0688] Step 4:
[0689] The cloud server analyzes the received voice data using Pyttsx3 to extract voice features, including patterns such as tone, accent, and speaking habits. The extracted voice feature data is then used to generate the voice message.
[0690] Step 5:
[0691] The text message entered by the user is sent to a cloud server, which uses a generative AI model to generate a text message with natural wording. In this process, a more emotionally charged message is generated based on the user's requested phrases.
[0692] Step 6:
[0693] The cloud server converts the generated text message into a voice message using the Pyttsx3 speech synthesis engine, which generates a voice message that reflects the analyzed voice characteristics, ensuring that the voice message sounds as close as possible to the voice of the deceased person that the user is familiar with.
[0694] Step 7:
[0695] The cloud server generates a video using the generated voice message and the extracted facial features. Using tools such as OpenCV and MoviePy, the video is created so that the face appears to be moving. The video is designed to play back the message entered by the user emotionally.
[0696] Step 8:
[0697] The cloud server saves the generated video in cloud storage and issues a security token to protect the data. This token is used to verify the user's access rights when playing the video. Storing the video in cloud storage ensures secure data management.
[0698] Step 9:
[0699] The user then accesses the cloud service again and plays the generated video in streaming format. The video data is provided from the cloud storage and played in real time on the user's device. At this time, the user's access authority is confirmed using a security token.
[0700] This series of processing steps generates an emotional video message that reproduces the voices and faces of loved ones, allowing users to watch it safely.
[0701] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0702] This invention relates to a system that generates video messages based on photographic and audio data from the deceased's life, and also describes an embodiment that combines an emotion engine. This system allows users to recreate the voice and face of a loved one and receive a heartfelt message that reflects their emotions.
[0703] System configuration
[0704] This system consists of a user's device, a cloud server, cloud storage, and an emotion engine. Each component will be explained below.
[0705] User's device
[0706] The user's device has the function of uploading photo data and audio data to the cloud server, the function of playing streaming videos provided by the cloud server, and the function of collecting the user's emotional data and sending it to the cloud server.
[0707] Cloud Server
[0708] The cloud server receives and analyzes photo and audio data. It also generates text messages based on the analysis results and converts them into audio messages. It also uses an emotion engine to recognize the user's emotions, adjusts the message content and tone as necessary, and generates videos.
[0709] Cloud Storage
[0710] Cloud storage is a database that securely stores the generated videos and provides them to users in streaming format. By using security tokens, user access rights are confirmed and secure data management is achieved.
[0711] Emotion Engine
[0712] The emotion engine analyzes emotional data sent from the user's device and recognizes the user's emotional state in real time. Based on this information, it adjusts the tone and content of generated messages, audio, and video.
[0713] Program processing
[0714] The program in this system plays four main roles: user, terminal, server, and emotion engine. The processing flow and specific operations are explained below.
[0715] First, the user accesses the cloud service and logs in. Next, the user uploads photos and recorded voices of their deceased family members from their device. The device converts these data into appropriate formats (e.g., JPEG and WAV) and sends them to the cloud server.
[0716] Once the server receives the uploaded data, it first performs image analysis. The image analysis module extracts facial features and generates a high-resolution image. Then, the voice analysis module analyzes the audio data and extracts voice characteristics, including dialect and speech patterns.
[0717] Next, based on the text message entered by the user, the server uses a generative AI model to generate a naturally-phrased text message, which is then converted into a voice message using a speech synthesis engine, reflecting the analyzed voice characteristics.
[0718] Furthermore, the emotion engine receives emotion data sent from the user's device and analyzes the user's emotional state. Based on the results, the generated message content and voice tone are adjusted. Specifically, different messages and tones are adopted depending on changes in emotion.
[0719] The server then uses the generated voice message and facial features to generate a video that appears to be moving, which is then stored in cloud storage and a security token is issued to protect the data.
[0720] Finally, the user accesses the cloud service again and plays the generated video in streaming format. The server verifies the user's access privileges and then plays the video securely using a security token.
[0721] Specific examples
[0722] 1. The user logs in to the cloud service and uploads a photo of their deceased mother and her last recorded message (e.g., a birthday message) from their device to the cloud server.
[0723] 2. The server analyzes the uploaded photo, extracts the mother's facial features, and generates a high-resolution image. At the same time, it performs audio analysis to extract the mother's tone of voice, accent, and speaking habits.
[0724] 3. The user requests a video of their mother wishing them a happy birthday. Based on this request, the server uses a generative AI model to generate a text message and converts it into a voice message.
[0725] 4. At this time, the emotion engine analyzes the user's emotions and, for example, if it recognizes that the user is emotional, it adjusts the voice tone to a gentler one.
[0726] 5. The server generates a video of the mother speaking using the generated voice message and the analyzed high-resolution facial features.
[0727] 6. The server stores the generated video in cloud storage and issues a security token to protect the data.
[0728] 7. The user accesses the cloud service to play the generated video and watch the heartwarming message from their deceased mother.
[0729] The above is a specific description of an embodiment of the "My Album in Action" system that combines the emotion engine of the present invention.
[0730] The processing flow will be explained below.
[0731] Step 1: Log in and upload data
[0732] 1. The user accesses the cloud service and logs in by entering their user ID and password.
[0733] 2. The device sends the login information to the cloud server for authentication.
[0734] 3. The server verifies the authentication information and, if successful, grants access to the user.
[0735] 4. The user selects the photo and audio data of the deceased family member on the device and uploads it to the cloud server.
[0736] 5. The device converts the selected data into the appropriate format (e.g., JPEG and WAV) and sends it to the cloud server.
[0737] 6. The server stores the received photo and audio data in cloud storage.
[0738] Step 2: Data analysis
[0739] 1. The server passes the uploaded photo data to the image analysis module.
[0740] 2. The server uses an image analysis module to extract facial features (such as the position of the eyes, nose, and mouth) of family members from the photo.
[0741] 3. The server uses super-resolution technology to generate a high-resolution facial image based on the extracted feature points.
[0742] 4. The server passes the uploaded voice data to the voice analysis module.
[0743] 5. The server uses a speech analysis module to extract tone, rate, accent, and dialect and diction patterns from the speech data.
[0744] Step 3: Collect and analyze emotion data
[0745] 1. The user interacts with the cloud service interface and gives permission to collect emotion data.
[0746] 2. The device collects emotional data such as the user's facial expressions and tone of voice and sends it to a cloud server.
[0747] 3. The server passes the received emotion data to the emotion engine and analyzes the user's emotional state.
[0748] Step 4: Message Generation
[0749] 1. The user enters the content of the message they want to generate (e.g., "Happy Birthday") through the cloud service interface.
[0750] 2. The server passes the input text message to a generative AI model (e.g., GPT-4) to generate a naturally phrased text message.
[0751] 3. The server uses a speech synthesis engine to generate a voice message that sounds like the family member's voice based on the generated text message and the extracted voice characteristics.
[0752] 4. The server adjusts the tone of the voice and the content of the message based on the user's emotional data provided by the emotion engine.
[0753] Step 5: Video Generation
[0754] 1. The server uses a high-resolution facial image and the generated voice message to generate a video that matches the facial movements and voice using ravioli generation technology (e.g., Facial Animation Technology).
[0755] 2. The server stores the generated video in cloud storage and applies appropriate encryption processes.
[0756] Step 6: Streaming
[0757] 1. The user clicks the play button to request playback of the video generated from the cloud service.
[0758] 2. The server verifies the user's access rights, issues a security token, and sends it to the user's device.
[0759] 3. The device receives the security token and sends a request to the server again to start streaming the video.
[0760] 4. The server verifies the security token and sends the video data to the device in streaming format.
[0761] 5. The device plays the received video data, allowing the user to watch the video.
[0762] The above is a description of the specific operations divided into processing steps of the "Moving My Album" system.
[0763] Example 2
[0764] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0765] In recent years, there has been a demand for preserving memories of deceased loved ones and receiving them as emotionally warm messages. However, conventional technologies have struggled to reproduce the voice and appearance of the deceased in a natural way. It has also been difficult to generate messages and videos that correspond to the user's emotional state. Furthermore, ensuring the security of the generated data and managing access rights are also important issues.
[0766] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0767] In this invention, the server includes means for acquiring image data of the deceased person, means for acquiring audio data of the deceased person, means for analyzing the acquired image data and audio data and extracting facial and vocal characteristics of the person, means for generating a natural language message based on the extracted facial and vocal characteristics and converting the message into synthetic speech, means for analyzing the user's emotional data and reflecting the emotional state in the generation process, means for generating a video using the generated synthetic speech and the analyzed facial characteristics, and means for saving the generated video in a storage device and making it playable in streaming format. This makes it possible to reproduce the voice and appearance of the deceased person with high accuracy, and to easily generate and safely manage warm video messages that correspond to the user's emotional state.
[0768] "Image data from before death" refers to photographs and image files taken of a person before their death.
[0769] "Pre-death acoustic data" refers to voices and audio files recorded during a person's lifetime.
[0770] "Means of acquisition" refers to the functions and interfaces that allow users to select and upload data.
[0771] "Means for analysis" refers to the algorithms and software used to analyze the acquired data and extract the necessary information.
[0772] "Facial features" refer to identifiable parts or attributes of a person's face that are extracted from an image.
[0773] "Voice features" refer to identifiable parts or attributes such as tone, accent, and catchphrases extracted from voice data.
[0774] A "natural language message" refers to a text message that is written in a way that is natural for humans to understand.
[0775] "Synthetic voice" refers to data generated as voice based on a text message.
[0776] "Emotional Data" refers to information collected to describe a user's emotional state.
[0777] The "generation process" refers to a series of steps to generate messages and videos based on analyzed data.
[0778] "Storage device" refers to a storage system for safely storing the generated video.
[0779] "Streaming format" refers to a format that allows users to watch videos in real time over the Internet.
[0780] This invention relates to a system that generates video messages based on image and audio data from a person's life, and by combining it with an emotion engine, it allows users to receive the messages as warm messages. This system consists of a user's device, a cloud server, cloud storage, and an emotion engine.
[0781] System configuration
[0782] User's device
[0783] The user's device has the function of uploading image data and audio data to the cloud server, the function of playing streaming videos provided by the cloud server, and the function of collecting the user's emotional data and sending it to the cloud server.
[0784] Cloud Server
[0785] The cloud server is equipped with software and hardware for receiving and analyzing image and acoustic data. The analysis uses OpenCV, an image processing software, and NVIDIA NeMo, a voice analysis software. Based on the text message entered by the user, a generative AI model (e.g., a generative AI model) is used to generate a text message with natural wording.
[0786] The server converts the generated text message into a voice message using a speech synthesis engine, incorporating the analyzed voice characteristics. This process uses the speech synthesis technology of Azure Cognitive Services.
[0787] Furthermore, the Emotion Analysis Engine analyzes the user's emotional data and adjusts the generated message content and voice tone based on the user's emotional state. For example, if the user is emotional, the voice tone will be made gentler.
[0788] The server uses the generated voice message and high-resolution facial features to generate a video that appears to show a moving face using DeepFaceLab. The video is encoded using FFmpeg. The video is then stored in cloud storage, where data is protected by a security token.
[0789] Cloud Storage
[0790] Cloud storage is a database that securely stores the generated videos and provides them to users in streaming format. By using security tokens, user access rights are confirmed and secure data management is achieved.
[0791] Emotion Engine
[0792] The emotion engine analyzes emotional data sent from the user's device and recognizes the user's emotional state in real time. Based on this information, it adjusts the tone and content of generated messages, audio, and video.
[0793] Specific examples
[0794] A user logs in to the cloud service and uploads a photo (JPEG format) of a deceased family member and audio data (WAV format) from their device to the cloud server. The server analyzes the uploaded data using OpenCV and NVIDIA NeMo to extract facial and vocal characteristics. Next, the generative AI model converts the user's requested message (e.g., "Happy Birthday") into natural-sounding text, which is then converted into a voice message using Azure Cognitive Services' speech synthesis engine. The emotion engine also analyzes the user's emotional data and adjusts the message content and voice tone. Finally, the server generates a video using DeepFaceLab and saves it in cloud storage. The user can then access the cloud service again to play the generated video in streaming format.
[0795] Example prompt: "Generate a sweet message from my late mother wishing me a happy birthday."
[0796] The above is a specific description of an embodiment of the "My Album in Action" system that combines the emotion engine of the present invention.
[0797] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0798] Step 1:
[0799] A user logs in
[0800] A user accesses a cloud service using a browser or application and enters a user ID and password, which authenticates the user and allows them to log in to the cloud service.
[0801] Input: User ID, Password
[0802] Output: Authentication token, indicating successful login
[0803] Step 2:
[0804] The user uploads photos and audio data.
[0805] The user selects a photo (JPEG format) and audio data (WAV format) of the deceased family member from the device and clicks the upload button. The device converts the data into the appropriate format (JPEG and WAV) and sends it to the cloud server.
[0806] Input: Photo data (JPEG), audio data (WAV)
[0807] Output: Photo and audio data uploaded to the cloud server
[0808] Step 3:
[0809] The server receives the photo and audio data.
[0810] The server receives the photo and audio data sent from the device and temporarily stores it in the specified storage. At this stage, the server checks the integrity of the data to ensure there is no unauthorized data.
[0811] Input: Photo data and audio data uploaded from the device
[0812] Output: Temporarily saved photo and audio data
[0813] Step 4:
[0814] The server performs image analysis
[0815] The server analyzes the received photo data using OpenCV, specifically applying a facial recognition algorithm to extract facial features.
[0816] Input: Temporarily saved photo data
[0817] Output: Facial features (position, shape, landmark points, etc.)
[0818] Step 5:
[0819] The server performs voice analysis
[0820] The server uses NVIDIA NeMo to analyze the received voice data, extracting voice tone, accent, and speaking habits.
[0821] Input: Temporarily saved audio data
[0822] Output: Voice characteristics (tone, accent, speech patterns)
[0823] Step 6:
[0824] The server generates a text message
[0825] Based on a prompt phrase entered by the user (e.g., "Say happy birthday"), the server uses a generative AI model (e.g., GPT-3) to generate a natural-sounding text message.
[0826] Input: The prompt text entered by the user
[0827] Output: The generated text message
[0828] Step 7:
[0829] The server converts it into a voice message
[0830] Based on the generated text message, the server uses Azure Cognitive Services' speech synthesis engine to convert it into a voice message, which reflects the analyzed voice characteristics.
[0831] Input: Generated text message, voice features
[0832] Output: Synthesized voice message
[0833] Step 8:
[0834] The emotion engine analyzes the emotion data
[0835] The emotion engine analyzes emotion data sent from the user's device and recognizes the user's emotional state in real time.
[0836] Input: User emotion data (facial expressions, tone of voice, etc.)
[0837] Output: Emotional state analysis result
[0838] Step 9:
[0839] The server generates a video based on the generated data.
[0840] The server uses DeepFaceLab to generate a video that appears to show a moving face based on the generated voice message and high-resolution facial features, and encodes the video using FFmpeg.
[0841] Input: Synthesized voice message, facial features
[0842] Output: Generated video file
[0843] Step 10:
[0844] The server saves the video to cloud storage.
[0845] The generated video is stored in cloud storage. A security token is issued to verify the user's access rights and protect the data.
[0846] Input: Generated video file
[0847] Output: Video stored in cloud storage, security token
[0848] Step 11:
[0849] Playing user-generated videos
[0850] The user then accesses the cloud service again and plays the video. The server then verifies the user's access privileges and uses the security token to play the video securely in streaming format.
[0851] Input: Security token, saved video file
[0852] Output: Streamed video
[0853] The above are the detailed processing steps of this system.
[0854] (Application example 2)
[0855] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0856] In order to recreate the memory of the deceased and move the bereaved family, it is desirable to provide more realistic and emotionally relevant video messages. However, current technology has problems in that it is difficult to adjust the tone and content of the message according to the user's emotional state, and the generated videos lack a sense of realism. In addition, there is a lack of appropriate display methods for viewing moving videos in real time. There is a need to solve these problems.
[0857] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0858] In this invention, the server includes means for uploading photo data of the person before death, means for uploading audio data of the person before death, means for analyzing the uploaded photo data and audio data and extracting facial and vocal characteristics of the person, means for generating a text message based on the extracted facial and vocal characteristics and converting the text message into an audio message, means for generating a video using the generated audio message and the analyzed facial characteristics, means for storing the generated video in cloud storage and making it playable in streaming format, means for collecting emotional data and analyzing the data to adjust the tone of the text message or audio message, and means for displaying the generated video using a head-mounted display, thereby making it possible to provide a realistic video message that matches the emotional state of the user.
[0859] "Photo data from before death" refers to image files of a person taken before death.
[0860] "Voice data from before death" refers to an audio file of a person's voice recorded before death.
[0861] The "means for uploading" refers to a method and apparatus for transmitting data from a terminal to a cloud server.
[0862] "Means for analyzing" refers to methods and devices that process uploaded data to extract information.
[0863] "Facial features" are characteristic parts of a person's face that are extracted from the analyzed image data.
[0864] "Voice features" are characteristics of a person's voice extracted from analyzed voice data.
[0865] A "means for generating a text message" is a method and device that creates a sentence based on input information.
[0866] A "means for converting to voice message" is a method and apparatus for converting a text message to voice data.
[0867] The "means for generating video" refers to a method and device for creating a video file based on image data and audio data.
[0868] "Cloud storage means" refers to a method and device for storing data on a remote server.
[0869] "Means for enabling playback in streaming format" refers to a method and apparatus for playing back stored video data in real time via the Internet.
[0870] A "means for collecting emotional data" is a method and apparatus for collecting a user's emotional state.
[0871] A "means for analyzing emotion data" is a method and apparatus for processing collected emotion data to make sense of the information.
[0872] "Means for adjusting the tone of text and voice messages" refers to a method and apparatus for changing the tone of message content based on analyzed emotional data.
[0873] A "head-mounted display" is a display device worn on the head, and is used to display videos and images.
[0874] This invention relates to a system for generating moving video messages using photographic and audio data from a person's lifetime. Furthermore, by collecting and analyzing the user's emotional data, the system adjusts the message content and tone to generate more personalized and emotional videos. A specific embodiment of this system is described below.
[0875] System configuration
[0876] This system consists of a user's device, a cloud server, cloud storage, an emotion engine, and a head-mounted display.
[0877] User's device
[0878] The user's device has the function of uploading photo data and audio data to the cloud server, the function of playing streaming videos provided by the cloud server, and the function of collecting the user's emotional data and sending it to the cloud server.
[0879] Cloud Server
[0880] The cloud server receives and analyzes photo and audio data. It also generates text messages based on the analysis results and converts them into audio messages. It also uses an emotion engine to recognize the user's emotions, adjusts the message content and tone as necessary, and generates videos.
[0881] Cloud Storage
[0882] Cloud storage is a database that securely stores the generated videos and provides them to users in streaming format. By using security tokens, user access rights are confirmed and secure data management is achieved.
[0883] Emotion Engine
[0884] The emotion engine analyzes emotional data sent from the user's device and recognizes the user's emotional state in real time. Based on this information, it adjusts the tone and content of generated messages, audio, and video.
[0885] head-mounted display
[0886] A head-mounted display is a device worn on the head that displays images to the user. It displays video received from a cloud server in real time, providing the user with a realistic viewing experience.
[0887] Program processing
[0888] The program in this system mainly plays four main roles: user, terminal, server, and emotion engine.
[0889] First, the user accesses the cloud service and logs in. The user uploads photos and audio data from their device to the cloud server. The device converts the data into an appropriate format (e.g., JPEG and WAV) and sends it to the cloud server.
[0890] Once the cloud server receives the uploaded data, it first performs image analysis. The image analysis module extracts facial features and generates a high-resolution image. The voice analysis module then analyzes the audio data and extracts voice characteristics, including dialect and speech patterns.
[0891] Next, based on the text message entered by the user, the cloud server uses a generative AI model to generate a naturally-phrased text message, which is then converted into a voice message using a speech synthesis engine, reflecting the analyzed voice characteristics.
[0892] Furthermore, the emotion engine receives emotion data sent from the user's device and analyzes the user's emotional state. Based on the results, it adjusts the generated message content and voice tone. Specifically, it adopts different messages and tones depending on changes in the user's emotions.
[0893] The cloud server then uses the generated voice message and facial features to generate a video that appears to be moving, which is then stored in cloud storage and a security token is issued to protect the data.
[0894] Finally, the user wears a head-mounted display and watches the generated video. The cloud server verifies the user's access rights and then plays the video securely using a security token.
[0895] Specific examples
[0896] A user visits a funeral home and uploads a photo of their deceased father and a recorded voice message (e.g., a birthday message). The cloud server analyzes the uploaded photo, extracts the father's facial features, and generates a high-resolution image. At the same time, it performs voice analysis to extract the father's tone of voice, accent, and catchphrases.
[0897] A user can request a generative AI model to generate a video of their father saying, "Take care." The emotion engine then analyzes the user's emotions and adjusts the tone of the voice to be gentler if it detects that the user is emotional, for example.
[0898] The cloud server then uses the generated voice message and analyzed high-resolution facial features to generate a video of the father speaking, which is then stored in cloud storage and a security token is issued to protect the data.
[0899] Finally, the user can wear a head-mounted display to watch the generated video and receive a heartwarming message from their deceased father.
[0900] Example prompts for generative AI models
[0901] Below is an example of a prompt sentence for the generative AI model (generative AI model).
[0902] "Dear visitors, wipe away your tears and upload your favorite photos and audio recordings here. We will recreate the voice and face of your loved one and deliver a heartwarming video message. Let's share this special moment together. We will recreate phrases such as 'Thank you' and 'Take care' in the video."
[0903] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0904] Step 1:
[0905] The user accesses the cloud service and logs in.
[0906] Input: User authentication information (user ID, password)
[0907] Output: Authentication result (success / failure)
[0908] Specific behavior:
[0909] The user's device sends authentication information to the cloud server, which checks it against the database and, if authentication is successful, proceeds to the next step.
[0910] Step 2:
[0911] Users upload photo data and audio data from their pre-death lives from their devices to a cloud server.
[0912] Input: Photo data from before death (JPEG), audio data from before death (WAV)
[0913] Output: Data stored on the cloud server
[0914] Specific behavior:
[0915] Users simply select photos and audio data from their device and click the upload button. The device then sends the data to the cloud server, which stores it securely.
[0916] Step 3:
[0917] The cloud server analyzes the uploaded photo data.
[0918] Input: Photo data from before death (JPEG)
[0919] Output: Facial feature data (facial feature points, facial expressions)
[0920] Specific behavior:
[0921] The server uses an image analysis module (e.g., OpenCV) to extract facial features from the photo data and generate high-resolution facial images.
[0922] Step 4:
[0923] The cloud server analyzes the voice data from the person's life.
[0924] Input: Audio data from before death (WAV)
[0925] Output: Voice characteristics data (tone of voice, accent, catchphrases)
[0926] Specific behavior:
[0927] The server uses a voice analysis module (e.g., librosa) to extract voice features from the audio data.
[0928] Step 5:
[0929] The cloud server generates a message using a generative AI model based on the text message entered by the user.
[0930] Input: User-entered text messages, facial feature data, and voice feature data
[0931] Output: Naturally-phrased text message
[0932] Specific behavior:
[0933] The server sends a prompt to a generative AI model (e.g., GPT-3) to generate a naturally phrased text message.
[0934] Step 6:
[0935] The cloud server converts the generated text message into a voice message.
[0936] Input: Naturally-phrased text messages, voice feature data
[0937] Output: Voice message (audio file)
[0938] Specific behavior:
[0939] The server uses a speech synthesis engine (e.g., Watson Text to Speech) to convert the text message into speech, and the generated voice message reflects the voice characteristics.
[0940] Step 7:
[0941] The emotion engine analyzes the user's emotion data and transmits the emotional state to the cloud server.
[0942] Input: User emotion data (real-time emotional state)
[0943] Output: Emotion analysis result (emotional state)
[0944] Specific behavior:
[0945] The emotion engine analyzes the emotion data collected from the user's device and sends the results to a cloud server.
[0946] Step 8:
[0947] The cloud server adjusts the message content and voice tone based on the results of emotion analysis.
[0948] Input: Emotion analysis results, voice message, facial feature data
[0949] Output: Adjusted message content, audio tone
[0950] Specific behavior:
[0951] The results of the emotion engine are reflected and the tone of the generated text and voice messages is adjusted to match the user's emotions.
[0952] Step 9:
[0953] The cloud server generates a video using the generated voice message and facial features.
[0954] Input: Facial feature data, voice message
[0955] Output: Video file
[0956] Specific behavior:
[0957] The server uses a video generation module (for example, Adobe After Effects API) to generate a video in which a person's face moves based on the voice message and facial feature data.
[0958] Step 10:
[0959] The cloud server stores the generated video in cloud storage and issues a security token to protect the data.
[0960] Input: Video file
[0961] Output: Saved video file, security token
[0962] Specific behavior:
[0963] The server stores the video file in cloud storage (e.g., AWS S3) and issues a security token for access control.
[0964] Step 11:
[0965] The user wears a head-mounted display and watches videos generated from a cloud server.
[0966] Input: Security Token
[0967] Output: Video Streaming
[0968] Specific behavior:
[0969] The user accesses the cloud service again and watches the video in real time using a head-mounted display (e.g., HoloLens, Oculus Rift). The cloud server uses a security token to verify the user's access privileges and plays the video securely.
[0970] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0971] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0972] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0973] [Third embodiment]
[0974] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0975] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0976] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0977] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0978] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0979] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0980] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0981] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0982] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0983] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0984] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0985] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0986] This invention provides a system for generating video messages based on photographic and audio data of loved ones, and describes an embodiment of the system. By implementing this system, users can recreate the voices and faces of loved ones and receive them as heartwarming messages.
[0987] System configuration
[0988] This system consists of a user terminal, a cloud server, and a storage system. Each component is explained below.
[0989] User's device
[0990] The user's device has the function of uploading photo data and audio data to the cloud server, and also has the function of playing streaming video provided by the cloud server.
[0991] Cloud Server
[0992] The cloud server has the functions of receiving and analyzing photo data and audio data, generating text messages based on the analysis results and converting them into audio messages, and generating videos based on the generated audio messages and facial features, and storing them in cloud storage after ensuring appropriate security.
[0993] Cloud Storage
[0994] Cloud storage is a database that securely stores the generated videos and provides them to users in streaming format. By using security tokens, user access rights are confirmed and secure data management is achieved.
[0995] Program processing
[0996] The program in this system plays three main roles: user, terminal, and server. The process flow and specific operations are explained below.
[0997] First, the user accesses the cloud service and logs in. Next, the user uploads photos and recorded voices of their deceased family members from their device. The device receives this data, converts it into an appropriate format (e.g., JPEG and WAV), and sends it to the cloud server.
[0998] The server then receives the uploaded data and performs image analysis. The image analysis module extracts facial features and generates a high-resolution image. The audio analysis module then analyzes the audio data and extracts voice characteristics, including dialect and speech patterns.
[0999] Next, based on the text message entered by the user, the server uses a generative AI model to generate a naturally-phrased text message, which is then converted into a voice message using a speech synthesis engine, reflecting the analyzed voice characteristics.
[1000] The server then uses the generated voice message and facial features to generate a video that appears to be moving, which is then stored in cloud storage and a security token is issued to protect the data.
[1001] Finally, the user accesses the cloud service again and plays the generated video in streaming format. The server verifies the user's access privileges and then plays the video securely using a security token.
[1002] Specific examples
[1003] 1. The user logs in to the cloud service and uploads a photo of their deceased mother and her last recorded message (e.g., a birthday message) from their device to the cloud server.
[1004] 2. The server analyzes the uploaded photo, extracts the mother's facial features, and generates a high-resolution image. At the same time, it performs audio analysis to extract the mother's tone of voice, accent, and speaking habits.
[1005] 3. The user requests a video of their mother wishing them a happy birthday. Based on this request, the server uses a generative AI model to generate a text message and converts it into a voice message.
[1006] 4. The server generates a video of the mother speaking using the generated voice message and the analyzed high-resolution facial features.
[1007] 5. The server stores the generated video in cloud storage and issues a security token to protect the data.
[1008] 6. The user accesses the cloud service to play the generated video and watch the heartwarming message from their deceased mother.
[1009] The above is a specific description of the embodiment of the "Animated My Album" system of the present invention.
[1010] The processing flow will be explained below.
[1011] Step 1: Log in and upload data
[1012] 1. The user accesses the cloud service and logs in by entering their user ID and password.
[1013] 2. The device sends the login information to the cloud server for authentication.
[1014] 3. The server verifies the authentication information and, if successful, grants access to the user.
[1015] 4. The user selects the photo and audio data of the deceased family member on the device and uploads it to the cloud server.
[1016] 5. The device converts the selected data into the appropriate format (e.g., JPEG and WAV) and sends it to the cloud server.
[1017] 6. The server stores the received photo and audio data in cloud storage.
[1018] Step 2: Data analysis
[1019] 1. The server passes the uploaded photo data to the image analysis module.
[1020] 2. The server uses an image analysis module to extract facial features (such as the position of the eyes, nose, and mouth) of family members from the photo.
[1021] 3. The server uses super-resolution technology to generate a high-resolution facial image based on the extracted feature points.
[1022] 4. The server passes the uploaded voice data to the voice analysis module.
[1023] 5. The server uses a speech analysis module to extract tone, rate, accent, and dialect and diction patterns from the speech data.
[1024] Step 3: Message Generation
[1025] 1. The user enters the content of the message they want to generate (e.g., "Happy Birthday") through the cloud service interface.
[1026] 2. The server passes the input text message to a generative AI model (e.g., GPT-4) to generate a naturally phrased text message.
[1027] 3. The server uses a speech synthesis engine to generate a voice message that sounds like the family member's voice based on the generated text message and the extracted voice characteristics.
[1028] Step 4: Video Generation
[1029] 1. The server uses a high-resolution facial image and the generated voice message to generate a video that matches the facial movements and voice using ravioli generation technology (e.g., Facial Animation Technology).
[1030] 2. The server stores the generated video in cloud storage and applies appropriate encryption processes.
[1031] Step 5: Streaming
[1032] 1. The user clicks the play button to request playback of the video generated from the cloud service.
[1033] 2. The server verifies the user's access rights, issues a security token, and sends it to the user's device.
[1034] 3. The device receives the security token and sends a request to the server again to start streaming the video.
[1035] 4. The server verifies the security token and sends the video data to the device in streaming format.
[1036] 5. The device plays the received video data, allowing the user to view the video.
[1037] The above is a description of the specific operations divided into each processing step.
[1038] Example 1
[1039] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1040] Conventional commemorative video generation systems that use photographs and audio data have difficulty adequately reproducing the face and voice of the deceased as desired by users, and there are also issues with the quality and realism of the generated videos. As a result, it is difficult for users to receive the heartwarming messages of their loved ones in a realistic way. Furthermore, there are insufficient aspects of data protection and security, which increases the risk of personal information being leaked.
[1041] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1042] In this invention, the server includes: means for a user to access digital data storage and upload photo data and audio data; means for converting the uploaded photo data and audio data into an appropriate format; means for transmitting the converted photo data and audio data to a cloud server; means for the cloud server to analyze the received photo data and audio data and extract facial and vocal characteristics of a person; means for generating a text message using a generative AI model based on the extracted facial and vocal characteristics and converting the text message into an audio message; means for generating a video using the generated audio message and the analyzed facial characteristics; and means for saving the generated video in cloud storage and making it playable in streaming format. This allows users to safely and easily generate high-quality, realistic commemorative videos and receive warm messages from their loved ones in real time.
[1043] "User" refers to an individual who uses the system and uploads photo data and audio data.
[1044] "Digital data storage" refers to a system that stores and makes accessible information electronically.
[1045] "Photo data" refers to still images saved in a format such as JPEG.
[1046] "Audio Data" refers to recorded audio information stored in a format such as WAV.
[1047] "Format" refers to a standard or method for organizing data into a specific form.
[1048] "Cloud server" refers to a remote server that stores, processes, and manages data over the Internet.
[1049] "Analysis" refers to the process of examining received data in detail and extracting features and patterns.
[1050] "Facial features of a person" refers to information about the position and shape of the eyes, nose, mouth, etc. extracted from photographic data.
[1051] "Voice characteristics" refers to information such as voice tone, pronunciation habits, etc. extracted from audio data.
[1052] A "generative AI model" refers to an artificial intelligence model that uses machine learning algorithms to generate new data and information.
[1053] "Text Message" means a message expressed as written words or sentences.
[1054] "Voice message" refers to a message that is an audio representation of a text message.
[1055] "Video" refers to a media format that reproduces movement and sound through the combination of multiple still images and audio.
[1056] "Cloud storage" refers to remote storage that stores and manages data via the Internet.
[1057] "Streaming format" refers to a format in which data is played in real time without downloading.
[1058] "Security token" refers to identification information issued to verify access rights and prevent unauthorized access.
[1059] This invention provides a system for generating video messages based on photographic and audio data of loved ones, and describes an embodiment of the system. By implementing this system, users can recreate the voices and faces of loved ones and receive them as heartwarming messages.
[1060] System configuration
[1061] This system consists of a user terminal, a cloud server, and a cloud storage system. Each component is explained below.
[1062] User's device
[1063] The user's device has the function of uploading photo data and audio data to the cloud server, and also has the function of playing streaming video provided by the cloud server.
[1064] Cloud Server
[1065] The cloud server receives and analyzes photo and audio data. It also generates text messages based on the analysis results and converts them into audio messages. It also generates videos based on the generated audio messages and facial features and saves them in cloud storage. A generative AI model is used to generate text messages with natural wording.
[1066] Cloud Storage
[1067] Cloud storage is a database that securely stores the generated videos and provides them to users in streaming format. By using security tokens, user access rights are confirmed and secure data management is achieved.
[1068] What the program does
[1069] The program in this system plays three main roles: user, terminal, and server. The specific operation is explained below.
[1070] First, users access the cloud service and log in. Then, they upload photos and audio recordings of their deceased family members from their devices. The devices receive these data, convert them into appropriate formats (e.g., JPEG and WAV), and send them to the cloud server.
[1071] The server then receives the uploaded data and first performs image analysis. The image analysis module extracts facial features and generates high-resolution images. The voice analysis module then analyzes the audio data and extracts voice characteristics, including dialect and speech patterns.
[1072] Next, based on the text message entered by the user, the server uses a generative AI model to generate a naturally-phrased text message. The text message is then converted into a voice message using a speech synthesis engine, reflecting the analyzed voice characteristics. An example of a prompt is a request such as, "Please generate a message for my mother celebrating her birthday using her voice and characteristics."
[1073] The server then uses the generated voice message and facial features to generate a video that appears to be moving, which is then stored in cloud storage and a security token is issued to protect the data.
[1074] Finally, the user accesses the cloud service again and plays the generated video in streaming format. The server verifies the user's access privileges and then plays the video securely using a security token.
[1075] Specific examples
[1076] The user logs in to the cloud service and uploads a photo of their deceased mother and her last recorded message (e.g., a birthday message) from their device to the cloud server.
[1077] The server analyzes the uploaded photo, extracts the mother's facial features, and generates a high-resolution image, while also performing audio analysis to extract the mother's tone of voice, accent, and speaking habits.
[1078] A user requests a video of their mother wishing them a happy birthday. Based on this request, the server uses a generative AI model to generate a text message and convert it into an audio message.
[1079] The server uses the generated audio message and analyzed high-resolution facial features to generate a video of the mother speaking.
[1080] The server stores the generated video in cloud storage and issues a security token to protect the data.
[1081] The user accesses the cloud service to play the generated video and watch the heartwarming message from their deceased mother.
[1082] The above is a specific description of the embodiment of the "Animated My Album" system of the present invention.
[1083] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1084] Step 1:
[1085] A user logs in to the cloud service and uploads photos and audio data. The inputs include the user's account information, the JPEG photo data to be uploaded, and the WAV audio data. The output is then prepared on the user's device.
[1086] Step 2:
[1087] The device sends the photo data and audio data received from the user to the server. The inputs are uploaded JPEG photo data and WAV audio data. Specifically, the device sends this data to the cloud server using an HTTP POST request, and the cloud server receives the data as output.
[1088] Step 3:
[1089] The server receives the photo data and audio data and begins analyzing it. The input is JPEG photo data and WAV audio data. Specifically, the server uses an image analysis module to extract facial features from the photo data. A face detection algorithm is used for this analysis, and facial features are obtained as the output.
[1090] Step 4:
[1091] The server analyzes the audio data and extracts voice characteristics. The input is audio data in WAV format. Specifically, the server uses acoustic analysis tools to extract voice characteristics such as tone, accent, and catchphrases from the audio data. The voice characteristics are obtained as output.
[1092] Step 5:
[1093] The server converts the text message entered by the user into a natural-sounding text message based on the generative AI model. The inputs include a text message provided by the user and a prompt: "Generate a message to celebrate my mother's birthday using her voice and characteristics." The output is a natural-sounding text message generated by the generative AI model.
[1094] Step 6:
[1095] The server converts the text message into a voice message using a speech synthesis engine. The inputs are the generated text message and the analyzed voice characteristics. In concrete terms, the server converts the text message into voice using the speech synthesis engine, and the output is the voice message.
[1096] Step 7:
[1097] The server generates a video based on the generated voice message and facial features. The inputs are the voice message and facial features. Specifically, it starts the video generation engine, synchronizes the facial movements with the voice, and generates a video that appears to be speaking as output.
[1098] Step 8:
[1099] The server stores the generated video in cloud storage and issues a security token. The input is the generated video data. Specifically, the server encrypts the video data, stores it in cloud storage, and issues a security token to verify the user's access rights. The output is the securely stored video data and a security token.
[1100] Step 9:
[1101] The user accesses the cloud service again and plays the generated video in streaming format. The inputs are the user's account information and a security token. Specifically, the server verifies the security token and grants the user permission to play the video. The output is the video played in streaming format and provided to the user.
[1102] (Application example 1)
[1103] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1104] Current memorial videos and photo albums are limited to still images or simple slideshows, making it difficult for bereaved families to visually and powerfully reproduce the voices and facial movements of their loved ones. Furthermore, privacy and security concerns pose a demand for a reliable system for safely creating, storing, and playing video content. Therefore, a video message generation system that can convey more personal emotions and deepen emotional connections is needed.
[1105] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1106] In this invention, the server is installed on a smartphone or tablet and includes means for generating, saving, and playing back video messages of memories based on the voices and photos of loved ones, means for analyzing uploaded photo and audio data to extract facial and vocal characteristics of people, and means for saving the generated videos in cloud storage and making them playable in streaming format. This makes it possible to generate memorial videos of loved ones and strengthen emotional ties in a safe and reliable manner.
[1107] "Photographic data from before death" refers to digital data of photographs taken of the deceased person before death.
[1108] "Voice data from before death" refers to digital data of the voice of a deceased person recorded before death.
[1109] "Uploading means" refers to a method or device for transmitting data owned by a user to a cloud server via the Internet.
[1110] "Facial features" are characteristics or specific points of a face extracted by facial recognition technology and are used to identify an individual.
[1111] "Voice features" are characteristics of a voice extracted by voice recognition technology and are used to identify an individual's voice.
[1112] A "text message" is character data in a natural language, and is generated as a message from a loved one.
[1113] A "voice message" is data obtained by converting a generated text message into voice.
[1114] The "means for generating moving images" refers to a method or device that converts still images into dynamic moving images based on facial features and voice messages.
[1115] "Cloud storage" is a data storage area on a remote server for storing files on the Internet.
[1116] "Streaming format" is a technology that allows data to be played in real time while being downloaded.
[1117] A "smartphone or tablet" is a type of mobile device that has advanced computing power and internet connectivity.
[1118] "Memory Video Message" is an emotional video message created using the face and voice of the deceased.
[1119] A "security token" is temporary authentication information used to verify access privileges.
[1120] A "dialect and speech database" is a collection of data that records the speech characteristics of a particular region or particular individual.
[1121] This invention provides a system for generating video messages based on photographic and audio data of loved ones, allowing users to receive emotionally rich messages by recreating the voices and faces of loved ones.
[1122] System configuration
[1123] This system consists of a user device, a cloud server, and cloud storage. Each component is explained below.
[1124] User's device
[1125] The user's device is a smartphone or tablet, which has the function of uploading photo data and audio data to a cloud server.
[1126] It also has the ability to play streaming videos provided by cloud servers.
[1127] Cloud Server
[1128] The cloud server has the function of receiving and analyzing the photo data and audio data.
[1129] It also generates text messages based on the analysis results and converts them into voice messages, and generates videos based on the generated voice messages and facial features, saving them to cloud storage.
[1130] The main software used includes Amazon Rekognition for image analysis, Pyttsx3 for speech synthesis, and OpenCV and MoviePy for video generation.
[1131] Cloud Storage
[1132] Cloud storage is a database that safely stores the generated videos and provides them to users in streaming format.
[1133] By using security tokens, user access rights can be confirmed and data management can be ensured safely.
[1134] Program processing
[1135] On the user's device:
[1136] Users access the cloud service, log in, and upload photos and audio data of their precious family members from their devices. At this time, the photo data is converted to JPEG format, and the audio data is converted to WAV format.
[1137] Cloud Server:
[1138] After receiving the uploaded data, an image analysis module extracts facial features and generates a high-resolution image.
[1139] A speech analysis module analyzes the audio data and extracts vocal characteristics, in this case reflecting specific patterns using a database of dialects and accents.
[1140] Based on the text message entered by the user, a generative AI model is used to generate a naturally-phrased text message, which is then converted into a voice message using a speech synthesis engine, incorporating the analyzed voice characteristics.
[1141] Finally, the generated voice message and facial features are used to generate a video that makes the face appear to be moving.
[1142] Cloud Storage:
[1143] The generated video is stored in cloud storage and a security token is issued to protect the data.
[1144] When a user plays a video, the cloud server checks the user's access rights and then provides the video securely in streaming format.
[1145] Specific examples
[1146] Example 1:
[1147] A user logs in to a cloud service and uploads a photo of their deceased mother and her last recorded message (e.g., a birthday message). The cloud server analyzes the received image data using Amazon Rekognition to extract the mother's facial features. Next, the audio data is analyzed using Pyttsx3 to extract the mother's tone and accent. When the user requests a message saying "Happy Birthday," the generative AI model generates a naturally phrased text message, which is then converted into an audio message using Pyttsx3. Finally, a video of the mother speaking is generated using OpenCV and MoviePy and saved to cloud storage. The user then accesses the cloud service again to play the generated video.
[1148] Prompt for the generative AI model:
[1149] prompt:
[1150] "Generate a text message when a user requests a message like 'Happy Birthday.' The generated text should be natural-sounding and emotive."
[1151] The above is a specific description for carrying out the invention.
[1152] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1153] Step 1:
[1154] A user accesses a cloud service using a smartphone or tablet and logs in. The login information is sent to an authentication server, which returns an authentication result. This identifies the user and allows them to use the system.
[1155] Step 2:
[1156] Users upload photos and audio data of their deceased family members. The photo files (JPEG format) and audio files (WAV format) selected by the user are sent via the Internet to a cloud server, which receives the data and formats it appropriately.
[1157] Step 3:
[1158] The cloud server analyzes the received photo data. First, it uses Amazon Rekognition to analyze the image and extract facial features. This includes information such as the position of the eyes, nose, and mouth, as well as facial contours and skin tone. The extracted facial feature data is used in the next processing step.
[1159] Step 4:
[1160] The cloud server analyzes the received voice data using Pyttsx3 to extract voice features, including patterns such as tone, accent, and speaking habits. The extracted voice feature data is then used to generate the voice message.
[1161] Step 5:
[1162] The text message entered by the user is sent to a cloud server, which uses a generative AI model to generate a text message with natural wording. In this process, a more emotionally charged message is generated based on the user's requested phrases.
[1163] Step 6:
[1164] The cloud server converts the generated text message into a voice message using the Pyttsx3 speech synthesis engine, which generates a voice message that reflects the analyzed voice characteristics, ensuring that the voice message sounds as close as possible to the voice of the deceased person that the user is familiar with.
[1165] Step 7:
[1166] The cloud server generates a video using the generated voice message and the extracted facial features. Using tools such as OpenCV and MoviePy, the video is created so that the face appears to be moving. The video is designed to play back the message entered by the user emotionally.
[1167] Step 8:
[1168] The cloud server saves the generated video in cloud storage and issues a security token to protect the data. This token is used to verify the user's access rights when playing the video. Storing the video in cloud storage ensures secure data management.
[1169] Step 9:
[1170] The user then accesses the cloud service again and plays the generated video in streaming format. The video data is provided from the cloud storage and played in real time on the user's device. At this time, the user's access authority is confirmed using a security token.
[1171] This series of processing steps generates an emotional video message that reproduces the voices and faces of loved ones, allowing users to watch it safely.
[1172] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1173] This invention relates to a system that generates video messages based on photographic and audio data from the deceased's life, and also describes an embodiment that combines an emotion engine. This system allows users to recreate the voice and face of a loved one and receive a heartfelt message that reflects their emotions.
[1174] System configuration
[1175] This system consists of a user's device, a cloud server, cloud storage, and an emotion engine. Each component will be explained below.
[1176] User's device
[1177] The user's device has the function of uploading photo data and audio data to the cloud server, the function of playing streaming videos provided by the cloud server, and the function of collecting the user's emotional data and sending it to the cloud server.
[1178] Cloud Server
[1179] The cloud server receives and analyzes photo and audio data. It also generates text messages based on the analysis results and converts them into audio messages. It also uses an emotion engine to recognize the user's emotions, adjusts the message content and tone as necessary, and generates videos.
[1180] Cloud Storage
[1181] Cloud storage is a database that securely stores the generated videos and provides them to users in streaming format. By using security tokens, user access rights are confirmed and secure data management is achieved.
[1182] Emotion Engine
[1183] The emotion engine analyzes emotional data sent from the user's device and recognizes the user's emotional state in real time. Based on this information, it adjusts the tone and content of generated messages, audio, and video.
[1184] Program processing
[1185] The program in this system plays four main roles: user, terminal, server, and emotion engine. The processing flow and specific operations are explained below.
[1186] First, the user accesses the cloud service and logs in. Next, the user uploads photos and recorded voices of their deceased family members from their device. The device converts these data into appropriate formats (e.g., JPEG and WAV) and sends them to the cloud server.
[1187] Once the server receives the uploaded data, it first performs image analysis. The image analysis module extracts facial features and generates a high-resolution image. Then, the voice analysis module analyzes the audio data and extracts voice characteristics, including dialect and speech patterns.
[1188] Next, based on the text message entered by the user, the server uses a generative AI model to generate a naturally-phrased text message, which is then converted into a voice message using a speech synthesis engine, reflecting the analyzed voice characteristics.
[1189] Furthermore, the emotion engine receives emotion data sent from the user's device and analyzes the user's emotional state. Based on the results, the generated message content and voice tone are adjusted. Specifically, different messages and tones are adopted depending on changes in emotion.
[1190] The server then uses the generated voice message and facial features to generate a video that appears to be moving, which is then stored in cloud storage and a security token is issued to protect the data.
[1191] Finally, the user accesses the cloud service again and plays the generated video in streaming format. The server verifies the user's access privileges and then plays the video securely using a security token.
[1192] Specific examples
[1193] 1. The user logs in to the cloud service and uploads a photo of their deceased mother and her last recorded message (e.g., a birthday message) from their device to the cloud server.
[1194] 2. The server analyzes the uploaded photo, extracts the mother's facial features, and generates a high-resolution image. At the same time, it performs audio analysis to extract the mother's tone of voice, accent, and speaking habits.
[1195] 3. The user requests a video of their mother wishing them a happy birthday. Based on this request, the server uses a generative AI model to generate a text message and converts it into a voice message.
[1196] 4. At this time, the emotion engine analyzes the user's emotions and, for example, if it recognizes that the user is emotional, it adjusts the voice tone to a gentler one.
[1197] 5. The server generates a video of the mother speaking using the generated voice message and the analyzed high-resolution facial features.
[1198] 6. The server stores the generated video in cloud storage and issues a security token to protect the data.
[1199] 7. The user accesses the cloud service to play the generated video and watch the heartwarming message from their deceased mother.
[1200] The above is a specific description of an embodiment of the "My Album in Action" system that combines the emotion engine of the present invention.
[1201] The processing flow will be explained below.
[1202] Step 1: Log in and upload data
[1203] 1. The user accesses the cloud service and logs in by entering their user ID and password.
[1204] 2. The device sends the login information to the cloud server for authentication.
[1205] 3. The server verifies the authentication information and, if successful, grants access to the user.
[1206] 4. The user selects the photo and audio data of the deceased family member on the device and uploads it to the cloud server.
[1207] 5. The device converts the selected data into the appropriate format (e.g., JPEG and WAV) and sends it to the cloud server.
[1208] 6. The server stores the received photo and audio data in cloud storage.
[1209] Step 2: Data analysis
[1210] 1. The server passes the uploaded photo data to the image analysis module.
[1211] 2. The server uses an image analysis module to extract facial features (such as the position of the eyes, nose, and mouth) of family members from the photo.
[1212] 3. The server uses super-resolution technology to generate a high-resolution facial image based on the extracted feature points.
[1213] 4. The server passes the uploaded voice data to the voice analysis module.
[1214] 5. The server uses a speech analysis module to extract tone, rate, accent, and dialect and diction patterns from the speech data.
[1215] Step 3: Collect and analyze emotion data
[1216] 1. The user interacts with the cloud service interface and gives permission to collect emotion data.
[1217] 2. The device collects emotional data such as the user's facial expressions and tone of voice and sends it to a cloud server.
[1218] 3. The server passes the received emotion data to the emotion engine and analyzes the user's emotional state.
[1219] Step 4: Message Generation
[1220] 1. The user enters the content of the message they want to generate (e.g., "Happy Birthday") through the cloud service interface.
[1221] 2. The server passes the input text message to a generative AI model (e.g., GPT-4) to generate a naturally phrased text message.
[1222] 3. The server uses a speech synthesis engine to generate a voice message that sounds like the family member's voice based on the generated text message and the extracted voice characteristics.
[1223] 4. The server adjusts the tone of the voice and the content of the message based on the user's emotional data provided by the emotion engine.
[1224] Step 5: Video Generation
[1225] 1. The server uses a high-resolution facial image and the generated voice message to generate a video that matches the facial movements and voice using ravioli generation technology (e.g., Facial Animation Technology).
[1226] 2. The server stores the generated video in cloud storage and applies appropriate encryption processes.
[1227] Step 6: Streaming
[1228] 1. The user clicks the play button to request playback of the video generated from the cloud service.
[1229] 2. The server verifies the user's access rights, issues a security token, and sends it to the user's device.
[1230] 3. The device receives the security token and sends a request to the server again to start streaming the video.
[1231] 4. The server verifies the security token and sends the video data to the device in streaming format.
[1232] 5. The device plays the received video data, allowing the user to watch the video.
[1233] The above is a description of the specific operations divided into processing steps of the "Moving My Album" system.
[1234] Example 2
[1235] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1236] In recent years, there has been a demand for preserving memories of deceased loved ones and receiving them as emotionally warm messages. However, conventional technologies have struggled to reproduce the voice and appearance of the deceased in a natural way. It has also been difficult to generate messages and videos that correspond to the user's emotional state. Furthermore, ensuring the security of the generated data and managing access rights are also important issues.
[1237] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1238] In this invention, the server includes means for acquiring image data of the deceased person, means for acquiring audio data of the deceased person, means for analyzing the acquired image data and audio data and extracting facial and vocal characteristics of the person, means for generating a natural language message based on the extracted facial and vocal characteristics and converting the message into synthetic speech, means for analyzing the user's emotional data and reflecting the emotional state in the generation process, means for generating a video using the generated synthetic speech and the analyzed facial characteristics, and means for saving the generated video in a storage device and making it playable in streaming format. This makes it possible to reproduce the voice and appearance of the deceased person with high accuracy, and to easily generate and safely manage warm video messages that correspond to the user's emotional state.
[1239] "Image data from before death" refers to photographs and image files taken of a person before their death.
[1240] "Pre-death acoustic data" refers to voices and audio files recorded during a person's lifetime.
[1241] "Means of acquisition" refers to the functions and interfaces that allow users to select and upload data.
[1242] "Means for analysis" refers to the algorithms and software used to analyze the acquired data and extract the necessary information.
[1243] "Facial features" refer to identifiable parts or attributes of a person's face that are extracted from an image.
[1244] "Voice features" refer to identifiable parts or attributes such as tone, accent, and catchphrases extracted from voice data.
[1245] A "natural language message" refers to a text message that is written in a way that is natural for humans to understand.
[1246] "Synthetic voice" refers to data generated as voice based on a text message.
[1247] "Emotional Data" refers to information collected to describe a user's emotional state.
[1248] The "generation process" refers to a series of steps to generate messages and videos based on analyzed data.
[1249] "Storage device" refers to a storage system for safely storing the generated video.
[1250] "Streaming format" refers to a format that allows users to watch videos in real time over the Internet.
[1251] This invention relates to a system that generates video messages based on image and audio data from a person's life, and by combining it with an emotion engine, it allows users to receive the messages as warm messages. This system consists of a user's device, a cloud server, cloud storage, and an emotion engine.
[1252] System configuration
[1253] User's device
[1254] The user's device has the function of uploading image data and audio data to the cloud server, the function of playing streaming videos provided by the cloud server, and the function of collecting the user's emotional data and sending it to the cloud server.
[1255] Cloud Server
[1256] The cloud server is equipped with software and hardware for receiving and analyzing image and acoustic data. The analysis uses OpenCV, an image processing software, and NVIDIA NeMo, a voice analysis software. Based on the text message entered by the user, a generative AI model (e.g., a generative AI model) is used to generate a text message with natural wording.
[1257] The server converts the generated text message into a voice message using a speech synthesis engine, incorporating the analyzed voice characteristics. This process uses the speech synthesis technology of Azure Cognitive Services.
[1258] Furthermore, the Emotion Analysis Engine analyzes the user's emotional data and adjusts the generated message content and voice tone based on the user's emotional state. For example, if the user is emotional, the voice tone will be made gentler.
[1259] The server uses the generated voice message and high-resolution facial features to generate a video that appears to show a moving face using DeepFaceLab. The video is encoded using FFmpeg. The video is then stored in cloud storage, where data is protected by a security token.
[1260] Cloud Storage
[1261] Cloud storage is a database that securely stores the generated videos and provides them to users in streaming format. By using security tokens, user access rights are confirmed and secure data management is achieved.
[1262] Emotion Engine
[1263] The emotion engine analyzes emotional data sent from the user's device and recognizes the user's emotional state in real time. Based on this information, it adjusts the tone and content of generated messages, audio, and video.
[1264] Specific examples
[1265] A user logs in to the cloud service and uploads a photo (JPEG format) of a deceased family member and audio data (WAV format) from their device to the cloud server. The server analyzes the uploaded data using OpenCV and NVIDIA NeMo to extract facial and vocal characteristics. Next, the generative AI model converts the user's requested message (e.g., "Happy Birthday") into natural-sounding text, which is then converted into a voice message using Azure Cognitive Services' speech synthesis engine. The emotion engine also analyzes the user's emotional data and adjusts the message content and voice tone. Finally, the server generates a video using DeepFaceLab and saves it in cloud storage. The user can then access the cloud service again to play the generated video in streaming format.
[1266] Example prompt: "Generate a sweet message from my late mother wishing me a happy birthday."
[1267] The above is a specific description of an embodiment of the "My Album in Action" system that combines the emotion engine of the present invention.
[1268] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1269] Step 1:
[1270] A user logs in
[1271] A user accesses a cloud service using a browser or application and enters a user ID and password, which authenticates the user and allows them to log in to the cloud service.
[1272] Input: User ID, Password
[1273] Output: Authentication token, indicating successful login
[1274] Step 2:
[1275] The user uploads photos and audio data.
[1276] The user selects a photo (JPEG format) and audio data (WAV format) of the deceased family member from the device and clicks the upload button. The device converts the data into the appropriate format (JPEG and WAV) and sends it to the cloud server.
[1277] Input: Photo data (JPEG), audio data (WAV)
[1278] Output: Photo and audio data uploaded to the cloud server
[1279] Step 3:
[1280] The server receives the photo and audio data.
[1281] The server receives the photo and audio data sent from the device and temporarily stores it in the specified storage. At this stage, the server checks the integrity of the data to ensure there is no unauthorized data.
[1282] Input: Photo data and audio data uploaded from the device
[1283] Output: Temporarily saved photo and audio data
[1284] Step 4:
[1285] The server performs image analysis
[1286] The server analyzes the received photo data using OpenCV, specifically applying a facial recognition algorithm to extract facial features.
[1287] Input: Temporarily saved photo data
[1288] Output: Facial features (position, shape, landmark points, etc.)
[1289] Step 5:
[1290] The server performs voice analysis
[1291] The server uses NVIDIA NeMo to analyze the received voice data, extracting voice tone, accent, and speaking habits.
[1292] Input: Temporarily saved audio data
[1293] Output: Voice characteristics (tone, accent, speech patterns)
[1294] Step 6:
[1295] The server generates a text message
[1296] Based on a prompt phrase entered by the user (e.g., "Say happy birthday"), the server uses a generative AI model (e.g., GPT-3) to generate a natural-sounding text message.
[1297] Input: The prompt text entered by the user
[1298] Output: The generated text message
[1299] Step 7:
[1300] The server converts it into a voice message
[1301] Based on the generated text message, the server uses Azure Cognitive Services' speech synthesis engine to convert it into a voice message, which reflects the analyzed voice characteristics.
[1302] Input: Generated text message, voice features
[1303] Output: Synthesized voice message
[1304] Step 8:
[1305] The emotion engine analyzes the emotion data
[1306] The emotion engine analyzes emotion data sent from the user's device and recognizes the user's emotional state in real time.
[1307] Input: User emotion data (facial expressions, tone of voice, etc.)
[1308] Output: Emotional state analysis result
[1309] Step 9:
[1310] The server generates a video based on the generated data.
[1311] The server uses DeepFaceLab to generate a video that appears to show a moving face based on the generated voice message and high-resolution facial features, and encodes the video using FFmpeg.
[1312] Input: Synthesized voice message, facial features
[1313] Output: Generated video file
[1314] Step 10:
[1315] The server saves the video to cloud storage.
[1316] The generated video is stored in cloud storage. A security token is issued to verify the user's access rights and protect the data.
[1317] Input: Generated video file
[1318] Output: Video stored in cloud storage, security token
[1319] Step 11:
[1320] Playing user-generated videos
[1321] The user then accesses the cloud service again and plays the video. The server then verifies the user's access privileges and uses the security token to play the video securely in streaming format.
[1322] Input: Security token, saved video file
[1323] Output: Streamed video
[1324] The above are the detailed processing steps of this system.
[1325] (Application example 2)
[1326] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1327] In order to recreate the memory of the deceased and move the bereaved family, it is desirable to provide more realistic and emotionally relevant video messages. However, current technology has problems in that it is difficult to adjust the tone and content of the message according to the user's emotional state, and the generated videos lack a sense of realism. In addition, there is a lack of appropriate display methods for viewing moving videos in real time. There is a need to solve these problems.
[1328] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1329] In this invention, the server includes means for uploading photo data of the person before death, means for uploading audio data of the person before death, means for analyzing the uploaded photo data and audio data and extracting facial and vocal characteristics of the person, means for generating a text message based on the extracted facial and vocal characteristics and converting the text message into an audio message, means for generating a video using the generated audio message and the analyzed facial characteristics, means for storing the generated video in cloud storage and making it playable in streaming format, means for collecting emotional data and analyzing the data to adjust the tone of the text message or audio message, and means for displaying the generated video using a head-mounted display, thereby making it possible to provide a realistic video message that matches the emotional state of the user.
[1330] "Photo data from before death" refers to image files of a person taken before death.
[1331] "Voice data from before death" refers to an audio file of a person's voice recorded before death.
[1332] The "means for uploading" refers to a method and apparatus for transmitting data from a terminal to a cloud server.
[1333] "Means for analyzing" refers to methods and devices that process uploaded data to extract information.
[1334] "Facial features" are characteristic parts of a person's face that are extracted from the analyzed image data.
[1335] "Voice features" are characteristics of a person's voice extracted from analyzed voice data.
[1336] A "means for generating a text message" is a method and device that creates a sentence based on input information.
[1337] A "means for converting to voice message" is a method and apparatus for converting a text message to voice data.
[1338] The "means for generating video" refers to a method and device for creating a video file based on image data and audio data.
[1339] "Cloud storage means" refers to a method and device for storing data on a remote server.
[1340] "Means for enabling playback in streaming format" refers to a method and apparatus for playing back stored video data in real time via the Internet.
[1341] A "means for collecting emotional data" is a method and apparatus for collecting a user's emotional state.
[1342] A "means for analyzing emotion data" is a method and apparatus for processing collected emotion data to make sense of the information.
[1343] "Means for adjusting the tone of text and voice messages" refers to a method and apparatus for changing the tone of message content based on analyzed emotional data.
[1344] A "head-mounted display" is a display device worn on the head, and is used to display videos and images.
[1345] This invention relates to a system for generating moving video messages using photographic and audio data from a person's lifetime. Furthermore, by collecting and analyzing the user's emotional data, the system adjusts the message content and tone to generate more personalized and emotional videos. A specific embodiment of this system is described below.
[1346] System configuration
[1347] This system consists of a user's device, a cloud server, cloud storage, an emotion engine, and a head-mounted display.
[1348] User's device
[1349] The user's device has the function of uploading photo data and audio data to the cloud server, the function of playing streaming videos provided by the cloud server, and the function of collecting the user's emotional data and sending it to the cloud server.
[1350] Cloud Server
[1351] The cloud server receives and analyzes photo and audio data. It also generates text messages based on the analysis results and converts them into audio messages. It also uses an emotion engine to recognize the user's emotions, adjusts the message content and tone as necessary, and generates videos.
[1352] Cloud Storage
[1353] Cloud storage is a database that securely stores the generated videos and provides them to users in streaming format. By using security tokens, user access rights are confirmed and secure data management is achieved.
[1354] Emotion Engine
[1355] The emotion engine analyzes emotional data sent from the user's device and recognizes the user's emotional state in real time. Based on this information, it adjusts the tone and content of generated messages, audio, and video.
[1356] head-mounted display
[1357] A head-mounted display is a device worn on the head that displays images to the user. It displays video received from a cloud server in real time, providing the user with a realistic viewing experience.
[1358] Program processing
[1359] The program in this system mainly plays four main roles: user, terminal, server, and emotion engine.
[1360] First, the user accesses the cloud service and logs in. The user uploads photos and audio data from their device to the cloud server. The device converts the data into an appropriate format (e.g., JPEG and WAV) and sends it to the cloud server.
[1361] Once the cloud server receives the uploaded data, it first performs image analysis. The image analysis module extracts facial features and generates a high-resolution image. The voice analysis module then analyzes the audio data and extracts voice characteristics, including dialect and speech patterns.
[1362] Next, based on the text message entered by the user, the cloud server uses a generative AI model to generate a naturally-phrased text message, which is then converted into a voice message using a speech synthesis engine, reflecting the analyzed voice characteristics.
[1363] Furthermore, the emotion engine receives emotion data sent from the user's device and analyzes the user's emotional state. Based on the results, it adjusts the generated message content and voice tone. Specifically, it adopts different messages and tones depending on changes in the user's emotions.
[1364] The cloud server then uses the generated voice message and facial features to generate a video that appears to be moving, which is then stored in cloud storage and a security token is issued to protect the data.
[1365] Finally, the user wears a head-mounted display and watches the generated video. The cloud server verifies the user's access rights and then plays the video securely using a security token.
[1366] Specific examples
[1367] A user visits a funeral home and uploads a photo of their deceased father and a recorded voice message (e.g., a birthday message). The cloud server analyzes the uploaded photo, extracts the father's facial features, and generates a high-resolution image. At the same time, it performs voice analysis to extract the father's tone of voice, accent, and catchphrases.
[1368] A user can request a generative AI model to generate a video of their father saying, "Take care." The emotion engine then analyzes the user's emotions and adjusts the tone of the voice to be gentler if it detects that the user is emotional, for example.
[1369] The cloud server then uses the generated voice message and analyzed high-resolution facial features to generate a video of the father speaking, which is then stored in cloud storage and a security token is issued to protect the data.
[1370] Finally, the user can wear a head-mounted display to watch the generated video and receive a heartwarming message from their deceased father.
[1371] Example prompts for generative AI models
[1372] Below is an example of a prompt sentence for the generative AI model (generative AI model).
[1373] "Dear visitors, wipe away your tears and upload your favorite photos and audio recordings here. We will recreate the voice and face of your loved one and deliver a heartwarming video message. Let's share this special moment together. We will recreate phrases such as 'Thank you' and 'Take care' in the video."
[1374] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1375] Step 1:
[1376] The user accesses the cloud service and logs in.
[1377] Input: User authentication information (user ID, password)
[1378] Output: Authentication result (success / failure)
[1379] Specific behavior:
[1380] The user's device sends authentication information to the cloud server, which checks it against the database and, if authentication is successful, proceeds to the next step.
[1381] Step 2:
[1382] Users upload photo data and audio data from their pre-death lives from their devices to a cloud server.
[1383] Input: Photo data from before death (JPEG), audio data from before death (WAV)
[1384] Output: Data stored on the cloud server
[1385] Specific behavior:
[1386] Users simply select photos and audio data from their device and click the upload button. The device then sends the data to the cloud server, which stores it securely.
[1387] Step 3:
[1388] The cloud server analyzes the uploaded photo data.
[1389] Input: Photo data from before death (JPEG)
[1390] Output: Facial feature data (facial feature points, facial expressions)
[1391] Specific behavior:
[1392] The server uses an image analysis module (e.g., OpenCV) to extract facial features from the photo data and generate high-resolution facial images.
[1393] Step 4:
[1394] The cloud server analyzes the voice data from the person's life.
[1395] Input: Audio data from before death (WAV)
[1396] Output: Voice characteristics data (tone of voice, accent, catchphrases)
[1397] Specific behavior:
[1398] The server uses a voice analysis module (e.g., librosa) to extract voice features from the audio data.
[1399] Step 5:
[1400] The cloud server generates a message using a generative AI model based on the text message entered by the user.
[1401] Input: User-entered text messages, facial feature data, and voice feature data
[1402] Output: Naturally-phrased text message
[1403] Specific behavior:
[1404] The server sends a prompt to a generative AI model (e.g., GPT-3) to generate a naturally phrased text message.
[1405] Step 6:
[1406] The cloud server converts the generated text message into a voice message.
[1407] Input: Naturally-phrased text messages, voice feature data
[1408] Output: Voice message (audio file)
[1409] Specific behavior:
[1410] The server uses a speech synthesis engine (e.g., Watson Text to Speech) to convert the text message into speech, and the generated voice message reflects the voice characteristics.
[1411] Step 7:
[1412] The emotion engine analyzes the user's emotion data and transmits the emotional state to the cloud server.
[1413] Input: User emotion data (real-time emotional state)
[1414] Output: Emotion analysis result (emotional state)
[1415] Specific behavior:
[1416] The emotion engine analyzes the emotion data collected from the user's device and sends the results to a cloud server.
[1417] Step 8:
[1418] The cloud server adjusts the message content and voice tone based on the results of emotion analysis.
[1419] Input: Emotion analysis results, voice message, facial feature data
[1420] Output: Adjusted message content, audio tone
[1421] Specific behavior:
[1422] The results of the emotion engine are reflected and the tone of the generated text and voice messages is adjusted to match the user's emotions.
[1423] Step 9:
[1424] The cloud server generates a video using the generated voice message and facial features.
[1425] Input: Facial feature data, voice message
[1426] Output: Video file
[1427] Specific behavior:
[1428] The server uses a video generation module (for example, Adobe After Effects API) to generate a video in which a person's face moves based on the voice message and facial feature data.
[1429] Step 10:
[1430] The cloud server stores the generated video in cloud storage and issues a security token to protect the data.
[1431] Input: Video file
[1432] Output: Saved video file, security token
[1433] Specific behavior:
[1434] The server stores the video file in cloud storage (e.g., AWS S3) and issues a security token for access control.
[1435] Step 11:
[1436] The user wears a head-mounted display and watches videos generated from a cloud server.
[1437] Input: Security Token
[1438] Output: Video Streaming
[1439] Specific behavior:
[1440] The user accesses the cloud service again and watches the video in real time using a head-mounted display (e.g., HoloLens, Oculus Rift). The cloud server uses a security token to verify the user's access privileges and plays the video securely.
[1441] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1442] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1443] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1444] [Fourth embodiment]
[1445] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1446] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1447] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1448] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1449] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1450] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1451] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1452] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1453] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1454] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1455] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1456] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1457] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1458] This invention provides a system for generating video messages based on photographic and audio data of loved ones, and describes an embodiment of the system. By implementing this system, users can recreate the voices and faces of loved ones and receive them as heartwarming messages.
[1459] System configuration
[1460] This system consists of a user terminal, a cloud server, and a storage system. Each component is explained below.
[1461] User's device
[1462] The user's device has the function of uploading photo data and audio data to the cloud server, and also has the function of playing streaming video provided by the cloud server.
[1463] Cloud Server
[1464] The cloud server has the functions of receiving and analyzing photo data and audio data, generating text messages based on the analysis results and converting them into audio messages, and generating videos based on the generated audio messages and facial features, and storing them in cloud storage after ensuring appropriate security.
[1465] Cloud Storage
[1466] Cloud storage is a database that securely stores the generated videos and provides them to users in streaming format. By using security tokens, user access rights are confirmed and secure data management is achieved.
[1467] Program processing
[1468] The program in this system plays three main roles: user, terminal, and server. The process flow and specific operations are explained below.
[1469] First, the user accesses the cloud service and logs in. Next, the user uploads photos and recorded voices of their deceased family members from their device. The device receives this data, converts it into an appropriate format (e.g., JPEG and WAV), and sends it to the cloud server.
[1470] The server then receives the uploaded data and performs image analysis. The image analysis module extracts facial features and generates a high-resolution image. The audio analysis module then analyzes the audio data and extracts voice characteristics, including dialect and speech patterns.
[1471] Next, based on the text message entered by the user, the server uses a generative AI model to generate a naturally-phrased text message, which is then converted into a voice message using a speech synthesis engine, reflecting the analyzed voice characteristics.
[1472] The server then uses the generated voice message and facial features to generate a video that appears to be moving, which is then stored in cloud storage and a security token is issued to protect the data.
[1473] Finally, the user accesses the cloud service again and plays the generated video in streaming format. The server verifies the user's access privileges and then plays the video securely using a security token.
[1474] Specific examples
[1475] 1. The user logs in to the cloud service and uploads a photo of their deceased mother and her last recorded message (e.g., a birthday message) from their device to the cloud server.
[1476] 2. The server analyzes the uploaded photo, extracts the mother's facial features, and generates a high-resolution image. At the same time, it performs audio analysis to extract the mother's tone of voice, accent, and speaking habits.
[1477] 3. The user requests a video of their mother wishing them a happy birthday. Based on this request, the server uses a generative AI model to generate a text message and converts it into a voice message.
[1478] 4. The server generates a video of the mother speaking using the generated voice message and the analyzed high-resolution facial features.
[1479] 5. The server stores the generated video in cloud storage and issues a security token to protect the data.
[1480] 6. The user accesses the cloud service to play the generated video and watch the heartwarming message from their deceased mother.
[1481] The above is a specific description of the embodiment of the "Animated My Album" system of the present invention.
[1482] The processing flow will be explained below.
[1483] Step 1: Log in and upload data
[1484] 1. The user accesses the cloud service and logs in by entering their user ID and password.
[1485] 2. The device sends the login information to the cloud server for authentication.
[1486] 3. The server verifies the authentication information and, if successful, grants access to the user.
[1487] 4. The user selects the photo and audio data of the deceased family member on the device and uploads it to the cloud server.
[1488] 5. The device converts the selected data into the appropriate format (e.g., JPEG and WAV) and sends it to the cloud server.
[1489] 6. The server stores the received photo and audio data in cloud storage.
[1490] Step 2: Data analysis
[1491] 1. The server passes the uploaded photo data to the image analysis module.
[1492] 2. The server uses an image analysis module to extract facial features (such as the position of the eyes, nose, and mouth) of family members from the photo.
[1493] 3. The server uses super-resolution technology to generate a high-resolution facial image based on the extracted feature points.
[1494] 4. The server passes the uploaded voice data to the voice analysis module.
[1495] 5. The server uses a speech analysis module to extract tone, rate, accent, and dialect and diction patterns from the speech data.
[1496] Step 3: Message Generation
[1497] 1. The user enters the content of the message they want to generate (e.g., "Happy Birthday") through the cloud service interface.
[1498] 2. The server passes the input text message to a generative AI model (e.g., GPT-4) to generate a naturally phrased text message.
[1499] 3. The server uses a speech synthesis engine to generate a voice message that sounds like the family member's voice based on the generated text message and the extracted voice characteristics.
[1500] Step 4: Video Generation
[1501] 1. The server uses a high-resolution facial image and the generated voice message to generate a video that matches the facial movements and voice using ravioli generation technology (e.g., Facial Animation Technology).
[1502] 2. The server stores the generated video in cloud storage and applies appropriate encryption processes.
[1503] Step 5: Streaming
[1504] 1. The user clicks the play button to request playback of the video generated from the cloud service.
[1505] 2. The server verifies the user's access rights, issues a security token, and sends it to the user's device.
[1506] 3. The device receives the security token and sends a request to the server again to start streaming the video.
[1507] 4. The server verifies the security token and sends the video data to the device in streaming format.
[1508] 5. The device plays the received video data, allowing the user to view the video.
[1509] The above is a description of the specific operations divided into each processing step.
[1510] Example 1
[1511] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1512] Conventional commemorative video generation systems that use photographs and audio data have difficulty adequately reproducing the face and voice of the deceased as desired by users, and there are also issues with the quality and realism of the generated videos. As a result, it is difficult for users to receive the heartwarming messages of their loved ones in a realistic way. Furthermore, there are insufficient aspects of data protection and security, which increases the risk of personal information being leaked.
[1513] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1514] In this invention, the server includes: means for a user to access digital data storage and upload photo data and audio data; means for converting the uploaded photo data and audio data into an appropriate format; means for transmitting the converted photo data and audio data to a cloud server; means for the cloud server to analyze the received photo data and audio data and extract facial and vocal characteristics of a person; means for generating a text message using a generative AI model based on the extracted facial and vocal characteristics and converting the text message into an audio message; means for generating a video using the generated audio message and the analyzed facial characteristics; and means for saving the generated video in cloud storage and making it playable in streaming format. This allows users to safely and easily generate high-quality, realistic commemorative videos and receive warm messages from their loved ones in real time.
[1515] "User" refers to an individual who uses the system and uploads photo data and audio data.
[1516] "Digital data storage" refers to a system that stores and makes accessible information electronically.
[1517] "Photo data" refers to still images saved in a format such as JPEG.
[1518] "Audio Data" refers to recorded audio information stored in a format such as WAV.
[1519] "Format" refers to a standard or method for organizing data into a specific form.
[1520] "Cloud server" refers to a remote server that stores, processes, and manages data over the Internet.
[1521] "Analysis" refers to the process of examining received data in detail and extracting features and patterns.
[1522] "Facial features of a person" refers to information about the position and shape of the eyes, nose, mouth, etc. extracted from photographic data.
[1523] "Voice characteristics" refers to information such as voice tone, pronunciation habits, etc. extracted from audio data.
[1524] A "generative AI model" refers to an artificial intelligence model that uses machine learning algorithms to generate new data and information.
[1525] "Text Message" means a message expressed as written words or sentences.
[1526] "Voice message" refers to a message that is an audio representation of a text message.
[1527] "Video" refers to a media format that reproduces movement and sound through the combination of multiple still images and audio.
[1528] "Cloud storage" refers to remote storage that stores and manages data via the Internet.
[1529] "Streaming format" refers to a format in which data is played in real time without downloading.
[1530] "Security token" refers to identification information issued to verify access rights and prevent unauthorized access.
[1531] This invention provides a system for generating video messages based on photographic and audio data of loved ones, and describes an embodiment of the system. By implementing this system, users can recreate the voices and faces of loved ones and receive them as heartwarming messages.
[1532] System configuration
[1533] This system consists of a user terminal, a cloud server, and a cloud storage system. Each component is explained below.
[1534] User's device
[1535] The user's device has the function of uploading photo data and audio data to the cloud server, and also has the function of playing streaming video provided by the cloud server.
[1536] Cloud Server
[1537] The cloud server receives and analyzes photo and audio data. It also generates text messages based on the analysis results and converts them into audio messages. It also generates videos based on the generated audio messages and facial features and saves them in cloud storage. A generative AI model is used to generate text messages with natural wording.
[1538] Cloud Storage
[1539] Cloud storage is a database that securely stores the generated videos and provides them to users in streaming format. By using security tokens, user access rights are confirmed and secure data management is achieved.
[1540] What the program does
[1541] The program in this system plays three main roles: user, terminal, and server. The specific operation is explained below.
[1542] First, users access the cloud service and log in. Then, they upload photos and audio recordings of their deceased family members from their devices. The devices receive these data, convert them into appropriate formats (e.g., JPEG and WAV), and send them to the cloud server.
[1543] The server then receives the uploaded data and first performs image analysis. The image analysis module extracts facial features and generates high-resolution images. The voice analysis module then analyzes the audio data and extracts voice characteristics, including dialect and speech patterns.
[1544] Next, based on the text message entered by the user, the server uses a generative AI model to generate a naturally-phrased text message. The text message is then converted into a voice message using a speech synthesis engine, reflecting the analyzed voice characteristics. An example of a prompt is a request such as, "Please generate a message for my mother celebrating her birthday using her voice and characteristics."
[1545] The server then uses the generated voice message and facial features to generate a video that appears to be moving, which is then stored in cloud storage and a security token is issued to protect the data.
[1546] Finally, the user accesses the cloud service again and plays the generated video in streaming format. The server verifies the user's access privileges and then plays the video securely using a security token.
[1547] Specific examples
[1548] The user logs in to the cloud service and uploads a photo of their deceased mother and her last recorded message (e.g., a birthday message) from their device to the cloud server.
[1549] The server analyzes the uploaded photo, extracts the mother's facial features, and generates a high-resolution image, while also performing audio analysis to extract the mother's tone of voice, accent, and speaking habits.
[1550] A user requests a video of their mother wishing them a happy birthday. Based on this request, the server uses a generative AI model to generate a text message and convert it into an audio message.
[1551] The server uses the generated audio message and analyzed high-resolution facial features to generate a video of the mother speaking.
[1552] The server stores the generated video in cloud storage and issues a security token to protect the data.
[1553] The user accesses the cloud service to play the generated video and watch the heartwarming message from their deceased mother.
[1554] The above is a specific description of the embodiment of the "Animated My Album" system of the present invention.
[1555] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1556] Step 1:
[1557] A user logs in to the cloud service and uploads photos and audio data. The inputs include the user's account information, the JPEG photo data to be uploaded, and the WAV audio data. The output is then prepared on the user's device.
[1558] Step 2:
[1559] The device sends the photo data and audio data received from the user to the server. The inputs are uploaded JPEG photo data and WAV audio data. Specifically, the device sends this data to the cloud server using an HTTP POST request, and the cloud server receives the data as output.
[1560] Step 3:
[1561] The server receives the photo data and audio data and begins analyzing it. The input is JPEG photo data and WAV audio data. Specifically, the server uses an image analysis module to extract facial features from the photo data. A face detection algorithm is used for this analysis, and facial features are obtained as the output.
[1562] Step 4:
[1563] The server analyzes the audio data and extracts voice characteristics. The input is audio data in WAV format. Specifically, the server uses acoustic analysis tools to extract voice characteristics such as tone, accent, and catchphrases from the audio data. The voice characteristics are obtained as output.
[1564] Step 5:
[1565] The server converts the text message entered by the user into a natural-sounding text message based on the generative AI model. The inputs include a text message provided by the user and a prompt: "Generate a message to celebrate my mother's birthday using her voice and characteristics." The output is a natural-sounding text message generated by the generative AI model.
[1566] Step 6:
[1567] The server converts the text message into a voice message using a speech synthesis engine. The inputs are the generated text message and the analyzed voice characteristics. In concrete terms, the server converts the text message into voice using the speech synthesis engine, and the output is the voice message.
[1568] Step 7:
[1569] The server generates a video based on the generated voice message and facial features. The inputs are the voice message and facial features. Specifically, it starts the video generation engine, synchronizes the facial movements with the voice, and generates a video that appears to be speaking as output.
[1570] Step 8:
[1571] The server stores the generated video in cloud storage and issues a security token. The input is the generated video data. Specifically, the server encrypts the video data, stores it in cloud storage, and issues a security token to verify the user's access rights. The output is the securely stored video data and a security token.
[1572] Step 9:
[1573] The user accesses the cloud service again and plays the generated video in streaming format. The inputs are the user's account information and a security token. Specifically, the server verifies the security token and grants the user permission to play the video. The output is the video played in streaming format and provided to the user.
[1574] (Application example 1)
[1575] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1576] Current memorial videos and photo albums are limited to still images or simple slideshows, making it difficult for bereaved families to visually and powerfully reproduce the voices and facial movements of their loved ones. Furthermore, privacy and security concerns pose a demand for a reliable system for safely creating, storing, and playing video content. Therefore, a video message generation system that can convey more personal emotions and deepen emotional connections is needed.
[1577] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1578] In this invention, the server is installed on a smartphone or tablet and includes means for generating, saving, and playing back video messages of memories based on the voices and photos of loved ones, means for analyzing uploaded photo and audio data to extract facial and vocal characteristics of people, and means for saving the generated videos in cloud storage and making them playable in streaming format. This makes it possible to generate memorial videos of loved ones and strengthen emotional ties in a safe and reliable manner.
[1579] "Photographic data from before death" refers to digital data of photographs taken of the deceased person before death.
[1580] "Voice data from before death" refers to digital data of the voice of a deceased person recorded before death.
[1581] "Uploading means" refers to a method or device for transmitting data owned by a user to a cloud server via the Internet.
[1582] "Facial features" are characteristics or specific points of a face extracted by facial recognition technology and are used to identify an individual.
[1583] "Voice features" are characteristics of a voice extracted by voice recognition technology and are used to identify an individual's voice.
[1584] A "text message" is character data in a natural language, and is generated as a message from a loved one.
[1585] A "voice message" is data obtained by converting a generated text message into voice.
[1586] The "means for generating moving images" refers to a method or device that converts still images into dynamic moving images based on facial features and voice messages.
[1587] "Cloud storage" is a data storage area on a remote server for storing files on the Internet.
[1588] "Streaming format" is a technology that allows data to be played in real time while being downloaded.
[1589] A "smartphone or tablet" is a type of mobile device that has advanced computing power and internet connectivity.
[1590] "Memory Video Message" is an emotional video message created using the face and voice of the deceased.
[1591] A "security token" is temporary authentication information used to verify access privileges.
[1592] A "dialect and speech database" is a collection of data that records the speech characteristics of a particular region or particular individual.
[1593] This invention provides a system for generating video messages based on photographic and audio data of loved ones, allowing users to receive emotionally rich messages by recreating the voices and faces of loved ones.
[1594] System configuration
[1595] This system consists of a user device, a cloud server, and cloud storage. Each component is explained below.
[1596] User's device
[1597] The user's device is a smartphone or tablet, which has the function of uploading photo data and audio data to a cloud server.
[1598] It also has the ability to play streaming videos provided by cloud servers.
[1599] Cloud Server
[1600] The cloud server has the function of receiving and analyzing the photo data and audio data.
[1601] It also generates text messages based on the analysis results and converts them into voice messages, and generates videos based on the generated voice messages and facial features, saving them to cloud storage.
[1602] The main software used includes Amazon Rekognition for image analysis, Pyttsx3 for speech synthesis, and OpenCV and MoviePy for video generation.
[1603] Cloud Storage
[1604] Cloud storage is a database that safely stores the generated videos and provides them to users in streaming format.
[1605] By using security tokens, user access rights can be confirmed and data management can be ensured safely.
[1606] Program processing
[1607] On the user's device:
[1608] Users access the cloud service, log in, and upload photos and audio data of their precious family members from their devices. At this time, the photo data is converted to JPEG format, and the audio data is converted to WAV format.
[1609] Cloud Server:
[1610] After receiving the uploaded data, an image analysis module extracts facial features and generates a high-resolution image.
[1611] A speech analysis module analyzes the audio data and extracts vocal characteristics, in this case reflecting specific patterns using a database of dialects and accents.
[1612] Based on the text message entered by the user, a generative AI model is used to generate a naturally-phrased text message, which is then converted into a voice message using a speech synthesis engine, incorporating the analyzed voice characteristics.
[1613] Finally, the generated voice message and facial features are used to generate a video that makes the face appear to be moving.
[1614] Cloud Storage:
[1615] The generated video is stored in cloud storage and a security token is issued to protect the data.
[1616] When a user plays a video, the cloud server checks the user's access rights and then provides the video securely in streaming format.
[1617] Specific examples
[1618] Example 1:
[1619] A user logs in to a cloud service and uploads a photo of their deceased mother and her last recorded message (e.g., a birthday message). The cloud server analyzes the received image data using Amazon Rekognition to extract the mother's facial features. Next, the audio data is analyzed using Pyttsx3 to extract the mother's tone and accent. When the user requests a message saying "Happy Birthday," the generative AI model generates a naturally phrased text message, which is then converted into an audio message using Pyttsx3. Finally, a video of the mother speaking is generated using OpenCV and MoviePy and saved to cloud storage. The user then accesses the cloud service again to play the generated video.
[1620] Prompt for the generative AI model:
[1621] prompt:
[1622] "Generate a text message when a user requests a message like 'Happy Birthday.' The generated text should be natural-sounding and emotive."
[1623] The above is a specific description for carrying out the invention.
[1624] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1625] Step 1:
[1626] A user accesses a cloud service using a smartphone or tablet and logs in. The login information is sent to an authentication server, which returns an authentication result. This identifies the user and allows them to use the system.
[1627] Step 2:
[1628] Users upload photos and audio data of their deceased family members. The photo files (JPEG format) and audio files (WAV format) selected by the user are sent via the Internet to a cloud server, which receives the data and formats it appropriately.
[1629] Step 3:
[1630] The cloud server analyzes the received photo data. First, it uses Amazon Rekognition to analyze the image and extract facial features. This includes information such as the position of the eyes, nose, and mouth, as well as facial contours and skin tone. The extracted facial feature data is used in the next processing step.
[1631] Step 4:
[1632] The cloud server analyzes the received voice data using Pyttsx3 to extract voice features, including patterns such as tone, accent, and speaking habits. The extracted voice feature data is then used to generate the voice message.
[1633] Step 5:
[1634] The text message entered by the user is sent to a cloud server, which uses a generative AI model to generate a text message with natural wording. In this process, a more emotionally charged message is generated based on the user's requested phrases.
[1635] Step 6:
[1636] The cloud server converts the generated text message into a voice message using the Pyttsx3 speech synthesis engine, which generates a voice message that reflects the analyzed voice characteristics, ensuring that the voice message sounds as close as possible to the voice of the deceased person that the user is familiar with.
[1637] Step 7:
[1638] The cloud server generates a video using the generated voice message and the extracted facial features. Using tools such as OpenCV and MoviePy, the video is created so that the face appears to be moving. The video is designed to play back the message entered by the user emotionally.
[1639] Step 8:
[1640] The cloud server saves the generated video in cloud storage and issues a security token to protect the data. This token is used to verify the user's access rights when playing the video. Storing the video in cloud storage ensures secure data management.
[1641] Step 9:
[1642] The user then accesses the cloud service again and plays the generated video in streaming format. The video data is provided from the cloud storage and played in real time on the user's device. At this time, the user's access authority is confirmed using a security token.
[1643] This series of processing steps generates an emotional video message that reproduces the voices and faces of loved ones, allowing users to watch it safely.
[1644] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1645] This invention relates to a system that generates video messages based on photographic and audio data from the deceased's life, and also describes an embodiment that combines an emotion engine. This system allows users to recreate the voice and face of a loved one and receive a heartfelt message that reflects their emotions.
[1646] System configuration
[1647] This system consists of a user's device, a cloud server, cloud storage, and an emotion engine. Each component will be explained below.
[1648] User's device
[1649] The user's device has the function of uploading photo data and audio data to the cloud server, the function of playing streaming videos provided by the cloud server, and the function of collecting the user's emotional data and sending it to the cloud server.
[1650] Cloud Server
[1651] The cloud server receives and analyzes photo and audio data. It also generates text messages based on the analysis results and converts them into audio messages. It also uses an emotion engine to recognize the user's emotions, adjusts the message content and tone as necessary, and generates videos.
[1652] Cloud Storage
[1653] Cloud storage is a database that securely stores the generated videos and provides them to users in streaming format. By using security tokens, user access rights are confirmed and secure data management is achieved.
[1654] Emotion Engine
[1655] The emotion engine analyzes emotional data sent from the user's device and recognizes the user's emotional state in real time. Based on this information, it adjusts the tone and content of generated messages, audio, and video.
[1656] Program processing
[1657] The program in this system plays four main roles: user, terminal, server, and emotion engine. The processing flow and specific operations are explained below.
[1658] First, the user accesses the cloud service and logs in. Next, the user uploads photos and recorded voices of their deceased family members from their device. The device converts these data into appropriate formats (e.g., JPEG and WAV) and sends them to the cloud server.
[1659] Once the server receives the uploaded data, it first performs image analysis. The image analysis module extracts facial features and generates a high-resolution image. Then, the voice analysis module analyzes the audio data and extracts voice characteristics, including dialect and speech patterns.
[1660] Next, based on the text message entered by the user, the server uses a generative AI model to generate a naturally-phrased text message, which is then converted into a voice message using a speech synthesis engine, reflecting the analyzed voice characteristics.
[1661] Furthermore, the emotion engine receives emotion data sent from the user's device and analyzes the user's emotional state. Based on the results, the generated message content and voice tone are adjusted. Specifically, different messages and tones are adopted depending on changes in emotion.
[1662] The server then uses the generated voice message and facial features to generate a video that appears to be moving, which is then stored in cloud storage and a security token is issued to protect the data.
[1663] Finally, the user accesses the cloud service again and plays the generated video in streaming format. The server verifies the user's access privileges and then plays the video securely using a security token.
[1664] Specific examples
[1665] 1. The user logs in to the cloud service and uploads a photo of their deceased mother and her last recorded message (e.g., a birthday message) from their device to the cloud server.
[1666] 2. The server analyzes the uploaded photo, extracts the mother's facial features, and generates a high-resolution image. At the same time, it performs audio analysis to extract the mother's tone of voice, accent, and speaking habits.
[1667] 3. The user requests a video of their mother wishing them a happy birthday. Based on this request, the server uses a generative AI model to generate a text message and converts it into a voice message.
[1668] 4. At this time, the emotion engine analyzes the user's emotions and, for example, if it recognizes that the user is emotional, it adjusts the voice tone to a gentler one.
[1669] 5. The server generates a video of the mother speaking using the generated voice message and the analyzed high-resolution facial features.
[1670] 6. The server stores the generated video in cloud storage and issues a security token to protect the data.
[1671] 7. The user accesses the cloud service to play the generated video and watch the heartwarming message from their deceased mother.
[1672] The above is a specific description of an embodiment of the "My Album in Action" system that combines the emotion engine of the present invention.
[1673] The processing flow will be explained below.
[1674] Step 1: Log in and upload data
[1675] 1. The user accesses the cloud service and logs in by entering their user ID and password.
[1676] 2. The device sends the login information to the cloud server for authentication.
[1677] 3. The server verifies the authentication information and, if successful, grants access to the user.
[1678] 4. The user selects the photo and audio data of the deceased family member on the device and uploads it to the cloud server.
[1679] 5. The device converts the selected data into the appropriate format (e.g., JPEG and WAV) and sends it to the cloud server.
[1680] 6. The server stores the received photo and audio data in cloud storage.
[1681] Step 2: Data analysis
[1682] 1. The server passes the uploaded photo data to the image analysis module.
[1683] 2. The server uses an image analysis module to extract facial features (such as the position of the eyes, nose, and mouth) of family members from the photo.
[1684] 3. The server uses super-resolution technology to generate a high-resolution facial image based on the extracted feature points.
[1685] 4. The server passes the uploaded voice data to the voice analysis module.
[1686] 5. The server uses a speech analysis module to extract tone, rate, accent, and dialect and diction patterns from the speech data.
[1687] Step 3: Collect and analyze emotion data
[1688] 1. The user interacts with the cloud service interface and gives permission to collect emotion data.
[1689] 2. The device collects emotional data such as the user's facial expressions and tone of voice and sends it to a cloud server.
[1690] 3. The server passes the received emotion data to the emotion engine and analyzes the user's emotional state.
[1691] Step 4: Message Generation
[1692] 1. The user enters the content of the message they want to generate (e.g., "Happy Birthday") through the cloud service interface.
[1693] 2. The server passes the input text message to a generative AI model (e.g., GPT-4) to generate a naturally phrased text message.
[1694] 3. The server uses a speech synthesis engine to generate a voice message that sounds like the family member's voice based on the generated text message and the extracted voice characteristics.
[1695] 4. The server adjusts the tone of the voice and the content of the message based on the user's emotional data provided by the emotion engine.
[1696] Step 5: Video Generation
[1697] 1. The server uses a high-resolution facial image and the generated voice message to generate a video that matches the facial movements and voice using ravioli generation technology (e.g., Facial Animation Technology).
[1698] 2. The server stores the generated video in cloud storage and applies appropriate encryption processes.
[1699] Step 6: Streaming
[1700] 1. The user clicks the play button to request playback of the video generated from the cloud service.
[1701] 2. The server verifies the user's access rights, issues a security token, and sends it to the user's device.
[1702] 3. The device receives the security token and sends a request to the server again to start streaming the video.
[1703] 4. The server verifies the security token and sends the video data to the device in streaming format.
[1704] 5. The device plays the received video data, allowing the user to watch the video.
[1705] The above is a description of the specific operations divided into processing steps of the "Moving My Album" system.
[1706] Example 2
[1707] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1708] In recent years, there has been a demand for preserving memories of deceased loved ones and receiving them as emotionally warm messages. However, conventional technologies have struggled to reproduce the voice and appearance of the deceased in a natural way. It has also been difficult to generate messages and videos that correspond to the user's emotional state. Furthermore, ensuring the security of the generated data and managing access rights are also important issues.
[1709] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1710] In this invention, the server includes means for acquiring image data of the deceased person, means for acquiring audio data of the deceased person, means for analyzing the acquired image data and audio data and extracting facial and vocal characteristics of the person, means for generating a natural language message based on the extracted facial and vocal characteristics and converting the message into synthetic speech, means for analyzing the user's emotional data and reflecting the emotional state in the generation process, means for generating a video using the generated synthetic speech and the analyzed facial characteristics, and means for saving the generated video in a storage device and making it playable in streaming format. This makes it possible to reproduce the voice and appearance of the deceased person with high accuracy, and to easily generate and safely manage warm video messages that correspond to the user's emotional state.
[1711] "Image data from before death" refers to photographs and image files taken of a person before their death.
[1712] "Pre-death acoustic data" refers to voices and audio files recorded during a person's lifetime.
[1713] "Means of acquisition" refers to the functions and interfaces that allow users to select and upload data.
[1714] "Means for analysis" refers to the algorithms and software used to analyze the acquired data and extract the necessary information.
[1715] "Facial features" refer to identifiable parts or attributes of a person's face that are extracted from an image.
[1716] "Voice features" refer to identifiable parts or attributes such as tone, accent, and catchphrases extracted from voice data.
[1717] A "natural language message" refers to a text message that is written in a way that is natural for humans to understand.
[1718] "Synthetic voice" refers to data generated as voice based on a text message.
[1719] "Emotional Data" refers to information collected to describe a user's emotional state.
[1720] The "generation process" refers to a series of steps to generate messages and videos based on analyzed data.
[1721] "Storage device" refers to a storage system for safely storing the generated video.
[1722] "Streaming format" refers to a format that allows users to watch videos in real time over the Internet.
[1723] This invention relates to a system that generates video messages based on image and audio data from a person's life, and by combining it with an emotion engine, it allows users to receive the messages as warm messages. This system consists of a user's device, a cloud server, cloud storage, and an emotion engine.
[1724] System configuration
[1725] User's device
[1726] The user's device has the function of uploading image data and audio data to the cloud server, the function of playing streaming videos provided by the cloud server, and the function of collecting the user's emotional data and sending it to the cloud server.
[1727] Cloud Server
[1728] The cloud server is equipped with software and hardware for receiving and analyzing image and acoustic data. The analysis uses OpenCV, an image processing software, and NVIDIA NeMo, a voice analysis software. Based on the text message entered by the user, a generative AI model (e.g., a generative AI model) is used to generate a text message with natural wording.
[1729] The server converts the generated text message into a voice message using a speech synthesis engine, incorporating the analyzed voice characteristics. This process uses the speech synthesis technology of Azure Cognitive Services.
[1730] Furthermore, the Emotion Analysis Engine analyzes the user's emotional data and adjusts the generated message content and voice tone based on the user's emotional state. For example, if the user is emotional, the voice tone will be made gentler.
[1731] The server uses the generated voice message and high-resolution facial features to generate a video that appears to show a moving face using DeepFaceLab. The video is encoded using FFmpeg. The video is then stored in cloud storage, where data is protected by a security token.
[1732] Cloud Storage
[1733] Cloud storage is a database that securely stores the generated videos and provides them to users in streaming format. By using security tokens, user access rights are confirmed and secure data management is achieved.
[1734] Emotion Engine
[1735] The emotion engine analyzes emotional data sent from the user's device and recognizes the user's emotional state in real time. Based on this information, it adjusts the tone and content of generated messages, audio, and video.
[1736] Specific examples
[1737] A user logs in to the cloud service and uploads a photo (JPEG format) of a deceased family member and audio data (WAV format) from their device to the cloud server. The server analyzes the uploaded data using OpenCV and NVIDIA NeMo to extract facial and vocal characteristics. Next, the generative AI model converts the user's requested message (e.g., "Happy Birthday") into natural-sounding text, which is then converted into a voice message using Azure Cognitive Services' speech synthesis engine. The emotion engine also analyzes the user's emotional data and adjusts the message content and voice tone. Finally, the server generates a video using DeepFaceLab and saves it in cloud storage. The user can then access the cloud service again to play the generated video in streaming format.
[1738] Example prompt: "Generate a sweet message from my late mother wishing me a happy birthday."
[1739] The above is a specific description of an embodiment of the "My Album in Action" system that combines the emotion engine of the present invention.
[1740] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1741] Step 1:
[1742] A user logs in
[1743] A user accesses a cloud service using a browser or application and enters a user ID and password, which authenticates the user and allows them to log in to the cloud service.
[1744] Input: User ID, Password
[1745] Output: Authentication token, indicating successful login
[1746] Step 2:
[1747] The user uploads photos and audio data.
[1748] The user selects a photo (JPEG format) and audio data (WAV format) of the deceased family member from the device and clicks the upload button. The device converts the data into the appropriate format (JPEG and WAV) and sends it to the cloud server.
[1749] Input: Photo data (JPEG), audio data (WAV)
[1750] Output: Photo and audio data uploaded to the cloud server
[1751] Step 3:
[1752] The server receives the photo and audio data.
[1753] The server receives the photo and audio data sent from the device and temporarily stores it in the specified storage. At this stage, the server checks the integrity of the data to ensure there is no unauthorized data.
[1754] Input: Photo data and audio data uploaded from the device
[1755] Output: Temporarily saved photo and audio data
[1756] Step 4:
[1757] The server performs image analysis
[1758] The server analyzes the received photo data using OpenCV, specifically applying a facial recognition algorithm to extract facial features.
[1759] Input: Temporarily saved photo data
[1760] Output: Facial features (position, shape, landmark points, etc.)
[1761] Step 5:
[1762] The server performs voice analysis
[1763] The server uses NVIDIA NeMo to analyze the received voice data, extracting voice tone, accent, and speaking habits.
[1764] Input: Temporarily saved audio data
[1765] Output: Voice characteristics (tone, accent, speech patterns)
[1766] Step 6:
[1767] The server generates a text message
[1768] Based on a prompt phrase entered by the user (e.g., "Say happy birthday"), the server uses a generative AI model (e.g., GPT-3) to generate a natural-sounding text message.
[1769] Input: The prompt text entered by the user
[1770] Output: The generated text message
[1771] Step 7:
[1772] The server converts it into a voice message
[1773] Based on the generated text message, the server uses Azure Cognitive Services' speech synthesis engine to convert it into a voice message, which reflects the analyzed voice characteristics.
[1774] Input: Generated text message, voice features
[1775] Output: Synthesized voice message
[1776] Step 8:
[1777] The emotion engine analyzes the emotion data
[1778] The emotion engine analyzes emotion data sent from the user's device and recognizes the user's emotional state in real time.
[1779] Input: User emotion data (facial expressions, tone of voice, etc.)
[1780] Output: Emotional state analysis result
[1781] Step 9:
[1782] The server generates a video based on the generated data.
[1783] The server uses DeepFaceLab to generate a video that appears to show a moving face based on the generated voice message and high-resolution facial features, and encodes the video using FFmpeg.
[1784] Input: Synthesized voice message, facial features
[1785] Output: Generated video file
[1786] Step 10:
[1787] The server saves the video to cloud storage.
[1788] The generated video is stored in cloud storage. A security token is issued to verify the user's access rights and protect the data.
[1789] Input: Generated video file
[1790] Output: Video stored in cloud storage, security token
[1791] Step 11:
[1792] Playing user-generated videos
[1793] The user then accesses the cloud service again and plays the video. The server then verifies the user's access privileges and uses the security token to play the video securely in streaming format.
[1794] Input: Security token, saved video file
[1795] Output: Streamed video
[1796] The above are the detailed processing steps of this system.
[1797] (Application example 2)
[1798] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1799] In order to recreate the memory of the deceased and move the bereaved family, it is desirable to provide more realistic and emotionally relevant video messages. However, current technology has problems in that it is difficult to adjust the tone and content of the message according to the user's emotional state, and the generated videos lack a sense of realism. In addition, there is a lack of appropriate display methods for viewing moving videos in real time. There is a need to solve these problems.
[1800] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1801] In this invention, the server includes means for uploading photo data of the person before death, means for uploading audio data of the person before death, means for analyzing the uploaded photo data and audio data and extracting facial and vocal characteristics of the person, means for generating a text message based on the extracted facial and vocal characteristics and converting the text message into an audio message, means for generating a video using the generated audio message and the analyzed facial characteristics, means for storing the generated video in cloud storage and making it playable in streaming format, means for collecting emotional data and analyzing the data to adjust the tone of the text message or audio message, and means for displaying the generated video using a head-mounted display, thereby making it possible to provide a realistic video message that matches the emotional state of the user.
[1802] "Photo data from before death" refers to image files of a person taken before death.
[1803] "Voice data from before death" refers to an audio file of a person's voice recorded before death.
[1804] The "means for uploading" refers to a method and apparatus for transmitting data from a terminal to a cloud server.
[1805] "Means for analyzing" refers to methods and devices that process uploaded data to extract information.
[1806] "Facial features" are characteristic parts of a person's face that are extracted from the analyzed image data.
[1807] "Voice features" are characteristics of a person's voice extracted from analyzed voice data.
[1808] A "means for generating a text message" is a method and device that creates a sentence based on input information.
[1809] A "means for converting to voice message" is a method and apparatus for converting a text message to voice data.
[1810] The "means for generating video" refers to a method and device for creating a video file based on image data and audio data.
[1811] "Cloud storage means" refers to a method and device for storing data on a remote server.
[1812] "Means for enabling playback in streaming format" refers to a method and apparatus for playing back stored video data in real time via the Internet.
[1813] A "means for collecting emotional data" is a method and apparatus for collecting a user's emotional state.
[1814] A "means for analyzing emotion data" is a method and apparatus for processing collected emotion data to make sense of the information.
[1815] "Means for adjusting the tone of text and voice messages" refers to a method and apparatus for changing the tone of message content based on analyzed emotional data.
[1816] A "head-mounted display" is a display device worn on the head, and is used to display videos and images.
[1817] This invention relates to a system for generating moving video messages using photographic and audio data from a person's lifetime. Furthermore, by collecting and analyzing the user's emotional data, the system adjusts the message content and tone to generate more personalized and emotional videos. A specific embodiment of this system is described below.
[1818] System configuration
[1819] This system consists of a user's device, a cloud server, cloud storage, an emotion engine, and a head-mounted display.
[1820] User's device
[1821] The user's device has the function of uploading photo data and audio data to the cloud server, the function of playing streaming videos provided by the cloud server, and the function of collecting the user's emotional data and sending it to the cloud server.
[1822] Cloud Server
[1823] The cloud server receives and analyzes photo and audio data. It also generates text messages based on the analysis results and converts them into audio messages. It also uses an emotion engine to recognize the user's emotions, adjusts the message content and tone as necessary, and generates videos.
[1824] Cloud Storage
[1825] Cloud storage is a database that securely stores the generated videos and provides them to users in streaming format. By using security tokens, user access rights are confirmed and secure data management is achieved.
[1826] Emotion Engine
[1827] The emotion engine analyzes emotional data sent from the user's device and recognizes the user's emotional state in real time. Based on this information, it adjusts the tone and content of generated messages, audio, and video.
[1828] head-mounted display
[1829] A head-mounted display is a device worn on the head that displays images to the user. It displays video received from a cloud server in real time, providing the user with a realistic viewing experience.
[1830] Program processing
[1831] The program in this system mainly plays four main roles: user, terminal, server, and emotion engine.
[1832] First, the user accesses the cloud service and logs in. The user uploads photos and audio data from their device to the cloud server. The device converts the data into an appropriate format (e.g., JPEG and WAV) and sends it to the cloud server.
[1833] Once the cloud server receives the uploaded data, it first performs image analysis. The image analysis module extracts facial features and generates a high-resolution image. The voice analysis module then analyzes the audio data and extracts voice characteristics, including dialect and speech patterns.
[1834] Next, based on the text message entered by the user, the cloud server uses a generative AI model to generate a naturally-phrased text message, which is then converted into a voice message using a speech synthesis engine, reflecting the analyzed voice characteristics.
[1835] Furthermore, the emotion engine receives emotion data sent from the user's device and analyzes the user's emotional state. Based on the results, it adjusts the generated message content and voice tone. Specifically, it adopts different messages and tones depending on changes in the user's emotions.
[1836] The cloud server then uses the generated voice message and facial features to generate a video that appears to be moving, which is then stored in cloud storage and a security token is issued to protect the data.
[1837] Finally, the user wears a head-mounted display and watches the generated video. The cloud server verifies the user's access rights and then plays the video securely using a security token.
[1838] Specific examples
[1839] A user visits a funeral home and uploads a photo of their deceased father and a recorded voice message (e.g., a birthday message). The cloud server analyzes the uploaded photo, extracts the father's facial features, and generates a high-resolution image. At the same time, it performs voice analysis to extract the father's tone of voice, accent, and catchphrases.
[1840] A user can request a generative AI model to generate a video of their father saying, "Take care." The emotion engine then analyzes the user's emotions and adjusts the tone of the voice to be gentler if it detects that the user is emotional, for example.
[1841] The cloud server then uses the generated voice message and analyzed high-resolution facial features to generate a video of the father speaking, which is then stored in cloud storage and a security token is issued to protect the data.
[1842] Finally, the user can wear a head-mounted display to watch the generated video and receive a heartwarming message from their deceased father.
[1843] Example prompts for generative AI models
[1844] Below is an example of a prompt sentence for the generative AI model (generative AI model).
[1845] "Dear visitors, wipe away your tears and upload your favorite photos and audio recordings here. We will recreate the voice and face of your loved one and deliver a heartwarming video message. Let's share this special moment together. We will recreate phrases such as 'Thank you' and 'Take care' in the video."
[1846] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1847] Step 1:
[1848] The user accesses the cloud service and logs in.
[1849] Input: User authentication information (user ID, password)
[1850] Output: Authentication result (success / failure)
[1851] Specific behavior:
[1852] The user's device sends authentication information to the cloud server, which checks it against the database and, if authentication is successful, proceeds to the next step.
[1853] Step 2:
[1854] Users upload photo data and audio data from their pre-death lives from their devices to a cloud server.
[1855] Input: Photo data from before death (JPEG), audio data from before death (WAV)
[1856] Output: Data stored on the cloud server
[1857] Specific behavior:
[1858] Users simply select photos and audio data from their device and click the upload button. The device then sends the data to the cloud server, which stores it securely.
[1859] Step 3:
[1860] The cloud server analyzes the uploaded photo data.
[1861] Input: Photo data from before death (JPEG)
[1862] Output: Facial feature data (facial feature points, facial expressions)
[1863] Specific behavior:
[1864] The server uses an image analysis module (e.g., OpenCV) to extract facial features from the photo data and generate high-resolution facial images.
[1865] Step 4:
[1866] The cloud server analyzes the voice data from the person's life.
[1867] Input: Audio data from before death (WAV)
[1868] Output: Voice characteristics data (tone of voice, accent, catchphrases)
[1869] Specific behavior:
[1870] The server uses a voice analysis module (e.g., librosa) to extract voice features from the audio data.
[1871] Step 5:
[1872] The cloud server generates a message using a generative AI model based on the text message entered by the user.
[1873] Input: User-entered text messages, facial feature data, and voice feature data
[1874] Output: Naturally-phrased text message
[1875] Specific behavior:
[1876] The server sends a prompt to a generative AI model (e.g., GPT-3) to generate a naturally phrased text message.
[1877] Step 6:
[1878] The cloud server converts the generated text message into a voice message.
[1879] Input: Naturally-phrased text messages, voice feature data
[1880] Output: Voice message (audio file)
[1881] Specific behavior:
[1882] The server uses a speech synthesis engine (e.g., Watson Text to Speech) to convert the text message into speech, and the generated voice message reflects the voice characteristics.
[1883] Step 7:
[1884] The emotion engine analyzes the user's emotion data and transmits the emotional state to the cloud server.
[1885] Input: User emotion data (real-time emotional state)
[1886] Output: Emotion analysis result (emotional state)
[1887] Specific behavior:
[1888] The emotion engine analyzes the emotion data collected from the user's device and sends the results to a cloud server.
[1889] Step 8:
[1890] The cloud server adjusts the message content and voice tone based on the results of emotion analysis.
[1891] Input: Emotion analysis results, voice message, facial feature data
[1892] Output: Adjusted message content, audio tone
[1893] Specific behavior:
[1894] The results of the emotion engine are reflected and the tone of the generated text and voice messages is adjusted to match the user's emotions.
[1895] Step 9:
[1896] The cloud server generates a video using the generated voice message and facial features.
[1897] Input: Facial feature data, voice message
[1898] Output: Video file
[1899] Specific behavior:
[1900] The server uses a video generation module (for example, Adobe After Effects API) to generate a video in which a person's face moves based on the voice message and facial feature data.
[1901] Step 10:
[1902] The cloud server stores the generated video in cloud storage and issues a security token to protect the data.
[1903] Input: Video file
[1904] Output: Saved video file, security token
[1905] Specific behavior:
[1906] The server stores the video file in cloud storage (e.g., AWS S3) and issues a security token for access control.
[1907] Step 11:
[1908] The user wears a head-mounted display and watches videos generated from a cloud server.
[1909] Input: Security Token
[1910] Output: Video Streaming
[1911] Specific behavior:
[1912] The user accesses the cloud service again and watches the video in real time using a head-mounted display (e.g., HoloLens, Oculus Rift). The cloud server uses a security token to verify the user's access privileges and plays the video securely.
[1913] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1914] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1915] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1916] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1917] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1918] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1919] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1920] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1921] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1922] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1923] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1924] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1925] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1926] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1927] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1928] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1929] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1930] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1931] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1932] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1933] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1934] The following is further disclosed regarding the above embodiment.
[1935] (Claim 1)
[1936] A way to upload photo data from before death,
[1937] A means to upload voice data from before death,
[1938] A means for analyzing the uploaded photo data and audio data and extracting facial features and voice features of a person;
[1939] means for generating a text message based on the extracted facial features and voice features and converting the text message into a voice message;
[1940] means for generating a video using the generated voice message and the analyzed facial features;
[1941] A means to store the generated video in cloud storage and make it playable in streaming format;
[1942] A system including:
[1943] (Claim 2)
[1944] 10. The system according to claim 1, further comprising means for issuing a security token and verifying a user's access rights during streaming playback.
[1945] (Claim 3)
[1946] 10. The system of claim 1, including means for using a database of dialects and dictionaries to reflect particular patterns in the analysis and generation of speech data.
[1947] "Example 1"
[1948] (Claim 1)
[1949] means for a user to access digital data storage and upload photographic and audio data;
[1950] A means for converting the uploaded photo data and audio data into an appropriate format;
[1951] means for transmitting the converted photo data and audio data to a cloud server;
[1952] A means for analyzing the photo data and audio data received by the cloud server and extracting facial features and voice features of a person;
[1953] a means for generating a text message using a generative AI model based on the extracted facial features and voice features, and converting the text message into a voice message;
[1954] means for generating a video using the generated voice message and the analyzed facial features;
[1955] A means to store the generated video in cloud storage and make it playable in streaming format;
[1956] A system including:
[1957] (Claim 2)
[1958] 10. The system according to claim 1, further comprising means for issuing a security token and verifying a user's access rights during streaming playback.
[1959] (Claim 3)
[1960] 10. The system of claim 1, including means for using a database of dialects and dictionaries to reflect particular patterns in the analysis and generation of speech data.
[1961] "Application Example 1"
[1962] (Claim 1)
[1963] A way to upload photo data from before death,
[1964] A means to upload voice data from before death,
[1965] A means for analyzing the uploaded photo data and audio data and extracting facial features and voice features of a person;
[1966] means for generating a text message based on the extracted facial features and voice features and converting the text message into a voice message;
[1967] means for generating a video using the generated voice message and the analyzed facial features;
[1968] A means to store the generated video in cloud storage and make it playable in streaming format;
[1969] The system is installed on a smartphone or tablet and includes means for generating, storing, and playing back video messages of memories based on the voices and photos of loved ones.
[1970] (Claim 2)
[1971] 10. The system according to claim 1, further comprising means for issuing a security token and verifying a user's access rights during streaming pla...
Claims
1. A way to upload photo data from before death, A means to upload voice data from before death, A means for analyzing the uploaded photo data and audio data and extracting facial features and voice features of a person; means for generating a text message based on the extracted facial features and voice features and converting the text message into a voice message; means for generating a video using the generated voice message and the analyzed facial features; A means to store the generated video in cloud storage and make it playable in streaming format; A system including:
2. 2. The system according to claim 1, further comprising means for issuing a security token and verifying a user's access authority during streaming playback.
3. 10. The system of claim 1, including means for using a database of dialects and dictionaries to reflect particular patterns in the analysis and generation of speech data.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A