System
Patent Information
- Application Number
- JP2024119155
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-24
- Publication Date
- 2026-02-05
Smart Images

Figure 2026018094000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] In recent years, with the advancement of technology, it has become common for many people to store digital data. However, there are limited appropriate means to re-experience past memories or specific historical events vividly and in detail. Furthermore, many existing systems lack data integration and consistency, making it difficult to provide personalized and realistic experiences. In this situation, there is a demand for re-experiences of past memories and historical events in detail and to provide experiences that are meaningful to individual users. [Means for solving the problem]
[0005] The present invention provides a system that receives past photos, documents, and audio recordings from a user and analyzes the data to extract metadata. It also includes a system that acquires related supplemental information from an external database and generates a scenario for a virtual time travel experience based on the analysis and supplemental information. This system aims to solve this problem by providing a system that generates visual and audio content based on the generated scenario and provides it to the user. By collecting and analyzing user feedback and incorporating it into the next experience generation, it is possible to provide a more accurate and personalized experience. Specifically, the system includes a combination of methods, such as recreating environmental sounds and conversations based on visual images and audio information to recreate specific dates and locations in the past based on the data received from the user, and adjusting weather data and newspaper articles obtained from an external database to ensure consistency within the scenario. In this way, the user can experience past events realistically and in detail.
[0006] A "user" is an individual or entity that accesses the system and provides historical media data.
[0007] "Terminal" means a device through which a user accesses the system, uploads data, and views experiences, and specifically includes smartphones, tablets, computers, VR goggles, etc.
[0008] A "server" is a computer system that performs the main calculation processing of the system, and is a device that receives data, analyzes data, generates scenarios, provides content, and so on.
[0009] "Data analysis" is the process of extracting metadata from user-provided media data (photos, documents, audio recordings, etc.) and converting it into meaningful information.
[0010] "Metadata" is additional information that accompanies media data such as photographs and audio recordings, and includes the date, time, location, speaker, content, and so on.
[0011] An "external database" is an information source that exists outside the system and is accessible, and includes weather data, newspaper articles, tourist guide information, and the like.
[0012] "Complementary information" is information added to metadata, and is data that makes the experience scenario consistent based on data obtained from external databases and analysis results.
[0013] "Scenario generation" is the process of designing the flow and content of a virtual time travel experience based on the analyzed metadata and complementary information.
[0014] "Visual content" is video data created based on the generated scenario, and is images and videos that allow users to visually re-experience past events and places.
[0015] "Audio content" refers to audio data created based on a generated scenario, and is audio and environmental sounds that allow the user to auditorily re-experience past events and environmental sounds.
[0016] "Feedback" refers to reaction data such as impressions, requests for improvement, and evaluations provided by users after experiencing the product. [Brief explanation of the drawings]
[0017] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0018] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0019] First, the terms used in the following description will be explained.
[0020] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0021] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0022] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0023] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0024] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0025] [First embodiment]
[0026] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0027] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0028] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0029] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0030] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0031] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0032] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0033] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0034] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0035] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0036] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0037] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0038] To implement the present invention, a user must provide historical photographs, documents, and audio recordings to the system, which then analyzes and integrates this data to generate a virtual time travel experience. The following describes in detail the mode for implementing the invention.
[0039] 1. System Overview
[0040] This system is based on users, terminals, and servers, and is composed of the following major functional modules:
[0041] 1. Data Collection and Analysis Module
[0042] 2. Experience Generation Module
[0043] 3. Interaction Module
[0044] 2. Program Processing
[0045] Data Acquisition and Analysis Module
[0046] First, users access the system using a terminal and create an account. After logging in, they upload photos, documents, and audio recordings (e.g., photos from a family trip or audio recordings of conversations) related to past memories.
[0047] The terminal transmits the uploaded data to the server.
[0048] The server securely stores the received data and analyzes it using AI algorithms. For example, it can recognize landscapes and people from photo data and extract the location and time of the photo. It also uses voice recognition technology to convert audio data into text.
[0049] External Data Collection
[0050] The server also retrieves relevant complementary information from external databases (e.g., weather databases, news archives), collecting data about the day's weather and important events.
[0051] This allows user-provided data and external information to be integrated consistently, laying the foundation for more realistic reconstructions of past events.
[0052] Experience Generation Module
[0053] The server combines the collected metadata with external information and uses AI models to generate virtual time-travel scenarios, such as a day in the life of a family vacation or a detailed description of specific events.
[0054] Based on the generated scenario, visual content (e.g., 3D models of old landscapes and people) and audio content (e.g., conversations and environmental sounds) are generated.
[0055] Interaction Module
[0056] The server transmits the completed visual and audio content to the terminal.
[0057] The device provides users with an experience through VR goggles and speakers, allowing them to virtually relive scenes from a past family trip, for example, and hear the conversations and environmental sounds from that time.
[0058] After the experience, users provide feedback via the device, including their impressions of the experience and requests for improvement.
[0059] The server analyzes the collected feedback and uses it to generate the next experience, improving the accuracy of the system and user satisfaction.
[0060] Specific examples
[0061] Example: Recreating family trip memories
[0062] 1. The user uploads photos, voice messages, and social media posts from tourist spots they have visited with their family to the system.
[0063] Examples: A photo taken at the beach on a summer day, or an audio recording of a family barbecue.
[0064] 2. The server analyzes this data and retrieves the day's weather (e.g., sunny with occasional clouds, temperature 30°C) and local newspaper articles (e.g., that a local festival was held) from an external database.
[0065] 3. The server creates a scenario of a day at the beach based on the retrieved metadata and complementary information. For example, spending the morning at the beach, having a family barbecue in the afternoon, and then going to a local festival.
[0066] 4. Based on the generated scenario, the server generates visual and audio content that realistically recreates the beach scenery and festival scenes.
[0067] 5. The device provides users with an experience through VR goggles, allowing them to experience the beach scenery right before their eyes and hear the conversations from past barbecues and the sounds of festivals in a realistic way.
[0068] 6. Users provide feedback after their experience, and the system uses that feedback to further improve the experience next time.
[0069] The present invention allows users to re-experience past memories in vivid detail, providing a meaningful virtual time travel experience for each individual user.
[0070] The processing flow will be explained below.
[0071] Step 1:
[0072] Users access the system using a terminal and log in. They select and upload past photos, documents, and audio recordings.
[0073] Step 2:
[0074] The device sends the data uploaded by the user to the server, including photo files, audio files, text files, etc.
[0075] Step 3:
[0076] The server stores the received data in a secure database and begins analyzing each piece of data, using AI algorithms to extract landscape and person metadata (e.g., date and time of photo capture, location), and convert the audio data into text.
[0077] Step 4:
[0078] The server makes requests to external APIs to retrieve relevant information from public databases, for example, weather data for a particular date and location, or newspaper articles for that day.
[0079] Step 5:
[0080] The server synthesizes the collected and analyzed data and generates scenarios for virtual time travel experiences, using AI models to create a coherent timeline based on the user's experiences. For example, it generates a scenario where the user arrives at the beach at 10:00 AM and has lunch at 1:00 PM.
[0081] Step 6:
[0082] The server generates visual content based on the scenario. It uses image generation models to create 3D models of past landscapes and people. Specifically, it realistically recreates beach sand, ocean views, and the user's family.
[0083] Step 7:
[0084] The server also generates audio content based on the scenario, using a speech generation model to realistically recreate past conversations and environmental sounds, providing an audio experience that includes, for example, the sounds of waves, wind, and conversations with family members.
[0085] Step 8:
[0086] The server then sends the finished visual and audio content to the device, which includes 4K resolution video and high-quality audio files to the user.
[0087] Step 9:
[0088] The device provides users with a virtual experience using VR goggles and high-quality speakers. Users can experience visual content through the VR goggles they wear and listen to audio content through the speakers.
[0089] Step 10:
[0090] After the experience, users can enter their feedback, such as their impressions of the experience or requests for improvements, into a dedicated form.
[0091] Step 11:
[0092] The terminal transmits the user's feedback to the server.
[0093] Step 12:
[0094] The server analyzes the received feedback and uses it as data to improve the system, improving the accuracy of the AI model and reflecting it in the next experience generation.
[0095] The above are the specific processing steps for allowing the user to vividly re-experience past memories and events.
[0096] Example 1
[0097] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0098] Previous technologies for recreating past memories lacked the accuracy and consistency of information analysis, making it difficult for users to vividly relive events that they actually experienced. Furthermore, integration with external information was insufficient, resulting in an inconsistent reproduction of past events. Furthermore, there was a lack of a way to effectively utilize user feedback to improve the experience.
[0099] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0100] In this invention, the server includes means for receiving past multimedia data from a user, means for analyzing the data and extracting metadata, means for obtaining related supplemental information from an external database, means for generating a scenario for a virtual time travel experience based on the analysis and supplemental information, means for providing prompts to a generative AI model to generate details of the scenario, means for generating visual and audio content based on the generated scenario, means for providing the generated content to the user, and means for collecting, analyzing, and reflecting user feedback in the generation of the next experience. This allows the user to vividly relive past memories and enjoy a realistic virtual time travel experience while maintaining consistency with external information. Furthermore, by utilizing feedback, the experience can be continuously improved.
[0101] A "user" is an individual or group who uses the system to relive past memories.
[0102] "Multimedia data" refers to digital data in various forms, such as photographs, documents, and audio recordings.
[0103] The "server" is a computer system that performs the core processing of the system, such as analyzing received data, obtaining supplementary information, generating scenarios, and creating visual and audio content.
[0104] "Metadata" refers to information that accompanies the main data, such as a photograph or audio data (e.g., the location where the photograph was taken, the time, recognized people or scenery, etc.).
[0105] An "external database" is a database located outside the system that is a source of information that provides supplementary information such as weather data and news data.
[0106] "Virtual time travel experience" refers to a virtual reality experience provided by a system that allows users to re-experience past events and environments through visual and audio experiences.
[0107] A "scenario" refers to a description or story that includes the specific sequence and details of past events that a user re-experiences.
[0108] "Generative AI model" refers to artificial intelligence technology for generating scenario details and descriptions based on prompt text.
[0109] A "prompt" is an input document that provides instructions to a generative AI model and provides a starting point for generating scenarios and explanatory text.
[0110] "Visual Content" refers to the visual elements, such as footage, images, and 3D models, provided in the virtual time travel experience.
[0111] "Audio Content" refers to the audio elements, such as music, sound effects, and dialogue, that are provided in a virtual time travel experience.
[0112] "Feedback" refers to information such as evaluations, impressions, and requests for improvement provided by users after experiencing a product.
[0113] The system of the present invention aims to provide a virtual time travel experience that allows users to vividly relive past memories. This system is mainly composed of three entities: a user, a terminal, and a server. Specific embodiments for implementing this system are described in detail below.
[0114] 1. Data Collection and Analysis
[0115] Users log in to the system via a dedicated application or website. After logging in, they upload multimedia data such as photos, documents, and audio recordings related to past memories via their device. The device automatically sends the uploaded data to the server, which stores it in secure storage. The server then uses AI algorithms to analyze the photo data, identifying scenes and people, and extracting metadata such as the location and time of the photo. Similarly, audio data is converted into text using voice recognition technology, and the content is analyzed.
[0116] 2. Acquiring and integrating external data
[0117] The server queries external databases (e.g., weather databases, news archives) based on the provided metadata to obtain complementary information. This allows for the collection of complementary information that is consistent with the user's multimedia data, creating a more realistic experience.
[0118] 3. Generation of Virtual Time Travel Scenario
[0119] The server combines the acquired metadata and complementary information to generate a virtual time travel scenario. This process utilizes a generative AI model, specifically using prompts such as:
[0120] "May 15, 2022, 3pm, on a beach in Tokyo"
[0121] Based on this prompt, the AI model generates a detailed scenario, such as "It's 3 p.m. and the beach is crowded with tourists, the temperature is 25 degrees, you can feel the sea breeze, and you can hear the laughter of family members all around."
[0122] 4. Visual and audio content generation
[0123] The server creates visual and audio content based on the generated scenario. The visual content is generated using a 3D graphics engine, and the audio content is reproduced realistically using voice synthesis technology.
[0124] 5. Providing a user experience
[0125] The device then provides the generated visual and audio content to the user through VR goggles and speakers, allowing the user to vividly and realistically re-experience past events, such as the beach scenery, conversations, and environmental sounds of a past family trip.
[0126] 6. Feedback Collection and Analysis
[0127] After the experience, the user provides feedback via their device. This feedback includes impressions of the experience and requests for improvements. The server collects and analyzes this feedback to help improve the system. By reflecting this feedback in the generation of the next experience, it becomes possible to increase user satisfaction.
[0128] In this way, the system of the present invention provides a virtual time travel experience that allows users to vividly relive past memories.
[0129] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0130] Step 1:
[0131] To log in to the system, users access a dedicated application or website. They enter their login information and complete authentication. Then, they upload multimedia data such as photos, documents, and audio recordings related to past memories. The input is various digital files selected by the user, and the output is data packets sent to the terminal.
[0132] Step 2:
[0133] The terminal sends the received multimedia data to the server. The terminal uses a data transfer protocol to send the files uploaded by the user to the server and add them to a queue for processing. The input is the file uploaded to the terminal, and the output is the data sent to the server.
[0134] Step 3:
[0135] The server stores the data in secure storage as soon as it receives it. Next, it uses AI algorithms to identify landscapes and people in the photo data and extract metadata such as the location and time of the photo. The audio data is converted to text using speech recognition technology and the content is analyzed. The input is the data sent from the device, and the output is the analyzed metadata and the converted audio data.
[0136] Step 4:
[0137] The server queries external databases based on the extracted metadata to obtain related complementary information. Specifically, it sends requests via API to weather databases, news archives, etc. to obtain the necessary information. The input is metadata (e.g., date and time of shooting, location), and the output is complementary information obtained from the external database.
[0138] Step 5:
[0139] The server combines internal data with external supplementary information and provides prompts to the generative AI model, which then generates a virtual time travel scenario. The prompts specify specific dates, times, and locations, and the AI model generates detailed scenarios based on them. The input is the prompts combined with metadata and supplementary information, and the output is the generated scenario.
[0140] For example, enter the prompt text as "May 15, 2022, 3:00 PM, on a beach in Tokyo."
[0141] Step 6:
[0142] The server creates visual and audio content based on the generated scenario. The visual content is realistically reproduced using a 3D graphics engine, and the audio content is realistically reproduced using speech synthesis technology. The input is the generated scenario, and the output is the visual and audio content provided to the user.
[0143] Step 7:
[0144] The terminal provides the user with visual and audio content sent from the server. Specifically, it uses VR goggles and speakers to make the experience feel realistic. The input is the content sent from the server, and the output is the virtual time travel experience provided to the user.
[0145] Step 8:
[0146] After completing the virtual time travel experience, the user provides feedback. The feedback includes impressions of the experience and suggestions for improvement, and is sent to the server via the terminal. The input is the user's impressions and suggestions for improvement, and the output is the feedback data sent to the server.
[0147] Step 9:
[0148] The server analyzes the received feedback and reflects it in the next experience generation, thereby improving user satisfaction. The input is the feedback data, and the output is the analysis results that will help improve the system.
[0149] (Application example 1)
[0150] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0151] Previous time travel experience systems have had limitations in the realism and detail of the visual and audio content when recreating past memories based on user-provided data. Furthermore, there has been a lack of methods to improve the consistency and accuracy of the generated content. This has resulted in a lack of realism and detailed reproduction in the virtual time travel experience that users can feel.
[0152] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0153] In this invention, the server includes a means for generating content to reproduce past memories provided by the user in more detail and more realistically, a means for creating optimal prompt sentences to be input into a generative AI model and generating content based on the prompt sentences, and a means for generating visual and audio data to reproduce a specific date and place in the past based on data provided by the user, thereby enabling the user to virtually re-experience past memories with high accuracy and realism.
[0154] A "user" is an individual who provides data such as photographs, documents, and audio recordings to use the system to recreate past memories as a virtual time travel experience.
[0155] "Photos" are image data taken by the user in the past, and serve as materials for generating visual content in the virtual time travel experience.
[0156] A "document" is text data created by a user in the past, and is used as auxiliary information in generating a scenario for virtual time travel.
[0157] "Audio recordings" are audio data recorded by users in the past, and are used to generate audio content for virtual time travel experiences.
[0158] "Metadata" is information extracted from photographs, documents, and audio recordings, and is additional information such as date, time, place, and people that is necessary to generate a scenario for a virtual time travel experience.
[0159] An "external database" is an external data source that provides additional information to complement the user's past memories, such as weather data or news articles.
[0160] The term "virtual time travel experience" refers to a user reliving past memories using virtual reality or augmented reality technology.
[0161] An "AI model" is an algorithm that uses machine learning techniques to analyze and integrate user-provided data to generate a virtual time travel experience.
[0162] A "prompt" is a text sentence input into a generative AI model that specifically instructs the scenario and visual and audio content of the virtual time travel experience.
[0163] "Visual content" means the visual representations, such as images and 3D models, that a user can view in the virtual time travel experience.
[0164] "Audio content" refers to the acoustic representation of conversations, environmental sounds, and other sounds that a user can hear during a virtual time travel experience.
[0165] "Feedback" refers to reaction data such as impressions and requests for improvement provided by a user after completing the virtual time travel experience.
[0166] "Content generation means" means the technical means for generating visual and audio content using AI models based on data provided by the user and complementary information from external databases.
[0167] The present invention is a system that analyzes past photographs, documents, and audio recordings provided by a user to generate a virtual time travel experience. This system is mainly composed of a user, a terminal, and a server, and is realized by the following processing steps.
[0168] 1. System Overview
[0169] The system consists of a data collection and analysis module, an experience generation module, and an interaction module.
[0170] 1.1 Data Collection and Analysis Module
[0171] Users upload photos, documents, and audio recordings related to past memories to the system. The device then sends the uploaded data to the server, which then analyzes it using AI algorithms. Scenes and people are recognized from the photos, and the location and time of the photo are extracted. Audio data is also converted into text using voice recognition technology.
[0172] 1.2 External Data Collection
[0173] The server retrieves relevant complementary information from external databases (e.g., weather databases, news archives), allowing the user-provided data to be integrated with external information to more realistically recreate past events.
[0174] 1.3 Experience Generation Module
[0175] The server integrates the collected metadata with external information and uses an AI model to generate a virtual time travel scenario, including a detailed description of a day in the family trip and specific events, and generates visual content (e.g., 3D models of historical landscapes and people) and audio content (e.g., conversations and environmental sounds).
[0176] 1.4 Interaction Modules
[0177] The server then sends the generated visual and audio content to the device, providing the experience to the user through VR goggles and speakers. The user can virtually relive scenes from a past family trip and hear the conversations and surrounding environmental sounds. After the experience, the user provides feedback, which the server analyzes and reflects in the generation of the next experience.
[0178] 2. Hardware and Software Used
[0179] The system uses the following hardware and software:
[0180] Hardware: Personal computer, VR goggles, smartphone, speakers
[0181] Software: Google Cloud Vision API (image analysis), Google Cloud Speech-to-Text API (voice recognition), cloud storage service
[0182] 3. Examples of concrete examples and prompts
[0183] Examples:
[0184] Users upload photos, voice messages, and social media posts from tourist spots they've visited with their family to the system. The system analyzes this data and retrieves the day's weather and newspaper articles from an external database. Based on the retrieved metadata and supplementary information, the system creates a scenario of a day at the beach and generates visual and audio content that realistically recreates beach scenes and festival scenes.
[0185] Example prompt sentence:
[0186] "Using user-provided data, recreate a family beach trip on August 15, 2005."
[0187] This system allows users to virtually re-experience past memories with high accuracy and realism.
[0188] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0189] Step 1:
[0190] Users access the system using a terminal and upload photos, documents, and audio recordings related to past memories. These data are used as input to generate a virtual time travel experience. The input for this step is the data provided by the user (photos, documents, audio recordings), and the output is this data sent to the server.
[0191] Step 2:
[0192] The device sends the data uploaded by the user to the server, which receives and securely stores this data. The input of this step is the data sent from the device, and the output is the data stored on the server.
[0193] Step 3:
[0194] The server analyzes the stored data and extracts metadata using AI algorithms. For example, it recognizes landscapes and people in photos and extracts the location and time of the photo, and converts audio data into text using voice recognition technology. The input for this step is the data stored on the server, and the output is the extracted metadata.
[0195] Step 4:
[0196] The server retrieves relevant complementary information from external databases (e.g., weather databases, news archives), providing complementary information that is consistent with the user's data. The input of this step is the extracted metadata, and the output is the retrieved complementary information.
[0197] Step 5:
[0198] The server integrates the collected metadata with external information and uses an AI model to generate a virtual time travel scenario. This scenario includes detailed descriptions of the day and specific events. The input of this step is the integrated metadata and complementary information, and the output is the generated scenario.
[0199] Step 6:
[0200] The server generates visual and audio content based on the generated scenario. The visual content includes 3D models of historical landscapes and people, and the audio content includes conversations and environmental sounds. The input of this step is the generated scenario, and the output is the visual and audio content.
[0201] Step 7:
[0202] The server transmits the generated visual and audio content to the terminal, where the user can experience virtual time travel using VR goggles and speakers. The input of this step is the generated visual and audio content, and the output is the provision of content to the user.
[0203] Step 8:
[0204] After the user finishes the virtual time travel experience, they provide feedback about the experience. The device sends this feedback to the server, which analyzes it. The input of this step is the user's feedback, and the output is the analyzed feedback data.
[0205] Step 9:
[0206] The server then refines the system based on the analyzed feedback, which then incorporates new knowledge into the next experience generation, improving the system's accuracy and user satisfaction. The input of this step is the analyzed feedback data, and the output is an improved system.
[0207] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0208] To implement the present invention, a user must provide past photographs, documents, and audio recordings to the system, which then analyzes and integrates this data to generate a virtual time travel experience. Furthermore, the present invention incorporates an emotion engine that recognizes the user's emotions, further personalizing the user's experience. The following describes in detail the modes for implementing the invention.
[0209] 1. System Overview
[0210] This system is based on users, terminals, and servers, and is composed of the following major functional modules:
[0211] 1. Data Collection and Analysis Module
[0212] 2. Experience Generation Module
[0213] 3. Interaction Module
[0214] 4. Emotion Engine
[0215] 2. Program Processing
[0216] Data Acquisition and Analysis Module
[0217] First, a user accesses the system using a terminal and logs in. The user uploads photos, documents, and audio recordings (e.g., photos from a family trip or audio of a conversation) related to past memories.
[0218] The terminal transmits the uploaded data to the server.
[0219] The server securely stores the received data and analyzes it using AI algorithms. For example, it can recognize landscapes and people from photo data and extract the location and time of the photo. It also uses voice recognition technology to convert audio data into text.
[0220] External Data Collection
[0221] The server makes requests to external APIs to retrieve relevant information from public databases, for example, weather data for a particular date and location, or newspaper articles for that day.
[0222] This allows user-provided data and external information to be integrated consistently, laying the foundation for more realistic reconstructions of past events.
[0223] Experience Generation Module
[0224] The server combines the collected metadata with external information and uses AI models to generate virtual time-travel scenarios, such as a day in the life of a family vacation or a detailed description of specific events.
[0225] Based on the generated scenario, visual content (e.g., 3D models of old landscapes and people) and audio content (e.g., conversations and environmental sounds) are generated.
[0226] Emotion Engine
[0227] The server recognizes emotions using photos, voice recordings, and text data provided by the user. The emotion engine analyzes the user's emotional state (e.g., joy, sadness, nostalgia) extracted from the data.
[0228] This allows the scenario of the virtual time travel experience to be adjusted according to the user's emotions, providing a more personalized experience.
[0229] Interaction Module
[0230] The server transmits the completed visual and audio content to the terminal.
[0231] The device provides users with an experience through VR goggles and speakers, allowing them to virtually relive scenes from a past family trip, for example, and hear the conversations and environmental sounds from that time.
[0232] Feedback collection
[0233] After the experience, users provide feedback via the device, including their impressions of the experience and requests for improvement.
[0234] The server analyzes the collected feedback and uses it to generate the next experience, improving the accuracy of the system and user satisfaction.
[0235] Specific examples
[0236] Example: Recreating family trip memories
[0237] 1. The user uploads photos, voice messages, and social media posts from tourist spots they have visited with their family to the system.
[0238] Examples: A photo taken at the beach on a summer day, or an audio recording of a family barbecue.
[0239] 2. The server analyzes this data and retrieves the day's weather (e.g., sunny with occasional clouds, temperature 30°C) and local newspaper articles (e.g., that a local festival was held) from an external database.
[0240] 3. The server creates a scenario of a day at the beach based on the retrieved metadata and complementary information. For example, spending the morning at the beach, having a family barbecue in the afternoon, and then going to a local festival.
[0241] 4. The server uses an emotion engine to read emotions from the user's uploaded data. Example: Identifying the user's emotion of "joy" from a photo of them at the beach.
[0242] 5. The server reflects the user's emotions in the generated scenario, for example, creating a scenario that emphasizes the joy of being at the beach or creating audio content that enhances the fun of having a barbecue.
[0243] 6. The server transmits the generated visual and audio content to the terminal.
[0244] 7. The device provides users with an experience through VR goggles, allowing them to experience the beach scenery right before their eyes and hear the conversations from past barbecues and the sounds of festivals in a realistic way.
[0245] 8. Users provide feedback after their experience, and the system uses that feedback to further improve the experience next time.
[0246] The present invention allows users to re-experience past memories in vivid detail, and the emotion engine provides a more personalized virtual time travel experience.
[0247] The processing flow will be explained below.
[0248] Step 1:
[0249] Users access the system using a terminal, create an account and log in. Users select past photos, documents and audio recordings and upload them to the system.
[0250] Step 2:
[0251] The device sends the data uploaded by the user to the server, including photo files, audio files, and text files.
[0252] Step 3:
[0253] The server stores the received data in a secure database, after which it begins analyzing each piece of data using AI algorithms.
[0254] Step 4:
[0255] The server analyzes the photo data and extracts metadata such as the landscape and people. For example, it uses image recognition technology to identify the location and date of the photo.
[0256] Step 5:
[0257] The server analyzes the audio data, converts the audio content into text, and uses voice recognition technology to identify the content of the conversation and who is speaking, saving it as metadata.
[0258] Step 6:
[0259] The server analyzes the text data provided by the user and applies emotion recognition algorithms to extract the user's emotional state (e.g., joy, sadness, nostalgia).
[0260] Step 7:
[0261] The server accesses external databases to gather additional information relevant to a particular date and location, such as weather data or newspaper articles for that day.
[0262] Step 8:
[0263] The server combines the collected metadata with external information to generate a scenario for a virtual time-travel experience, using an AI model to create a timeline with specific actions, such as "playing at the beach in the morning and having a family barbecue in the afternoon."
[0264] Step 9:
[0265] The server uses an emotion engine to reflect the user's emotions in the generated scenario. For example, it generates vivid and bright visuals to emphasize scenes in which the user felt "joy."
[0266] Step 10:
[0267] The server generates visual content, using a generative image model to create 3D models of historical landscapes and people, such as beach sand, ocean views, or even the user's family.
[0268] Step 11:
[0269] The server generates audio content, using a speech generation model to realistically recreate past conversations and environmental sounds (e.g., the sound of waves or family conversations).
[0270] Step 12:
[0271] The server then transmits the generated visual and audio content to the device, including 4K resolution video files and high-quality audio files.
[0272] Step 13:
[0273] The device provides the received content to the user, who then puts on the VR goggles and begins the experience using high-quality speakers.
[0274] Step 14:
[0275] Users can relive past memories through virtual experiences, such as a beach scene unfolding before their eyes, with the sound of waves and family conversations realistically recreated.
[0276] Step 15:
[0277] After the experience, users can enter their feedback, sending their impressions of the experience and requests for improvements to the system through a dedicated feedback form.
[0278] Step 16:
[0279] The terminal transmits the user's feedback to the server.
[0280] Step 17:
[0281] The server analyzes the received feedback and uses it as data to improve the system. Based on the analysis results, the accuracy of the AI model is improved and reflected in the next experience generation.
[0282] In this way, the present invention allows users to relive their past memories vividly and in detail, providing a personalized virtual time travel experience with an emotion engine.
[0283] Example 2
[0284] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0285] Conventional virtual experience systems have difficulty integrating user-provided data with external information to generate a realistic experience. Furthermore, because personalization that takes user emotions into account is not performed, it is not possible to provide an optimal experience for each individual user. Therefore, a system is needed that analyzes the diverse data provided by users and integrates it with external information to provide an individually optimized virtual experience that reflects the user's emotions.
[0286] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a means for receiving past image data, text data, and audio data from the user, a means for analyzing this data and extracting metadata, and a means for acquiring related supplementary information from an external information source. This makes it possible to provide a realistic and personalized virtual experience while taking into account the user's emotions.
[0287] A "user" is a person who uses the virtual experience system.
[0288] "Image data" refers to digital data that contains visual information such as photographs and illustrations.
[0289] "Text data" refers to digital data that contains information in the form of sentences or characters.
[0290] "Audio data" refers to digital data that contains auditory information such as human voices and environmental sounds.
[0291] "Metadata" is data that represents information related to data, and is information that serves as the basis for analysis and integration.
[0292] "External sources" are resources that provide additional information, such as databases or APIs outside the system.
[0293] A "scenario" is the story or sequence of events that make up the overall picture of a virtual experience.
[0294] "Visual content" refers to visual information such as images and videos that are presented to a user.
[0295] "Audio content" refers to auditory information such as music and sound effects that is presented to the user.
[0296] "Feedback" refers to impressions and suggestions for improvement provided by users of the virtual experience system.
[0297] "Emotion" refers to information that represents the user's psychological state, such as joy, sadness, or nostalgia.
[0298] "Personalization" is the process of optimizing the content and presentation of an experience for each individual user.
[0299] To implement this invention, the user must provide the system with previously captured image data, text data, and audio data, which is then analyzed and integrated to generate a virtual experience. In addition, the invention incorporates an emotion engine that recognizes the user's emotions, allowing for a more personalized experience.
[0300] System configuration
[0301] This system is based on users, terminals, and servers, and is composed of the following main functional modules:
[0302] 1. Data Collection and Analysis Module
[0303] 2. Experience Generation Module
[0304] 3. Interaction Module
[0305] 4. Emotion Engine
[0306] Data Acquisition and Analysis Module
[0307] First, the user accesses the system using a terminal and logs in. The user then uploads image data, text data, and audio data related to past memories.
[0308] The terminal transmits the uploaded data to the server.
[0309] The server securely stores the received data and analyzes each piece of data using AI algorithms. For example, it recognizes landscapes and people from photo data and extracts the location and time of the photo. It also uses voice recognition technology to convert audio data into text. The specific software used is an "image recognition API" for image analysis and a "voice recognition API" for audio analysis.
[0310] External Data Collection
[0311] The server sends requests to external APIs to retrieve relevant information from public databases, such as weather data for a specific date and location, or news articles for that day. Specifically, it uses the "Weather API" for weather information and the "News API" for news articles.
[0312] This allows user-provided data and external information to be integrated consistently, laying the foundation for more realistic reconstructions of past events.
[0313] Experience Generation Module
[0314] The server combines the collected metadata with external information and generates a scenario for the virtual experience using a generative AI model. For example, it creates a scenario that includes a detailed description of a day on a family trip or a specific event. The AI model used is a "generative AI model (such as GPT-4)."
[0315] Based on the generated scenario, visual content (e.g., 3D models of old landscapes and people) and audio content (e.g., conversations and environmental sounds) are generated. Visual content is generated using "visual content generation software (e.g., 3D modeling software)," and audio content is generated using "audio editing software."
[0316] Emotion Engine
[0317] The server recognizes emotions using image data, audio data, and text data provided by the user. The emotion engine analyzes the user's emotional state (e.g., joy, sadness, nostalgia) extracted from the data. An "emotion analysis API" is used for emotion analysis.
[0318] This allows the virtual experience scenario to be adjusted according to the user's emotions, providing a more personalized experience.
[0319] Interaction Module
[0320] The server transmits the completed visual and audio content to the terminal.
[0321] The device provides users with an experience through VR goggles and speakers, allowing them to virtually relive scenes from a past family trip, for example, and hear the conversations and environmental sounds from that time.
[0322] Feedback collection
[0323] After the experience, users provide feedback via their device, including their impressions of the experience and requests for improvement.
[0324] The server analyzes the collected feedback and uses it to generate the next experience, improving the accuracy of the system and user satisfaction.
[0325] Specific examples
[0326] Example: Recreating family trip memories
[0327] 1. The user uploads image data, voice messages, and social media posts of tourist spots visited with their family to the system.
[0328] Examples: a photo taken at a place you visited on a summer day, an audio recording of you having a meal with your family.
[0329] 2. The server analyzes this data and retrieves the day's weather (e.g., sunny with occasional clouds, temperature 30°C) and local news articles (e.g., local events being held) from an external database.
[0330] 3. The server creates a scenario of a day in a specific location based on the retrieved metadata and complementary information, e.g., visiting tourist spots in the morning, having a family meal in the afternoon, and then going to a local event.
[0331] 4. The server uses an emotion engine to read emotions from the user's uploaded data. For example, it identifies the user's "joy" emotion from a photo.
[0332] 5. The server reflects the user's emotions in the generated scenario. For example, it creates a scenario that emphasizes the joy of visiting a tourist spot or audio content that enhances the enjoyment of eating.
[0333] 6. The server transmits the generated visual and audio content to the terminal.
[0334] 7. The device provides users with an experience through VR goggles, allowing them to experience the scenery of tourist spots right before their eyes and hear the sounds of past conversations and events in a realistic way.
[0335] 8. Users provide feedback after their experience, and the system uses that feedback to further improve the experience next time.
[0336] Prompt Sentence Examples
[0337] Below is an example of a prompt sentence to input to the generative AI model.
[0338] I've uploaded photos and audio recordings from past family trips. Please use these to generate a virtual experience. I'd particularly like to emphasize the emotion of joy. Include supplemental information like the weather for that day and local events.
[0339] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0340] Step 1:
[0341] The user accesses the system using a terminal and logs in by entering their ID and password on the login page.
[0342] Input: ID, password
[0343] Output: User credentials
[0344] Specific operation: The device sends the login information to the server, and the server verifies the authentication information and allows the user to log in. If authentication is successful, the user's dashboard is displayed.
[0345] Step 2:
[0346] Users select image data, text data, and audio data related to past memories from the dashboard and press the upload button to upload this data to the system.
[0347] Input: image data, text data, audio data
[0348] Output: Upload data
[0349] Specific operation: The terminal divides the selected data into packets and sends them to the server.
[0350] Step 3:
[0351] The terminal transmits the uploaded data to the server.
[0352] Input: Upload data
[0353] Output: Transmitted data
[0354] Specific operation: The terminal converts the data into the appropriate format for each type (image, audio, text) and sends it to the server.
[0355] Step 4:
[0356] The server will securely store the received data in the specified directory.
[0357] Input: Send data
[0358] Output: Saved data
[0359] Specific operation: The server changes the file name and saves it in the appropriate folder depending on the type of data.
[0360] Step 5:
[0361] The server analyzes the data using AI algorithms such as image recognition, voice recognition, and text analysis.
[0362] Input: Saved data
[0363] Output: Metadata
[0364] Specific operation: Extracts scenery and people from photo data and converts audio data into text. Specifically, it uses "image recognition API" and "voice recognition API." Analyzed metadata includes location, time, and characters.
[0365] Step 6:
[0366] The server sends external API requests to retrieve relevant information from designated public sources.
[0367] Input: Metadata
[0368] Output: Complementary information
[0369] Specific operation: Compose a query for the information you want to obtain (weather, news articles, etc.) and access external sources. For example, use the "Weather API" for weather information and the "News API" for news articles.
[0370] Step 7:
[0371] The server integrates the collected metadata with external information and generates a virtual experience scenario using a generative AI model.
[0372] Input: Metadata, complementary information
[0373] Output: Scenario data
[0374] Specific operation: Automatically generates a storyboard based on the specified date and location based on user-provided data and external information. A generative AI model (such as GPT-4) is used as the generative AI model.
[0375] Step 8:
[0376] The server creates visual and audio content based on the generated scenario.
[0377] Input: Scenario data
[0378] Output: Visual content, audio content
[0379] What it does: Visual content is generated using visual content generation software (e.g., 3D modeling software) and audio content is generated using audio editing software, which generates the detailed elements of the virtual experience.
[0380] Step 9:
[0381] The server uses an emotion engine to analyze emotions from the image data, audio data, and text data provided by the user.
[0382] Input: image data, audio data, text data
[0383] Output: Emotion data
[0384] Specific operation: Emotion analysis uses the "Emotion Analysis API" to identify the user's psychological state. For example, emotions such as "joy" and "sadness" can be extracted from a photo.
[0385] Step 10:
[0386] Based on the analysis results, the server adjusts the experience scenario to match the user's emotions.
[0387] Input: Emotion data, scenario data
[0388] Output: personalized scenario data
[0389] Specific behavior: Adding music and effects that correspond to emotions and optimizing the content of the scenario for each individual user, providing an experience that emphasizes specific emotions.
[0390] Step 11:
[0391] The server transmits the completed visual and audio content to the terminal.
[0392] Input: Visual content, audio content
[0393] Output: Send content
[0394] Specific operation: The server acts as a file server and provides streaming or download links to devices.
[0395] Step 12:
[0396] The device provides users with an experience using VR goggles and speakers.
[0397] Input: Submit content
[0398] Output: Experience data
[0399] How it works: Users can enjoy a virtual experience through a dedicated VR app, and past experiences are realistically recreated for the user.
[0400] Step 13:
[0401] After the experience, users provide feedback via the device.
[0402] Input: Experience data
[0403] Output: Feedback data
[0404] What to do: Enter your thoughts and suggestions for improvement in the feedback form and submit it. Use the text fields and rating sliders.
[0405] Step 14:
[0406] The server analyzes the collected feedback using machine learning models and reflects it in improvements to the system.
[0407] Input: Feedback data
[0408] Output: Improvement data
[0409] Specific actions: Perform natural language processing of feedback to discover new areas for improvement. Specific software used includes "natural language processing software."
[0410] (Application example 2)
[0411] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0412] In today's world, there is a growing demand for new experiences that vividly recreate past memories. However, existing technologies simply display or play back past photographs and audio recordings, making it difficult to integrate this information and provide a deeply personalized virtual experience. Furthermore, they are unable to tailor the experience to the user's emotions, creating a need for a system that allows users to relive the emotions of past memories. The present invention aims to solve these problems and provide a virtual experience system that realistically recreates past experiences while recognizing the user's emotions.
[0413] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for receiving the user's past photos, documents, and voice recordings, analyzing this data, and extracting metadata; means for acquiring related complementary information from an external database; means for generating a scenario for a virtual time travel experience based on the analysis and complementary information; means for generating visual and audio content based on the generated scenario; and means for recognizing the user's emotions and adjusting the experience scenario and content accordingly. This enables the user to recreate realistic past experiences personalized according to their emotions.
[0414] "Means for recognizing a user's emotions and adjusting the experience scenario and content accordingly" refers to technology that uses artificial intelligence to analyze a user's emotional state from photos, documents, audio recordings, etc. provided by the user, and personalizes the content of the virtual experience based on those emotions.
[0415] "Generative AI models" are artificial intelligence algorithms and frameworks that perform natural language processing and image generation, examples of which include GPT-4 and DALL-E.
[0416] "Metadata" is data that provides information about the original data, such as the date and location when a photo was taken, and the contents of audio recordings.
[0417] "External databases" refer to data resources and databases that are publicly available on the Internet, including weather information and past newspaper articles.
[0418] A "virtual time travel experience scenario" is a scenario generated by artificial intelligence based on the user's past records and related information, and refers to a script or plan of action that allows the user to recreate past experiences in virtual reality.
[0419] "Visual and audio content" refers to the visual and audio elements that provide the user with a virtual experience, such as generated 3D models, ambient sounds, and dialogue.
[0420] "Means for collecting, analyzing, and reflecting feedback in generating the next experience" refers to a system that records the impressions and opinions provided by users after an experience, analyzes them using artificial intelligence, and improves the next experience scenario.
[0421] "Complementary information obtained from external databases" is additional information collected to complement the data provided by the user, examples of which include weather information and event information.
[0422] 1. System Overview
[0423] This system is based on the user, terminal, and server, and consists of the following main functional modules that analyze and integrate the user's past photos, documents, and audio recordings to generate and provide a virtual time travel experience.
[0424] 1.1 Data Collection and Analysis Module
[0425] The server receives photos, documents, and audio recordings uploaded by users using their devices. It analyzes this data and extracts metadata about location, time, and content. It uses Python and TensorFlow to apply AI algorithms to recognize scenes and people in photos and convert audio data into text.
[0426] 1.2 External Data Collection Module
[0427] The server retrieves relevant complementary information from external databases, such as the day's weather information or newspaper articles from public databases on the Internet, using a REST API to retrieve this external data.
[0428] 1.3 Experience Generation Module
[0429] The server combines the collected metadata with external information and generates a virtual time travel scenario using a generative AI model (e.g., GPT-4, DALL-E). Based on the generated scenario, visual and audio content such as 3D models and environmental sounds are generated. This process is performed using Unity or Unreal Engine.
[0430] 1.4 Emotion Engine
[0431] The server recognizes emotions from the data uploaded by the user. The emotion engine uses a natural language processing (NLP) library to analyze the user's emotional state and incorporate it into the scenario. For example, it identifies the emotion of joy from a photo and incorporates it into the scene.
[0432] 1.5 Interaction Module
[0433] The server then sends the completed visual and audio content to the terminal, providing the user with a virtual experience, which can be enjoyed through VR goggles or smart glasses.
[0434] 1.6 Feedback Collection Module
[0435] The server collects and analyzes feedback from users, including impressions of the experience and requests for improvement. The analyzed feedback is used to generate the next experience.
[0436] 2. Specific Examples
[0437] Example: Recreating family vacation memories
[0438] 1. The user uploads photos and voice messages of tourist spots that they have visited with their family to the system.
[0439] Example: Photos of tourist spots you visited on a summer day and a recording of the conversations you had at the time.
[0440] 2. The server analyzes this data and retrieves information about the day's weather and local events from an external database.
[0441] 3. The server generates a family trip experience scenario based on the acquired metadata and supplementary information, including, for example, scenes of play at tourist spots and events of the day.
[0442] 4. The server uses an emotion engine to read emotions from the user's uploaded data. For example, it identifies the user's emotion of "joy" from photos of tourist spots.
[0443] 5. The server reflects the user's emotions in the generated scenario, for example, creating a scenario that emphasizes joy or audio content that enhances enjoyment.
[0444] 6. The server transmits the generated visual and audio content to the terminal.
[0445] 7. The device provides users with an experience through VR goggles, allowing them to experience the scenery of tourist spots right before their eyes and hear past conversations and environmental sounds in a realistic way.
[0446] 8. Users provide feedback after their experience, and the server uses that feedback to further improve the experience next time.
[0447] 3. Examples of prompts
[0448] 1. To generate a scenario that recreates family vacation memories, the following prompt sentence is input into the generative AI model:
[0449] "Generate an experience scenario of a past family trip based on photos and audio recordings of tourist spots visited by the family. The scenario should include the flow of a day enjoyed at tourist spots visited on a summer day. In particular, the photos convey the emotion of 'joy,' so please reflect this emotion in the scenario."
[0450] As described above, the present invention realizes a system that realistically reproduces a user's past experiences and provides a personalized virtual time travel experience that responds to emotions.
[0451] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0452] Step 1:
[0453] Users log in to the system using a terminal and upload photos, documents, and audio recordings related to past memories. This data is input and sent from the terminal to the server. Specifically, users open the upload screen on their device, such as a smartphone or PC, select the target data using the file selection button, and press the upload button.
[0454] Step 2:
[0455] The server receives data sent by the user and stores it securely. The received data is the input, and storing it internally on the server is the output. Specifically, the server's file storage system stores the data and adds a metadata record to the database.
[0456] Step 3:
[0457] The server analyzes the stored data, recognizes scenes and people in the photos, and converts audio data into text. This is done using Python and TensorFlow. Photo and audio data are the input, and extracted metadata is the output. Specifically, the image recognition algorithm identifies scenes and people in the photos and generates associated tags. At the same time, the speech recognition algorithm converts audio into text.
[0458] Step 4:
[0459] The server sends a request to the public database's API to obtain related information from the external database. For example, to obtain weather information or event information for a specified date and time. The input is location and date information related to the user's input data, and the output is weather information and event information obtained based on that. Specifically, the server sends an HTTP request to the external API, parses the JSON response, and extracts the required information.
[0460] Step 5:
[0461] The server uses a generative AI model (e.g., GPT-4, DALL-E) to generate a scenario for a virtual time travel experience based on the analyzed metadata and the acquired external information. The input is the metadata and external information, and the output is a virtual experience scenario. Specifically, the server sends a prompt to the AI model to generate the scenario. This prompt includes the user's data and related information.
[0462] Step 6:
[0463] The server generates visual and audio content based on the scenario. For example, it creates 3D models and environmental sounds. The input is the generated scenario, and the output is the visual and audio content. Specifically, it uses Unity or Unreal Engine to model the visual content and the audio engine to generate the environmental sounds.
[0464] Step 7:
[0465] The server reads emotions from the user's uploaded data, analyzes the emotional state using an emotion engine, and reflects it in the generated scenario. The input is the user's photo and voice data, and the output is a scenario that reflects the emotions. Specifically, it uses an NLP library to perform emotion analysis and add emotional expressions to the scenario text.
[0466] Step 8:
[0467] The server sends the completed visual and audio content to the device and provides it to the user. The input is the generated content, and the output is the user's experience. Specifically, the content is streamed to the user's device and displayed and played to the user through the device's VR goggles or smart glasses.
[0468] Step 9:
[0469] After the experience, the user provides feedback through the device. The input is their impressions of the experience and requests for improvement, and the output is feedback data. Specifically, the user enters comments in the feedback form on the device and presses the send button.
[0470] Step 10:
[0471] The server analyzes the collected feedback and reflects it in the generation of the next experience. The input is feedback data, and the output is an improved experience scenario. Specifically, the server analyzes the feedback data with an analysis tool, obtains new insights, and reflects them in the next prompt.
[0472] Through the above processing steps, the user can recreate a realistic past experience that is personalized according to their emotions.
[0473] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0474] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0475] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0476] [Second embodiment]
[0477] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0478] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0479] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0480] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0481] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0482] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0483] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0484] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0485] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0486] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0487] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0488] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0489] To implement the present invention, a user must provide historical photographs, documents, and audio recordings to the system, which then analyzes and integrates this data to generate a virtual time travel experience. The following describes in detail the mode for implementing the invention.
[0490] 1. System Overview
[0491] This system is based on users, terminals, and servers, and is composed of the following major functional modules:
[0492] 1. Data Collection and Analysis Module
[0493] 2. Experience Generation Module
[0494] 3. Interaction Module
[0495] 2. Program Processing
[0496] Data Acquisition and Analysis Module
[0497] First, users access the system using a terminal and create an account. After logging in, they upload photos, documents, and audio recordings (e.g., photos from a family trip or audio recordings of conversations) related to past memories.
[0498] The terminal transmits the uploaded data to the server.
[0499] The server securely stores the received data and analyzes it using AI algorithms. For example, it can recognize landscapes and people from photo data and extract the location and time of the photo. It also uses voice recognition technology to convert audio data into text.
[0500] External Data Collection
[0501] The server also retrieves relevant complementary information from external databases (e.g., weather databases, news archives), collecting data about the day's weather and important events.
[0502] This allows user-provided data and external information to be integrated consistently, laying the foundation for more realistic reconstructions of past events.
[0503] Experience Generation Module
[0504] The server combines the collected metadata with external information and uses AI models to generate virtual time-travel scenarios, such as a day in the life of a family vacation or a detailed description of specific events.
[0505] Based on the generated scenario, visual content (e.g., 3D models of old landscapes and people) and audio content (e.g., conversations and environmental sounds) are generated.
[0506] Interaction Module
[0507] The server transmits the completed visual and audio content to the terminal.
[0508] The device provides users with an experience through VR goggles and speakers, allowing them to virtually relive scenes from a past family trip, for example, and hear the conversations and environmental sounds from that time.
[0509] After the experience, users provide feedback via the device, including their impressions of the experience and requests for improvement.
[0510] The server analyzes the collected feedback and uses it to generate the next experience, improving the accuracy of the system and user satisfaction.
[0511] Specific examples
[0512] Example: Recreating family trip memories
[0513] 1. The user uploads photos, voice messages, and social media posts from tourist spots they have visited with their family to the system.
[0514] Examples: A photo taken at the beach on a summer day, or an audio recording of a family barbecue.
[0515] 2. The server analyzes this data and retrieves the day's weather (e.g., sunny with occasional clouds, temperature 30°C) and local newspaper articles (e.g., that a local festival was held) from an external database.
[0516] 3. The server creates a scenario of a day at the beach based on the retrieved metadata and complementary information. For example, spending the morning at the beach, having a family barbecue in the afternoon, and then going to a local festival.
[0517] 4. Based on the generated scenario, the server generates visual and audio content that realistically recreates the beach scenery and festival scenes.
[0518] 5. The device provides users with an experience through VR goggles, allowing them to experience the beach scenery right before their eyes and hear the conversations from past barbecues and the sounds of festivals in a realistic way.
[0519] 6. Users provide feedback after their experience, and the system uses that feedback to further improve the experience next time.
[0520] The present invention allows users to re-experience past memories in vivid detail, providing a meaningful virtual time travel experience for each individual user.
[0521] The processing flow will be explained below.
[0522] Step 1:
[0523] Users access the system using a terminal and log in. They select and upload past photos, documents, and audio recordings.
[0524] Step 2:
[0525] The device sends the data uploaded by the user to the server, including photo files, audio files, text files, etc.
[0526] Step 3:
[0527] The server stores the received data in a secure database and begins analyzing each piece of data, using AI algorithms to extract landscape and person metadata (e.g., date and time of photo capture, location), and convert the audio data into text.
[0528] Step 4:
[0529] The server makes requests to external APIs to retrieve relevant information from public databases, for example, weather data for a particular date and location, or newspaper articles for that day.
[0530] Step 5:
[0531] The server synthesizes the collected and analyzed data and generates scenarios for virtual time travel experiences, using AI models to create a coherent timeline based on the user's experiences. For example, it generates a scenario where the user arrives at the beach at 10:00 AM and has lunch at 1:00 PM.
[0532] Step 6:
[0533] The server generates visual content based on the scenario. It uses image generation models to create 3D models of past landscapes and people. Specifically, it realistically recreates beach sand, ocean views, and the user's family.
[0534] Step 7:
[0535] The server also generates audio content based on the scenario, using a speech generation model to realistically recreate past conversations and environmental sounds, providing an audio experience that includes, for example, the sounds of waves, wind, and conversations with family members.
[0536] Step 8:
[0537] The server then sends the finished visual and audio content to the device, which includes 4K resolution video and high-quality audio files to the user.
[0538] Step 9:
[0539] The device provides users with a virtual experience using VR goggles and high-quality speakers. Users can experience visual content through the VR goggles they wear and listen to audio content through the speakers.
[0540] Step 10:
[0541] After the experience, users can enter their feedback, such as their impressions of the experience or requests for improvements, into a dedicated form.
[0542] Step 11:
[0543] The terminal transmits the user's feedback to the server.
[0544] Step 12:
[0545] The server analyzes the received feedback and uses it as data to improve the system, improving the accuracy of the AI model and reflecting it in the next experience generation.
[0546] The above are the specific processing steps for allowing the user to vividly re-experience past memories and events.
[0547] Example 1
[0548] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0549] Previous technologies for recreating past memories lacked the accuracy and consistency of information analysis, making it difficult for users to vividly relive events that they actually experienced. Furthermore, integration with external information was insufficient, resulting in an inconsistent reproduction of past events. Furthermore, there was a lack of a way to effectively utilize user feedback to improve the experience.
[0550] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0551] In this invention, the server includes means for receiving past multimedia data from a user, means for analyzing the data and extracting metadata, means for obtaining related supplemental information from an external database, means for generating a scenario for a virtual time travel experience based on the analysis and supplemental information, means for providing prompts to a generative AI model to generate details of the scenario, means for generating visual and audio content based on the generated scenario, means for providing the generated content to the user, and means for collecting, analyzing, and reflecting user feedback in the generation of the next experience. This allows the user to vividly relive past memories and enjoy a realistic virtual time travel experience while maintaining consistency with external information. Furthermore, by utilizing feedback, the experience can be continuously improved.
[0552] A "user" is an individual or group who uses the system to relive past memories.
[0553] "Multimedia data" refers to digital data in various forms, such as photographs, documents, and audio recordings.
[0554] The "server" is a computer system that performs the core processing of the system, such as analyzing received data, obtaining supplementary information, generating scenarios, and creating visual and audio content.
[0555] "Metadata" refers to information that accompanies the main data, such as a photograph or audio data (e.g., the location where the photograph was taken, the time, recognized people or scenery, etc.).
[0556] An "external database" is a database located outside the system that is a source of information that provides supplementary information such as weather data and news data.
[0557] "Virtual time travel experience" refers to a virtual reality experience provided by a system that allows users to re-experience past events and environments through visual and audio experiences.
[0558] A "scenario" refers to a description or story that includes the specific sequence and details of past events that a user re-experiences.
[0559] "Generative AI model" refers to artificial intelligence technology for generating scenario details and descriptions based on prompt text.
[0560] A "prompt" is an input document that provides instructions to a generative AI model and provides a starting point for generating scenarios and explanatory text.
[0561] "Visual Content" refers to the visual elements, such as footage, images, and 3D models, provided in the virtual time travel experience.
[0562] "Audio Content" refers to the audio elements, such as music, sound effects, and dialogue, that are provided in a virtual time travel experience.
[0563] "Feedback" refers to information such as evaluations, impressions, and requests for improvement provided by users after experiencing a product.
[0564] The system of the present invention aims to provide a virtual time travel experience that allows users to vividly relive past memories. This system is mainly composed of three entities: a user, a terminal, and a server. Specific embodiments for implementing this system are described in detail below.
[0565] 1. Data Collection and Analysis
[0566] Users log in to the system via a dedicated application or website. After logging in, they upload multimedia data such as photos, documents, and audio recordings related to past memories via their device. The device automatically sends the uploaded data to the server, which stores it in secure storage. The server then uses AI algorithms to analyze the photo data, identifying scenes and people, and extracting metadata such as the location and time of the photo. Similarly, audio data is converted into text using voice recognition technology, and the content is analyzed.
[0567] 2. Acquiring and integrating external data
[0568] The server queries external databases (e.g., weather databases, news archives) based on the provided metadata to obtain complementary information. This allows for the collection of complementary information that is consistent with the user's multimedia data, creating a more realistic experience.
[0569] 3. Generation of Virtual Time Travel Scenario
[0570] The server combines the acquired metadata and complementary information to generate a virtual time travel scenario. This process utilizes a generative AI model, specifically using prompts such as:
[0571] "May 15, 2022, 3pm, on a beach in Tokyo"
[0572] Based on this prompt, the AI model generates a detailed scenario, such as "It's 3 p.m. and the beach is crowded with tourists, the temperature is 25 degrees, you can feel the sea breeze, and you can hear the laughter of family members all around."
[0573] 4. Visual and audio content generation
[0574] The server creates visual and audio content based on the generated scenario. The visual content is generated using a 3D graphics engine, and the audio content is reproduced realistically using voice synthesis technology.
[0575] 5. Providing a user experience
[0576] The device then provides the generated visual and audio content to the user through VR goggles and speakers, allowing the user to vividly and realistically re-experience past events, such as the beach scenery, conversations, and environmental sounds of a past family trip.
[0577] 6. Feedback Collection and Analysis
[0578] After the experience, the user provides feedback via their device. This feedback includes impressions of the experience and requests for improvements. The server collects and analyzes this feedback to help improve the system. By reflecting this feedback in the generation of the next experience, it becomes possible to increase user satisfaction.
[0579] In this way, the system of the present invention provides a virtual time travel experience that allows users to vividly relive past memories.
[0580] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0581] Step 1:
[0582] To log in to the system, users access a dedicated application or website. They enter their login information and complete authentication. Then, they upload multimedia data such as photos, documents, and audio recordings related to past memories. The input is various digital files selected by the user, and the output is data packets sent to the terminal.
[0583] Step 2:
[0584] The terminal sends the received multimedia data to the server. The terminal uses a data transfer protocol to send the files uploaded by the user to the server and add them to a queue for processing. The input is the file uploaded to the terminal, and the output is the data sent to the server.
[0585] Step 3:
[0586] The server stores the data in secure storage as soon as it receives it. Next, it uses AI algorithms to identify landscapes and people in the photo data and extract metadata such as the location and time of the photo. The audio data is converted to text using speech recognition technology and the content is analyzed. The input is the data sent from the device, and the output is the analyzed metadata and the converted audio data.
[0587] Step 4:
[0588] The server queries external databases based on the extracted metadata to obtain related complementary information. Specifically, it sends requests via API to weather databases, news archives, etc. to obtain the necessary information. The input is metadata (e.g., date and time of shooting, location), and the output is complementary information obtained from the external database.
[0589] Step 5:
[0590] The server combines internal data with external supplementary information and provides prompts to the generative AI model, which then generates a virtual time travel scenario. The prompts specify specific dates, times, and locations, and the AI model generates detailed scenarios based on them. The input is the prompts combined with metadata and supplementary information, and the output is the generated scenario.
[0591] For example, enter the prompt text as "May 15, 2022, 3:00 PM, on a beach in Tokyo."
[0592] Step 6:
[0593] The server creates visual and audio content based on the generated scenario. The visual content is realistically reproduced using a 3D graphics engine, and the audio content is realistically reproduced using speech synthesis technology. The input is the generated scenario, and the output is the visual and audio content provided to the user.
[0594] Step 7:
[0595] The terminal provides the user with visual and audio content sent from the server. Specifically, it uses VR goggles and speakers to make the experience feel realistic. The input is the content sent from the server, and the output is the virtual time travel experience provided to the user.
[0596] Step 8:
[0597] After completing the virtual time travel experience, the user provides feedback. The feedback includes impressions of the experience and suggestions for improvement, and is sent to the server via the terminal. The input is the user's impressions and suggestions for improvement, and the output is the feedback data sent to the server.
[0598] Step 9:
[0599] The server analyzes the received feedback and reflects it in the next experience generation, thereby improving user satisfaction. The input is the feedback data, and the output is the analysis results that will help improve the system.
[0600] (Application example 1)
[0601] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0602] Previous time travel experience systems have had limitations in the realism and detail of the visual and audio content when recreating past memories based on user-provided data. Furthermore, there has been a lack of methods to improve the consistency and accuracy of the generated content. This has resulted in a lack of realism and detailed reproduction in the virtual time travel experience that users can feel.
[0603] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0604] In this invention, the server includes a means for generating content to reproduce past memories provided by the user in more detail and more realistically, a means for creating optimal prompt sentences to be input into a generative AI model and generating content based on the prompt sentences, and a means for generating visual and audio data to reproduce a specific date and place in the past based on data provided by the user, thereby enabling the user to virtually re-experience past memories with high accuracy and realism.
[0605] A "user" is an individual who provides data such as photographs, documents, and audio recordings to use the system to recreate past memories as a virtual time travel experience.
[0606] "Photos" are image data taken by the user in the past, and serve as materials for generating visual content in the virtual time travel experience.
[0607] A "document" is text data created by a user in the past, and is used as auxiliary information in generating a scenario for virtual time travel.
[0608] "Audio recordings" are audio data recorded by users in the past, and are used to generate audio content for virtual time travel experiences.
[0609] "Metadata" is information extracted from photographs, documents, and audio recordings, and is additional information such as date, time, place, and people that is necessary to generate a scenario for a virtual time travel experience.
[0610] An "external database" is an external data source that provides additional information to complement the user's past memories, such as weather data or news articles.
[0611] The term "virtual time travel experience" refers to a user reliving past memories using virtual reality or augmented reality technology.
[0612] An "AI model" is an algorithm that uses machine learning techniques to analyze and integrate user-provided data to generate a virtual time travel experience.
[0613] A "prompt" is a text sentence input into a generative AI model that specifically instructs the scenario and visual and audio content of the virtual time travel experience.
[0614] "Visual content" means the visual representations, such as images and 3D models, that a user can view in the virtual time travel experience.
[0615] "Audio content" refers to the acoustic representation of conversations, environmental sounds, and other sounds that a user can hear during a virtual time travel experience.
[0616] "Feedback" refers to reaction data such as impressions and requests for improvement provided by a user after completing the virtual time travel experience.
[0617] "Content generation means" means the technical means for generating visual and audio content using AI models based on data provided by the user and complementary information from external databases.
[0618] The present invention is a system that analyzes past photographs, documents, and audio recordings provided by a user to generate a virtual time travel experience. This system is mainly composed of a user, a terminal, and a server, and is realized by the following processing steps.
[0619] 1. System Overview
[0620] The system consists of a data collection and analysis module, an experience generation module, and an interaction module.
[0621] 1.1 Data Collection and Analysis Module
[0622] Users upload photos, documents, and audio recordings related to past memories to the system. The device then sends the uploaded data to the server, which then analyzes it using AI algorithms. Scenes and people are recognized from the photos, and the location and time of the photo are extracted. Audio data is also converted into text using voice recognition technology.
[0623] 1.2 External Data Collection
[0624] The server retrieves relevant complementary information from external databases (e.g., weather databases, news archives), allowing the user-provided data to be integrated with external information to more realistically recreate past events.
[0625] 1.3 Experience Generation Module
[0626] The server integrates the collected metadata with external information and uses an AI model to generate a virtual time travel scenario, including a detailed description of a day in the family trip and specific events, and generates visual content (e.g., 3D models of historical landscapes and people) and audio content (e.g., conversations and environmental sounds).
[0627] 1.4 Interaction Modules
[0628] The server then sends the generated visual and audio content to the device, providing the experience to the user through VR goggles and speakers. The user can virtually relive scenes from a past family trip and hear the conversations and surrounding environmental sounds. After the experience, the user provides feedback, which the server analyzes and reflects in the generation of the next experience.
[0629] 2. Hardware and Software Used
[0630] The system uses the following hardware and software:
[0631] Hardware: Personal computer, VR goggles, smartphone, speakers
[0632] Software: Google Cloud Vision API (image analysis), Google Cloud Speech-to-Text API (voice recognition), cloud storage service
[0633] 3. Examples of concrete examples and prompts
[0634] Examples:
[0635] Users upload photos, voice messages, and social media posts from tourist spots they've visited with their family to the system. The system analyzes this data and retrieves the day's weather and newspaper articles from an external database. Based on the retrieved metadata and supplementary information, the system creates a scenario of a day at the beach and generates visual and audio content that realistically recreates beach scenes and festival scenes.
[0636] Example prompt sentence:
[0637] "Using user-provided data, recreate a family beach trip on August 15, 2005."
[0638] This system allows users to virtually re-experience past memories with high accuracy and realism.
[0639] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0640] Step 1:
[0641] Users access the system using a terminal and upload photos, documents, and audio recordings related to past memories. These data are used as input to generate a virtual time travel experience. The input for this step is the data provided by the user (photos, documents, audio recordings), and the output is this data sent to the server.
[0642] Step 2:
[0643] The device sends the data uploaded by the user to the server, which receives and securely stores this data. The input of this step is the data sent from the device, and the output is the data stored on the server.
[0644] Step 3:
[0645] The server analyzes the stored data and extracts metadata using AI algorithms. For example, it recognizes landscapes and people in photos and extracts the location and time of the photo, and converts audio data into text using voice recognition technology. The input for this step is the data stored on the server, and the output is the extracted metadata.
[0646] Step 4:
[0647] The server retrieves relevant complementary information from external databases (e.g., weather databases, news archives), providing complementary information that is consistent with the user's data. The input of this step is the extracted metadata, and the output is the retrieved complementary information.
[0648] Step 5:
[0649] The server integrates the collected metadata with external information and uses an AI model to generate a virtual time travel scenario. This scenario includes detailed descriptions of the day and specific events. The input of this step is the integrated metadata and complementary information, and the output is the generated scenario.
[0650] Step 6:
[0651] The server generates visual and audio content based on the generated scenario. The visual content includes 3D models of historical landscapes and people, and the audio content includes conversations and environmental sounds. The input of this step is the generated scenario, and the output is the visual and audio content.
[0652] Step 7:
[0653] The server transmits the generated visual and audio content to the terminal, where the user can experience virtual time travel using VR goggles and speakers. The input of this step is the generated visual and audio content, and the output is the provision of content to the user.
[0654] Step 8:
[0655] After the user finishes the virtual time travel experience, they provide feedback about the experience. The device sends this feedback to the server, which analyzes it. The input of this step is the user's feedback, and the output is the analyzed feedback data.
[0656] Step 9:
[0657] The server then refines the system based on the analyzed feedback, which then incorporates new knowledge into the next experience generation, improving the system's accuracy and user satisfaction. The input of this step is the analyzed feedback data, and the output is an improved system.
[0658] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0659] To implement the present invention, a user must provide past photographs, documents, and audio recordings to the system, which then analyzes and integrates this data to generate a virtual time travel experience. Furthermore, the present invention incorporates an emotion engine that recognizes the user's emotions, further personalizing the user's experience. The following describes in detail the modes for implementing the invention.
[0660] 1. System Overview
[0661] This system is based on users, terminals, and servers, and is composed of the following major functional modules:
[0662] 1. Data Collection and Analysis Module
[0663] 2. Experience Generation Module
[0664] 3. Interaction Module
[0665] 4. Emotion Engine
[0666] 2. Program Processing
[0667] Data Acquisition and Analysis Module
[0668] First, a user accesses the system using a terminal and logs in. The user uploads photos, documents, and audio recordings (e.g., photos from a family trip or audio of a conversation) related to past memories.
[0669] The terminal transmits the uploaded data to the server.
[0670] The server securely stores the received data and analyzes it using AI algorithms. For example, it can recognize landscapes and people from photo data and extract the location and time of the photo. It also uses voice recognition technology to convert audio data into text.
[0671] External Data Collection
[0672] The server makes requests to external APIs to retrieve relevant information from public databases, for example, weather data for a particular date and location, or newspaper articles for that day.
[0673] This allows user-provided data and external information to be integrated consistently, laying the foundation for more realistic reconstructions of past events.
[0674] Experience Generation Module
[0675] The server combines the collected metadata with external information and uses AI models to generate virtual time-travel scenarios, such as a day in the life of a family vacation or a detailed description of specific events.
[0676] Based on the generated scenario, visual content (e.g., 3D models of old landscapes and people) and audio content (e.g., conversations and environmental sounds) are generated.
[0677] Emotion Engine
[0678] The server recognizes emotions using photos, voice recordings, and text data provided by the user. The emotion engine analyzes the user's emotional state (e.g., joy, sadness, nostalgia) extracted from the data.
[0679] This allows the scenario of the virtual time travel experience to be adjusted according to the user's emotions, providing a more personalized experience.
[0680] Interaction Module
[0681] The server transmits the completed visual and audio content to the terminal.
[0682] The device provides users with an experience through VR goggles and speakers, allowing them to virtually relive scenes from a past family trip, for example, and hear the conversations and environmental sounds from that time.
[0683] Feedback collection
[0684] After the experience, users provide feedback via the device, including their impressions of the experience and requests for improvement.
[0685] The server analyzes the collected feedback and uses it to generate the next experience, improving the accuracy of the system and user satisfaction.
[0686] Specific examples
[0687] Example: Recreating family trip memories
[0688] 1. The user uploads photos, voice messages, and social media posts from tourist spots they have visited with their family to the system.
[0689] Examples: A photo taken at the beach on a summer day, or an audio recording of a family barbecue.
[0690] 2. The server analyzes this data and retrieves the day's weather (e.g., sunny with occasional clouds, temperature 30°C) and local newspaper articles (e.g., that a local festival was held) from an external database.
[0691] 3. The server creates a scenario of a day at the beach based on the retrieved metadata and complementary information. For example, spending the morning at the beach, having a family barbecue in the afternoon, and then going to a local festival.
[0692] 4. The server uses an emotion engine to read emotions from the user's uploaded data. Example: Identifying the user's emotion of "joy" from a photo of them at the beach.
[0693] 5. The server reflects the user's emotions in the generated scenario, for example, creating a scenario that emphasizes the joy of being at the beach or creating audio content that enhances the fun of having a barbecue.
[0694] 6. The server transmits the generated visual and audio content to the terminal.
[0695] 7. The device provides users with an experience through VR goggles, allowing them to experience the beach scenery right before their eyes and hear the conversations from past barbecues and the sounds of festivals in a realistic way.
[0696] 8. Users provide feedback after their experience, and the system uses that feedback to further improve the experience next time.
[0697] The present invention allows users to re-experience past memories in vivid detail, and the emotion engine provides a more personalized virtual time travel experience.
[0698] The processing flow will be explained below.
[0699] Step 1:
[0700] Users access the system using a terminal, create an account and log in. Users select past photos, documents and audio recordings and upload them to the system.
[0701] Step 2:
[0702] The device sends the data uploaded by the user to the server, including photo files, audio files, and text files.
[0703] Step 3:
[0704] The server stores the received data in a secure database, after which it begins analyzing each piece of data using AI algorithms.
[0705] Step 4:
[0706] The server analyzes the photo data and extracts metadata such as the landscape and people. For example, it uses image recognition technology to identify the location and date of the photo.
[0707] Step 5:
[0708] The server analyzes the audio data, converts the audio content into text, and uses voice recognition technology to identify the content of the conversation and who is speaking, saving it as metadata.
[0709] Step 6:
[0710] The server analyzes the text data provided by the user and applies emotion recognition algorithms to extract the user's emotional state (e.g., joy, sadness, nostalgia).
[0711] Step 7:
[0712] The server accesses external databases to gather additional information relevant to a particular date and location, such as weather data or newspaper articles for that day.
[0713] Step 8:
[0714] The server combines the collected metadata with external information to generate a scenario for a virtual time-travel experience, using an AI model to create a timeline with specific actions, such as "playing at the beach in the morning and having a family barbecue in the afternoon."
[0715] Step 9:
[0716] The server uses an emotion engine to reflect the user's emotions in the generated scenario. For example, it generates vivid and bright visuals to emphasize scenes in which the user felt "joy."
[0717] Step 10:
[0718] The server generates visual content, using a generative image model to create 3D models of historical landscapes and people, such as beach sand, ocean views, or even the user's family.
[0719] Step 11:
[0720] The server generates audio content, using a speech generation model to realistically recreate past conversations and environmental sounds (e.g., the sound of waves or family conversations).
[0721] Step 12:
[0722] The server then transmits the generated visual and audio content to the device, including 4K resolution video files and high-quality audio files.
[0723] Step 13:
[0724] The device provides the received content to the user, who then puts on the VR goggles and begins the experience using high-quality speakers.
[0725] Step 14:
[0726] Users can relive past memories through virtual experiences, such as a beach scene unfolding before their eyes, with the sound of waves and family conversations realistically recreated.
[0727] Step 15:
[0728] After the experience, users can enter their feedback, sending their impressions of the experience and requests for improvements to the system through a dedicated feedback form.
[0729] Step 16:
[0730] The terminal transmits the user's feedback to the server.
[0731] Step 17:
[0732] The server analyzes the received feedback and uses it as data to improve the system. Based on the analysis results, the accuracy of the AI model is improved and reflected in the next experience generation.
[0733] In this way, the present invention allows users to relive their past memories vividly and in detail, providing a personalized virtual time travel experience with an emotion engine.
[0734] Example 2
[0735] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0736] Conventional virtual experience systems have difficulty integrating user-provided data with external information to generate a realistic experience. Furthermore, because personalization that takes user emotions into account is not performed, it is not possible to provide an optimal experience for each individual user. Therefore, a system is needed that analyzes the diverse data provided by users and integrates it with external information to provide an individually optimized virtual experience that reflects the user's emotions.
[0737] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a means for receiving past image data, text data, and audio data from the user, a means for analyzing this data and extracting metadata, and a means for acquiring related supplementary information from an external information source. This makes it possible to provide a realistic and personalized virtual experience while taking into account the user's emotions.
[0738] A "user" is a person who uses the virtual experience system.
[0739] "Image data" refers to digital data that contains visual information such as photographs and illustrations.
[0740] "Text data" refers to digital data that contains information in the form of sentences or characters.
[0741] "Audio data" refers to digital data that contains auditory information such as human voices and environmental sounds.
[0742] "Metadata" is data that represents information related to data, and is information that serves as the basis for analysis and integration.
[0743] "External sources" are resources that provide additional information, such as databases or APIs outside the system.
[0744] A "scenario" is the story or sequence of events that make up the overall picture of a virtual experience.
[0745] "Visual content" refers to visual information such as images and videos that are presented to a user.
[0746] "Audio content" refers to auditory information such as music and sound effects that is presented to the user.
[0747] "Feedback" refers to impressions and suggestions for improvement provided by users of the virtual experience system.
[0748] "Emotion" refers to information that represents the user's psychological state, such as joy, sadness, or nostalgia.
[0749] "Personalization" is the process of optimizing the content and presentation of an experience for each individual user.
[0750] To implement this invention, the user must provide the system with previously captured image data, text data, and audio data, which is then analyzed and integrated to generate a virtual experience. In addition, the invention incorporates an emotion engine that recognizes the user's emotions, allowing for a more personalized experience.
[0751] System configuration
[0752] This system is based on users, terminals, and servers, and is composed of the following main functional modules:
[0753] 1. Data Collection and Analysis Module
[0754] 2. Experience Generation Module
[0755] 3. Interaction Module
[0756] 4. Emotion Engine
[0757] Data Acquisition and Analysis Module
[0758] First, the user accesses the system using a terminal and logs in. The user then uploads image data, text data, and audio data related to past memories.
[0759] The terminal transmits the uploaded data to the server.
[0760] The server securely stores the received data and analyzes each piece of data using AI algorithms. For example, it recognizes landscapes and people from photo data and extracts the location and time of the photo. It also uses voice recognition technology to convert audio data into text. The specific software used is an "image recognition API" for image analysis and a "voice recognition API" for audio analysis.
[0761] External Data Collection
[0762] The server sends requests to external APIs to retrieve relevant information from public databases, such as weather data for a specific date and location, or news articles for that day. Specifically, it uses the "Weather API" for weather information and the "News API" for news articles.
[0763] This allows user-provided data and external information to be integrated consistently, laying the foundation for more realistic reconstructions of past events.
[0764] Experience Generation Module
[0765] The server combines the collected metadata with external information and generates a scenario for the virtual experience using a generative AI model. For example, it creates a scenario that includes a detailed description of a day on a family trip or a specific event. The AI model used is a "generative AI model (such as GPT-4)."
[0766] Based on the generated scenario, visual content (e.g., 3D models of old landscapes and people) and audio content (e.g., conversations and environmental sounds) are generated. Visual content is generated using "visual content generation software (e.g., 3D modeling software)," and audio content is generated using "audio editing software."
[0767] Emotion Engine
[0768] The server recognizes emotions using image data, audio data, and text data provided by the user. The emotion engine analyzes the user's emotional state (e.g., joy, sadness, nostalgia) extracted from the data. An "emotion analysis API" is used for emotion analysis.
[0769] This allows the virtual experience scenario to be adjusted according to the user's emotions, providing a more personalized experience.
[0770] Interaction Module
[0771] The server transmits the completed visual and audio content to the terminal.
[0772] The device provides users with an experience through VR goggles and speakers, allowing them to virtually relive scenes from a past family trip, for example, and hear the conversations and environmental sounds from that time.
[0773] Feedback collection
[0774] After the experience, users provide feedback via their device, including their impressions of the experience and requests for improvement.
[0775] The server analyzes the collected feedback and uses it to generate the next experience, improving the accuracy of the system and user satisfaction.
[0776] Specific examples
[0777] Example: Recreating family trip memories
[0778] 1. The user uploads image data, voice messages, and social media posts of tourist spots visited with their family to the system.
[0779] Examples: a photo taken at a place you visited on a summer day, an audio recording of you having a meal with your family.
[0780] 2. The server analyzes this data and retrieves the day's weather (e.g., sunny with occasional clouds, temperature 30°C) and local news articles (e.g., local events being held) from an external database.
[0781] 3. The server creates a scenario of a day in a specific location based on the retrieved metadata and complementary information, e.g., visiting tourist spots in the morning, having a family meal in the afternoon, and then going to a local event.
[0782] 4. The server uses an emotion engine to read emotions from the user's uploaded data. For example, it identifies the user's "joy" emotion from a photo.
[0783] 5. The server reflects the user's emotions in the generated scenario. For example, it creates a scenario that emphasizes the joy of a tourist spot or audio content that enhances the enjoyment of a meal.
[0784] 6. The server transmits the generated visual and audio content to the terminal.
[0785] 7. The device provides users with an experience through VR goggles, allowing them to experience the scenery of tourist spots right before their eyes and hear the sounds of past conversations and events in a realistic way.
[0786] 8. Users provide feedback after their experience, and the system uses that feedback to further improve the experience next time.
[0787] Prompt Sentence Examples
[0788] Below is an example of a prompt sentence to input to the generative AI model.
[0789] I've uploaded photos and audio recordings from past family trips. Please use these to generate a virtual experience. I'd particularly like to emphasize the emotion of joy. Include supplemental information like the weather for that day and local events.
[0790] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0791] Step 1:
[0792] The user accesses the system using a terminal and logs in by entering their ID and password on the login page.
[0793] Input: ID, password
[0794] Output: User credentials
[0795] Specific operation: The device sends the login information to the server, and the server verifies the authentication information and allows the user to log in. If authentication is successful, the user's dashboard is displayed.
[0796] Step 2:
[0797] Users select image data, text data, and audio data related to past memories from the dashboard and press the upload button to upload this data to the system.
[0798] Input: image data, text data, audio data
[0799] Output: Upload data
[0800] Specific operation: The terminal divides the selected data into packets and sends them to the server.
[0801] Step 3:
[0802] The terminal transmits the uploaded data to the server.
[0803] Input: Upload data
[0804] Output: Transmitted data
[0805] Specific operation: The terminal converts the data into the appropriate format for each type (image, audio, text) and sends it to the server.
[0806] Step 4:
[0807] The server will securely store the received data in the specified directory.
[0808] Input: Send data
[0809] Output: Saved data
[0810] Specific operation: The server changes the file name and saves it in the appropriate folder depending on the type of data.
[0811] Step 5:
[0812] The server analyzes the data using AI algorithms such as image recognition, voice recognition, and text analysis.
[0813] Input: Saved data
[0814] Output: Metadata
[0815] Specific operation: Extracts scenery and people from photo data and converts audio data into text. Specifically, it uses "image recognition API" and "voice recognition API." Analyzed metadata includes location, time, and characters.
[0816] Step 6:
[0817] The server sends external API requests to retrieve relevant information from designated public sources.
[0818] Input: Metadata
[0819] Output: Complementary information
[0820] Specific operation: Compose a query for the information you want to obtain (weather, news articles, etc.) and access external sources. For example, use the "Weather API" for weather information and the "News API" for news articles.
[0821] Step 7:
[0822] The server integrates the collected metadata with external information and generates a virtual experience scenario using a generative AI model.
[0823] Input: Metadata, complementary information
[0824] Output: Scenario data
[0825] Specific operation: Automatically generates a storyboard based on the specified date and location based on user-provided data and external information. A generative AI model (such as GPT-4) is used as the generative AI model.
[0826] Step 8:
[0827] The server creates visual and audio content based on the generated scenario.
[0828] Input: Scenario data
[0829] Output: Visual content, audio content
[0830] What it does: Visual content is generated using visual content generation software (e.g., 3D modeling software) and audio content is generated using audio editing software, which generates the detailed elements of the virtual experience.
[0831] Step 9:
[0832] The server uses an emotion engine to analyze emotions from the image data, audio data, and text data provided by the user.
[0833] Input: image data, audio data, text data
[0834] Output: Emotion data
[0835] Specific operation: Emotion analysis uses the "Emotion Analysis API" to identify the user's psychological state. For example, emotions such as "joy" and "sadness" can be extracted from a photo.
[0836] Step 10:
[0837] Based on the analysis results, the server adjusts the experience scenario to match the user's emotions.
[0838] Input: Emotion data, scenario data
[0839] Output: personalized scenario data
[0840] Specific behavior: Adding music and effects that correspond to emotions and optimizing the content of the scenario for each individual user, providing an experience that emphasizes specific emotions.
[0841] Step 11:
[0842] The server transmits the completed visual and audio content to the terminal.
[0843] Input: Visual content, audio content
[0844] Output: Send content
[0845] Specific operation: The server acts as a file server and provides streaming or download links to devices.
[0846] Step 12:
[0847] The device provides users with an experience using VR goggles and speakers.
[0848] Input: Submit content
[0849] Output: Experience data
[0850] How it works: Users can enjoy a virtual experience through a dedicated VR app, and past experiences are realistically recreated for the user.
[0851] Step 13:
[0852] After the experience, users provide feedback via the device.
[0853] Input: Experience data
[0854] Output: Feedback data
[0855] What to do: Enter your thoughts and suggestions for improvement in the feedback form and submit it. Use the text fields and rating sliders.
[0856] Step 14:
[0857] The server analyzes the collected feedback using machine learning models and reflects it in improvements to the system.
[0858] Input: Feedback data
[0859] Output: Improvement data
[0860] Specific actions: Perform natural language processing of feedback to discover new areas for improvement. Specific software used includes "natural language processing software."
[0861] (Application example 2)
[0862] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0863] In today's world, there is a growing demand for new experiences that vividly recreate past memories. However, existing technologies simply display or play back past photographs and audio recordings, making it difficult to integrate this information and provide a deeply personalized virtual experience. Furthermore, they are unable to tailor the experience to the user's emotions, creating a need for a system that allows users to relive the emotions of past memories. The present invention aims to solve these problems and provide a virtual experience system that realistically recreates past experiences while recognizing the user's emotions.
[0864] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for receiving the user's past photos, documents, and voice recordings, analyzing this data, and extracting metadata; means for acquiring related complementary information from an external database; means for generating a scenario for a virtual time travel experience based on the analysis and complementary information; means for generating visual and audio content based on the generated scenario; and means for recognizing the user's emotions and adjusting the experience scenario and content accordingly. This enables the user to recreate realistic past experiences personalized according to their emotions.
[0865] "Means for recognizing a user's emotions and adjusting the experience scenario and content accordingly" refers to technology that uses artificial intelligence to analyze a user's emotional state from photos, documents, audio recordings, etc. provided by the user, and personalizes the content of the virtual experience based on those emotions.
[0866] "Generative AI models" are artificial intelligence algorithms and frameworks that perform natural language processing and image generation, examples of which include GPT-4 and DALL-E.
[0867] "Metadata" is data that provides information about the original data, such as the date and location when a photo was taken, and the contents of audio recordings.
[0868] "External databases" refer to data resources and databases that are publicly available on the Internet, including weather information and past newspaper articles.
[0869] A "virtual time travel experience scenario" is a scenario generated by artificial intelligence based on the user's past records and related information, and refers to a script or plan of action that allows the user to recreate past experiences in virtual reality.
[0870] "Visual and audio content" refers to the visual and audio elements that provide the user with a virtual experience, such as generated 3D models, ambient sounds, and dialogue.
[0871] "Means for collecting, analyzing, and reflecting feedback in generating the next experience" refers to a system that records the impressions and opinions provided by users after an experience, analyzes them using artificial intelligence, and improves the next experience scenario.
[0872] "Complementary information obtained from external databases" is additional information collected to complement the data provided by the user, examples of which include weather information and event information.
[0873] 1. System Overview
[0874] This system is based on the user, terminal, and server, and consists of the following main functional modules that analyze and integrate the user's past photos, documents, and audio recordings to generate and provide a virtual time travel experience.
[0875] 1.1 Data Collection and Analysis Module
[0876] The server receives photos, documents, and audio recordings uploaded by users using their devices. It analyzes this data and extracts metadata about location, time, and content. It uses Python and TensorFlow to apply AI algorithms to recognize scenes and people in photos and convert audio data into text.
[0877] 1.2 External Data Collection Module
[0878] The server retrieves relevant complementary information from external databases, such as the day's weather information or newspaper articles from public databases on the Internet, using a REST API to retrieve this external data.
[0879] 1.3 Experience Generation Module
[0880] The server combines the collected metadata with external information and generates a virtual time travel scenario using a generative AI model (e.g., GPT-4, DALL-E). Based on the generated scenario, visual and audio content such as 3D models and environmental sounds are generated. This process is performed using Unity or Unreal Engine.
[0881] 1.4 Emotion Engine
[0882] The server recognizes emotions from the data uploaded by the user. The emotion engine uses a natural language processing (NLP) library to analyze the user's emotional state and incorporate it into the scenario. For example, it identifies the emotion of joy from a photo and incorporates it into the scene.
[0883] 1.5 Interaction Module
[0884] The server then sends the completed visual and audio content to the terminal, providing the user with a virtual experience, which can be enjoyed through VR goggles or smart glasses.
[0885] 1.6 Feedback Collection Module
[0886] The server collects and analyzes feedback from users, including impressions of the experience and requests for improvement. The analyzed feedback is used to generate the next experience.
[0887] 2. Specific Examples
[0888] Example: Recreating family vacation memories
[0889] 1. The user uploads photos and voice messages of tourist spots that they have visited with their family to the system.
[0890] Example: Photos of tourist spots you visited on a summer day and a recording of the conversations you had at the time.
[0891] 2. The server analyzes this data and retrieves information about the day's weather and local events from an external database.
[0892] 3. The server generates a family trip experience scenario based on the acquired metadata and supplementary information, including, for example, scenes of play at tourist spots and events of the day.
[0893] 4. The server uses an emotion engine to read emotions from the user's uploaded data. For example, it identifies the user's emotion of "joy" from photos of tourist spots.
[0894] 5. The server reflects the user's emotions in the generated scenario, for example, creating a scenario that emphasizes joy or audio content that enhances enjoyment.
[0895] 6. The server transmits the generated visual and audio content to the terminal.
[0896] 7. The device provides users with an experience through VR goggles, allowing them to experience the scenery of tourist spots right before their eyes and hear past conversations and environmental sounds in a realistic way.
[0897] 8. Users provide feedback after their experience, and the server uses that feedback to further improve the experience next time.
[0898] 3. Examples of prompts
[0899] 1. To generate a scenario that recreates family vacation memories, the following prompt sentence is input into the generative AI model:
[0900] "Generate an experience scenario of a past family trip based on photos and audio recordings of tourist spots visited by the family. The scenario should include the flow of a day enjoyed at tourist spots visited on a summer day. In particular, the photos convey the emotion of 'joy,' so please reflect this emotion in the scenario."
[0901] As described above, the present invention realizes a system that realistically reproduces a user's past experiences and provides a personalized virtual time travel experience that responds to emotions.
[0902] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0903] Step 1:
[0904] Users log in to the system using a terminal and upload photos, documents, and audio recordings related to past memories. This data is input and sent from the terminal to the server. Specifically, users open the upload screen on their device, such as a smartphone or PC, select the target data using the file selection button, and press the upload button.
[0905] Step 2:
[0906] The server receives data sent by the user and stores it securely. The received data is the input, and storing it internally on the server is the output. Specifically, the server's file storage system stores the data and adds a metadata record to the database.
[0907] Step 3:
[0908] The server analyzes the stored data, recognizes scenes and people in the photos, and converts audio data into text. This is done using Python and TensorFlow. Photo and audio data are the input, and extracted metadata is the output. Specifically, the image recognition algorithm identifies scenes and people in the photos and generates associated tags. At the same time, the speech recognition algorithm converts audio into text.
[0909] Step 4:
[0910] The server sends a request to the public database's API to obtain related information from the external database. For example, to obtain weather information or event information for a specified date and time. The input is location and date information related to the user's input data, and the output is weather information and event information obtained based on that. Specifically, the server sends an HTTP request to the external API, parses the JSON response, and extracts the required information.
[0911] Step 5:
[0912] The server uses a generative AI model (e.g., GPT-4, DALL-E) to generate a scenario for a virtual time travel experience based on the analyzed metadata and the acquired external information. The input is the metadata and external information, and the output is a virtual experience scenario. Specifically, the server sends a prompt to the AI model to generate the scenario. This prompt includes the user's data and related information.
[0913] Step 6:
[0914] The server generates visual and audio content based on the scenario. For example, it creates 3D models and environmental sounds. The input is the generated scenario, and the output is the visual and audio content. Specifically, it uses Unity or Unreal Engine to model the visual content and the audio engine to generate the environmental sounds.
[0915] Step 7:
[0916] The server reads emotions from the user's uploaded data, analyzes the emotional state using an emotion engine, and reflects it in the generated scenario. The input is the user's photo and voice data, and the output is a scenario that reflects the emotions. Specifically, it uses an NLP library to perform emotion analysis and add emotional expressions to the scenario text.
[0917] Step 8:
[0918] The server sends the completed visual and audio content to the device and provides it to the user. The input is the generated content, and the output is the user's experience. Specifically, the content is streamed to the user's device and displayed and played to the user through the device's VR goggles or smart glasses.
[0919] Step 9:
[0920] After the experience, the user provides feedback through the device. The input is their impressions of the experience and requests for improvement, and the output is feedback data. Specifically, the user enters their comments in the feedback form on the device and presses the send button.
[0921] Step 10:
[0922] The server analyzes the collected feedback and reflects it in the generation of the next experience. The input is feedback data, and the output is an improved experience scenario. Specifically, the server analyzes the feedback data with an analysis tool, obtains new insights, and reflects them in the next prompt.
[0923] Through the above processing steps, the user can recreate a realistic past experience that is personalized according to their emotions.
[0924] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0925] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0926] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0927] [Third embodiment]
[0928] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0929] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[0930] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0931] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0932] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0933] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0934] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0935] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0936] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0937] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0938] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0939] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0940] To implement the present invention, a user must provide historical photographs, documents, and audio recordings to the system, which then analyzes and integrates this data to generate a virtual time travel experience. The following describes in detail the mode for implementing the invention.
[0941] 1. System Overview
[0942] This system is based on users, terminals, and servers, and is composed of the following major functional modules:
[0943] 1. Data Collection and Analysis Module
[0944] 2. Experience Generation Module
[0945] 3. Interaction Module
[0946] 2. Program Processing
[0947] Data Acquisition and Analysis Module
[0948] First, users access the system using a terminal and create an account. After logging in, they upload photos, documents, and audio recordings (e.g., photos from a family trip or audio recordings of conversations) related to past memories.
[0949] The terminal transmits the uploaded data to the server.
[0950] The server securely stores the received data and analyzes it using AI algorithms. For example, it can recognize landscapes and people from photo data and extract the location and time of the photo. It also uses voice recognition technology to convert audio data into text.
[0951] External Data Collection
[0952] The server also retrieves relevant complementary information from external databases (e.g., weather databases, news archives), collecting data about the day's weather and important events.
[0953] This allows user-provided data and external information to be integrated consistently, laying the foundation for more realistic reconstructions of past events.
[0954] Experience Generation Module
[0955] The server combines the collected metadata with external information and uses AI models to generate virtual time-travel scenarios, such as a day in the life of a family vacation or a detailed description of specific events.
[0956] Based on the generated scenario, visual content (e.g., 3D models of old landscapes and people) and audio content (e.g., conversations and environmental sounds) are generated.
[0957] Interaction Module
[0958] The server transmits the completed visual and audio content to the terminal.
[0959] The device provides users with an experience through VR goggles and speakers, allowing them to virtually relive scenes from a past family trip, for example, and hear the conversations and environmental sounds from that time.
[0960] After the experience, users provide feedback via the device, including their impressions of the experience and requests for improvement.
[0961] The server analyzes the collected feedback and uses it to generate the next experience, improving the accuracy of the system and user satisfaction.
[0962] Specific examples
[0963] Example: Recreating family trip memories
[0964] 1. The user uploads photos, voice messages, and social media posts from tourist spots they have visited with their family to the system.
[0965] Examples: A photo taken at the beach on a summer day, or an audio recording of a family barbecue.
[0966] 2. The server analyzes this data and retrieves the day's weather (e.g., sunny with occasional clouds, temperature 30°C) and local newspaper articles (e.g., that a local festival was held) from an external database.
[0967] 3. The server creates a scenario of a day at the beach based on the retrieved metadata and complementary information. For example, spending the morning at the beach, having a family barbecue in the afternoon, and then going to a local festival.
[0968] 4. Based on the generated scenario, the server generates visual and audio content that realistically recreates the beach scenery and festival scenes.
[0969] 5. The device provides users with an experience through VR goggles, allowing them to experience the beach scenery right before their eyes and hear the conversations from past barbecues and the sounds of festivals in a realistic way.
[0970] 6. Users provide feedback after their experience, and the system uses that feedback to further improve the experience next time.
[0971] The present invention allows users to re-experience past memories in vivid detail, providing a meaningful virtual time travel experience for each individual user.
[0972] The processing flow will be explained below.
[0973] Step 1:
[0974] Users access the system using a terminal and log in. They select and upload past photos, documents, and audio recordings.
[0975] Step 2:
[0976] The device sends the data uploaded by the user to the server, including photo files, audio files, text files, etc.
[0977] Step 3:
[0978] The server stores the received data in a secure database and begins analyzing each piece of data, using AI algorithms to extract landscape and person metadata (e.g., date and time of photo capture, location), and convert the audio data into text.
[0979] Step 4:
[0980] The server makes requests to external APIs to retrieve relevant information from public databases, for example, weather data for a particular date and location, or newspaper articles for that day.
[0981] Step 5:
[0982] The server synthesizes the collected and analyzed data and generates scenarios for virtual time travel experiences, using AI models to create a coherent timeline based on the user's experiences. For example, it generates a scenario where the user arrives at the beach at 10:00 AM and has lunch at 1:00 PM.
[0983] Step 6:
[0984] The server generates visual content based on the scenario. It uses image generation models to create 3D models of past landscapes and people. Specifically, it realistically recreates beach sand, ocean views, and the user's family.
[0985] Step 7:
[0986] The server also generates audio content based on the scenario, using a speech generation model to realistically recreate past conversations and environmental sounds, providing an audio experience that includes, for example, the sounds of waves, wind, and conversations with family members.
[0987] Step 8:
[0988] The server then sends the finished visual and audio content to the device, which includes 4K resolution video and high-quality audio files to the user.
[0989] Step 9:
[0990] The device provides users with a virtual experience using VR goggles and high-quality speakers. Users can experience visual content through the VR goggles they wear and listen to audio content through the speakers.
[0991] Step 10:
[0992] After the experience, users can enter their feedback, such as their impressions of the experience or requests for improvements, into a dedicated form.
[0993] Step 11:
[0994] The terminal transmits the user's feedback to the server.
[0995] Step 12:
[0996] The server analyzes the received feedback and uses it as data to improve the system, improving the accuracy of the AI model and reflecting it in the next experience generation.
[0997] The above are the specific processing steps for allowing the user to vividly re-experience past memories and events.
[0998] Example 1
[0999] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1000] Previous technologies for recreating past memories lacked the accuracy and consistency of information analysis, making it difficult for users to vividly relive events that they actually experienced. Furthermore, integration with external information was insufficient, resulting in an inconsistent reproduction of past events. Furthermore, there was a lack of a way to effectively utilize user feedback to improve the experience.
[1001] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1002] In this invention, the server includes means for receiving past multimedia data from a user, means for analyzing the data and extracting metadata, means for obtaining related supplemental information from an external database, means for generating a scenario for a virtual time travel experience based on the analysis and supplemental information, means for providing prompts to a generative AI model to generate details of the scenario, means for generating visual and audio content based on the generated scenario, means for providing the generated content to the user, and means for collecting, analyzing, and reflecting user feedback in the generation of the next experience. This allows the user to vividly relive past memories and enjoy a realistic virtual time travel experience while maintaining consistency with external information. Furthermore, by utilizing feedback, the experience can be continuously improved.
[1003] A "user" is an individual or group who uses the system to relive past memories.
[1004] "Multimedia data" refers to digital data in various forms, such as photographs, documents, and audio recordings.
[1005] The "server" is a computer system that performs the core processing of the system, such as analyzing received data, obtaining supplementary information, generating scenarios, and creating visual and audio content.
[1006] "Metadata" refers to information that accompanies the main data, such as a photograph or audio data (e.g., the location where the photograph was taken, the time, recognized people or scenery, etc.).
[1007] An "external database" is a database located outside the system that is a source of information that provides supplementary information such as weather data and news data.
[1008] "Virtual time travel experience" refers to a virtual reality experience provided by a system that allows users to re-experience past events and environments through visual and audio experiences.
[1009] A "scenario" refers to a description or story that includes the specific sequence and details of past events that a user re-experiences.
[1010] "Generative AI model" refers to artificial intelligence technology for generating scenario details and descriptions based on prompt text.
[1011] A "prompt" is an input document that provides instructions to a generative AI model and provides a starting point for generating scenarios and explanatory text.
[1012] "Visual Content" refers to the visual elements, such as footage, images, and 3D models, provided in the virtual time travel experience.
[1013] "Audio Content" refers to the audio elements, such as music, sound effects, and dialogue, that are provided in a virtual time travel experience.
[1014] "Feedback" refers to information such as evaluations, impressions, and requests for improvement provided by users after experiencing a product.
[1015] The system of the present invention aims to provide a virtual time travel experience that allows users to vividly relive past memories. This system is mainly composed of three entities: a user, a terminal, and a server. Specific embodiments for implementing this system are described in detail below.
[1016] 1. Data Collection and Analysis
[1017] Users log in to the system via a dedicated application or website. After logging in, they upload multimedia data such as photos, documents, and audio recordings related to past memories via their device. The device automatically sends the uploaded data to the server, which stores it in secure storage. The server then uses AI algorithms to analyze the photo data, identifying scenes and people, and extracting metadata such as the location and time of the photo. Similarly, audio data is converted into text using voice recognition technology, and the content is analyzed.
[1018] 2. Acquiring and integrating external data
[1019] The server queries external databases (e.g., weather databases, news archives) based on the provided metadata to obtain complementary information. This allows for the collection of complementary information that is consistent with the user's multimedia data, creating a more realistic experience.
[1020] 3. Generation of Virtual Time Travel Scenario
[1021] The server combines the acquired metadata and complementary information to generate a virtual time travel scenario. This process utilizes a generative AI model, specifically using prompts such as:
[1022] "May 15, 2022, 3pm, on a beach in Tokyo"
[1023] Based on this prompt, the AI model generates a detailed scenario, such as "It's 3 p.m. and the beach is crowded with tourists, the temperature is 25 degrees, you can feel the sea breeze, and you can hear the laughter of family members all around."
[1024] 4. Visual and audio content generation
[1025] The server creates visual and audio content based on the generated scenario. The visual content is generated using a 3D graphics engine, and the audio content is reproduced realistically using voice synthesis technology.
[1026] 5. Providing a user experience
[1027] The device then provides the generated visual and audio content to the user through VR goggles and speakers, allowing the user to vividly and realistically re-experience past events, such as the beach scenery, conversations, and environmental sounds of a past family trip.
[1028] 6. Feedback Collection and Analysis
[1029] After the experience, the user provides feedback via their device. This feedback includes impressions of the experience and requests for improvements. The server collects and analyzes this feedback to help improve the system. By reflecting this feedback in the generation of the next experience, it becomes possible to increase user satisfaction.
[1030] In this way, the system of the present invention provides a virtual time travel experience that allows users to vividly relive past memories.
[1031] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1032] Step 1:
[1033] To log in to the system, users access a dedicated application or website. They enter their login information and complete authentication. Then, they upload multimedia data such as photos, documents, and audio recordings related to past memories. The input is various digital files selected by the user, and the output is data packets sent to the terminal.
[1034] Step 2:
[1035] The terminal sends the received multimedia data to the server. The terminal uses a data transfer protocol to send the files uploaded by the user to the server and add them to a queue for processing. The input is the file uploaded to the terminal, and the output is the data sent to the server.
[1036] Step 3:
[1037] The server stores the data in secure storage as soon as it receives it. Next, it uses AI algorithms to identify landscapes and people in the photo data and extract metadata such as the location and time of the photo. The audio data is converted to text using speech recognition technology and the content is analyzed. The input is the data sent from the device, and the output is the analyzed metadata and the converted audio data.
[1038] Step 4:
[1039] The server queries external databases based on the extracted metadata to obtain related complementary information. Specifically, it sends requests via API to weather databases, news archives, etc. to obtain the necessary information. The input is metadata (e.g., date and time of shooting, location), and the output is complementary information obtained from the external database.
[1040] Step 5:
[1041] The server combines internal data with external supplementary information and provides prompts to the generative AI model, which then generates a virtual time travel scenario. The prompts specify specific dates, times, and locations, and the AI model generates detailed scenarios based on them. The input is the prompts combined with metadata and supplementary information, and the output is the generated scenario.
[1042] For example, enter the prompt text as "May 15, 2022, 3:00 PM, on a beach in Tokyo."
[1043] Step 6:
[1044] The server creates visual and audio content based on the generated scenario. The visual content is realistically reproduced using a 3D graphics engine, and the audio content is realistically reproduced using speech synthesis technology. The input is the generated scenario, and the output is the visual and audio content provided to the user.
[1045] Step 7:
[1046] The terminal provides the user with visual and audio content sent from the server. Specifically, it uses VR goggles and speakers to make the experience feel realistic. The input is the content sent from the server, and the output is the virtual time travel experience provided to the user.
[1047] Step 8:
[1048] After completing the virtual time travel experience, the user provides feedback. The feedback includes impressions of the experience and suggestions for improvement, and is sent to the server via the terminal. The input is the user's impressions and suggestions for improvement, and the output is the feedback data sent to the server.
[1049] Step 9:
[1050] The server analyzes the received feedback and reflects it in the next experience generation, thereby improving user satisfaction. The input is the feedback data, and the output is the analysis results that will help improve the system.
[1051] (Application example 1)
[1052] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1053] Previous time travel experience systems have had limitations in the realism and detail of the visual and audio content when recreating past memories based on user-provided data. Furthermore, there has been a lack of methods to improve the consistency and accuracy of the generated content. This has resulted in a lack of realism and detailed reproduction in the virtual time travel experience that users can feel.
[1054] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1055] In this invention, the server includes a means for generating content to reproduce past memories provided by the user in more detail and more realistically, a means for creating optimal prompt sentences to be input into a generative AI model and generating content based on the prompt sentences, and a means for generating visual and audio data to reproduce a specific date and place in the past based on data provided by the user, thereby enabling the user to virtually re-experience past memories with high accuracy and realism.
[1056] A "user" is an individual who provides data such as photographs, documents, and audio recordings to use the system to recreate past memories as a virtual time travel experience.
[1057] "Photos" are image data taken by the user in the past, and serve as materials for generating visual content in the virtual time travel experience.
[1058] A "document" is text data created by a user in the past, and is used as auxiliary information in generating a scenario for virtual time travel.
[1059] "Audio recordings" are audio data recorded by users in the past, and are used to generate audio content for virtual time travel experiences.
[1060] "Metadata" is information extracted from photographs, documents, and audio recordings, and is additional information such as date, time, place, and people that is necessary to generate a scenario for a virtual time travel experience.
[1061] An "external database" is an external data source that provides additional information to complement the user's past memories, such as weather data or news articles.
[1062] The term "virtual time travel experience" refers to a user reliving past memories using virtual reality or augmented reality technology.
[1063] An "AI model" is an algorithm that uses machine learning techniques to analyze and integrate user-provided data to generate a virtual time travel experience.
[1064] A "prompt" is a text sentence input into a generative AI model that specifically instructs the scenario and visual and audio content of the virtual time travel experience.
[1065] "Visual content" means the visual representations, such as images and 3D models, that a user can view in the virtual time travel experience.
[1066] "Audio content" refers to the acoustic representation of conversations, environmental sounds, and other sounds that a user can hear during a virtual time travel experience.
[1067] "Feedback" refers to reaction data such as impressions and requests for improvement provided by a user after completing the virtual time travel experience.
[1068] "Content generation means" means the technical means for generating visual and audio content using AI models based on data provided by the user and complementary information from external databases.
[1069] The present invention is a system that analyzes past photographs, documents, and audio recordings provided by a user to generate a virtual time travel experience. This system is mainly composed of a user, a terminal, and a server, and is realized by the following processing steps.
[1070] 1. System Overview
[1071] The system consists of a data collection and analysis module, an experience generation module, and an interaction module.
[1072] 1.1 Data Collection and Analysis Module
[1073] Users upload photos, documents, and audio recordings related to past memories to the system. The device then sends the uploaded data to the server, which then analyzes it using AI algorithms. Scenes and people are recognized from the photos, and the location and time of the photo are extracted. Audio data is also converted into text using voice recognition technology.
[1074] 1.2 External Data Collection
[1075] The server retrieves relevant complementary information from external databases (e.g., weather databases, news archives), allowing the user-provided data to be integrated with external information to more realistically recreate past events.
[1076] 1.3 Experience Generation Module
[1077] The server integrates the collected metadata with external information and uses an AI model to generate a virtual time travel scenario, including a detailed description of a day in the family trip and specific events, and generates visual content (e.g., 3D models of historical landscapes and people) and audio content (e.g., conversations and environmental sounds).
[1078] 1.4 Interaction Modules
[1079] The server then sends the generated visual and audio content to the device, providing the experience to the user through VR goggles and speakers. The user can virtually relive scenes from a past family trip and hear the conversations and surrounding environmental sounds. After the experience, the user provides feedback, which the server analyzes and reflects in the generation of the next experience.
[1080] 2. Hardware and Software Used
[1081] The system uses the following hardware and software:
[1082] Hardware: Personal computer, VR goggles, smartphone, speakers
[1083] Software: Google Cloud Vision API (image analysis), Google Cloud Speech-to-Text API (voice recognition), cloud storage service
[1084] 3. Examples of concrete examples and prompts
[1085] Examples:
[1086] Users upload photos, voice messages, and social media posts from tourist spots they've visited with their family to the system. The system analyzes this data and retrieves the day's weather and newspaper articles from an external database. Based on the retrieved metadata and supplementary information, the system creates a scenario of a day at the beach and generates visual and audio content that realistically recreates beach scenes and festival scenes.
[1087] Example prompt sentence:
[1088] "Using user-provided data, recreate a family beach trip on August 15, 2005."
[1089] This system allows users to virtually re-experience past memories with high accuracy and realism.
[1090] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1091] Step 1:
[1092] Users access the system using a terminal and upload photos, documents, and audio recordings related to past memories. These data are used as input to generate a virtual time travel experience. The input for this step is the data provided by the user (photos, documents, audio recordings), and the output is this data sent to the server.
[1093] Step 2:
[1094] The device sends the data uploaded by the user to the server, which receives and securely stores this data. The input of this step is the data sent from the device, and the output is the data stored on the server.
[1095] Step 3:
[1096] The server analyzes the stored data and extracts metadata using AI algorithms. For example, it recognizes landscapes and people in photos and extracts the location and time of the photo, and converts audio data into text using voice recognition technology. The input for this step is the data stored on the server, and the output is the extracted metadata.
[1097] Step 4:
[1098] The server retrieves relevant complementary information from external databases (e.g., weather databases, news archives), providing complementary information that is consistent with the user's data. The input of this step is the extracted metadata, and the output is the retrieved complementary information.
[1099] Step 5:
[1100] The server integrates the collected metadata with external information and uses an AI model to generate a virtual time travel scenario. This scenario includes detailed descriptions of the day and specific events. The input of this step is the integrated metadata and complementary information, and the output is the generated scenario.
[1101] Step 6:
[1102] The server generates visual and audio content based on the generated scenario. The visual content includes 3D models of historical landscapes and people, and the audio content includes conversations and environmental sounds. The input of this step is the generated scenario, and the output is the visual and audio content.
[1103] Step 7:
[1104] The server transmits the generated visual and audio content to the terminal, where the user can experience virtual time travel using VR goggles and speakers. The input of this step is the generated visual and audio content, and the output is the provision of content to the user.
[1105] Step 8:
[1106] After the user finishes the virtual time travel experience, they provide feedback about the experience. The device sends this feedback to the server, which analyzes it. The input of this step is the user's feedback, and the output is the analyzed feedback data.
[1107] Step 9:
[1108] The server then refines the system based on the analyzed feedback, which then incorporates new knowledge into the next experience generation, improving the system's accuracy and user satisfaction. The input of this step is the analyzed feedback data, and the output is an improved system.
[1109] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1110] To implement the present invention, a user must provide past photographs, documents, and audio recordings to the system, which then analyzes and integrates this data to generate a virtual time travel experience. Furthermore, the present invention incorporates an emotion engine that recognizes the user's emotions, further personalizing the user's experience. The following describes in detail the modes for implementing the invention.
[1111] 1. System Overview
[1112] This system is based on users, terminals, and servers, and is composed of the following major functional modules:
[1113] 1. Data Collection and Analysis Module
[1114] 2. Experience Generation Module
[1115] 3. Interaction Module
[1116] 4. Emotion Engine
[1117] 2. Program Processing
[1118] Data Acquisition and Analysis Module
[1119] First, a user accesses the system using a terminal and logs in. The user uploads photos, documents, and audio recordings (e.g., photos from a family trip or audio of a conversation) related to past memories.
[1120] The terminal transmits the uploaded data to the server.
[1121] The server securely stores the received data and analyzes it using AI algorithms. For example, it can recognize landscapes and people from photo data and extract the location and time of the photo. It also uses voice recognition technology to convert audio data into text.
[1122] External Data Collection
[1123] The server makes requests to external APIs to retrieve relevant information from public databases, for example, weather data for a particular date and location, or newspaper articles for that day.
[1124] This allows user-provided data and external information to be integrated consistently, laying the foundation for more realistic reconstructions of past events.
[1125] Experience Generation Module
[1126] The server combines the collected metadata with external information and uses AI models to generate virtual time-travel scenarios, such as a day in the life of a family vacation or a detailed description of specific events.
[1127] Based on the generated scenario, visual content (e.g., 3D models of old landscapes and people) and audio content (e.g., conversations and environmental sounds) are generated.
[1128] Emotion Engine
[1129] The server recognizes emotions using photos, voice recordings, and text data provided by the user. The emotion engine analyzes the user's emotional state (e.g., joy, sadness, nostalgia) extracted from the data.
[1130] This allows the scenario of the virtual time travel experience to be adjusted according to the user's emotions, providing a more personalized experience.
[1131] Interaction Module
[1132] The server transmits the completed visual and audio content to the terminal.
[1133] The device provides users with an experience through VR goggles and speakers, allowing them to virtually relive scenes from a past family trip, for example, and hear the conversations and environmental sounds from that time.
[1134] Feedback collection
[1135] After the experience, users provide feedback via the device, including their impressions of the experience and requests for improvement.
[1136] The server analyzes the collected feedback and uses it to generate the next experience, improving the accuracy of the system and user satisfaction.
[1137] Specific examples
[1138] Example: Recreating family trip memories
[1139] 1. The user uploads photos, voice messages, and social media posts from tourist spots they have visited with their family to the system.
[1140] Examples: A photo taken at the beach on a summer day, or an audio recording of a family barbecue.
[1141] 2. The server analyzes this data and retrieves the day's weather (e.g., sunny with occasional clouds, temperature 30°C) and local newspaper articles (e.g., that a local festival was held) from an external database.
[1142] 3. The server creates a scenario of a day at the beach based on the retrieved metadata and complementary information. For example, spending the morning at the beach, having a family barbecue in the afternoon, and then going to a local festival.
[1143] 4. The server uses an emotion engine to read emotions from the user's uploaded data. Example: Identifying the user's emotion of "joy" from a photo of them at the beach.
[1144] 5. The server reflects the user's emotions in the generated scenario, for example, creating a scenario that emphasizes the joy of being at the beach or creating audio content that enhances the fun of having a barbecue.
[1145] 6. The server transmits the generated visual and audio content to the terminal.
[1146] 7. The device provides users with an experience through VR goggles, allowing them to experience the beach scenery right before their eyes and hear the conversations from past barbecues and the sounds of festivals in a realistic way.
[1147] 8. Users provide feedback after their experience, and the system uses that feedback to further improve the experience next time.
[1148] The present invention allows users to re-experience past memories in vivid detail, and the emotion engine provides a more personalized virtual time travel experience.
[1149] The processing flow will be explained below.
[1150] Step 1:
[1151] Users access the system using a terminal, create an account and log in. Users select past photos, documents and audio recordings and upload them to the system.
[1152] Step 2:
[1153] The device sends the data uploaded by the user to the server, including photo files, audio files, and text files.
[1154] Step 3:
[1155] The server stores the received data in a secure database, after which it begins analyzing each piece of data using AI algorithms.
[1156] Step 4:
[1157] The server analyzes the photo data and extracts metadata such as the landscape and people. For example, it uses image recognition technology to identify the location and date of the photo.
[1158] Step 5:
[1159] The server analyzes the audio data, converts the audio content into text, and uses voice recognition technology to identify the content of the conversation and who is speaking, saving it as metadata.
[1160] Step 6:
[1161] The server analyzes the text data provided by the user and applies emotion recognition algorithms to extract the user's emotional state (e.g., joy, sadness, nostalgia).
[1162] Step 7:
[1163] The server accesses external databases to gather additional information relevant to a particular date and location, such as weather data or newspaper articles for that day.
[1164] Step 8:
[1165] The server combines the collected metadata with external information to generate a scenario for a virtual time-travel experience, using an AI model to create a timeline with specific actions, such as "playing at the beach in the morning and having a family barbecue in the afternoon."
[1166] Step 9:
[1167] The server uses an emotion engine to reflect the user's emotions in the generated scenario. For example, it generates vivid and bright visuals to emphasize scenes in which the user felt "joy."
[1168] Step 10:
[1169] The server generates visual content, using a generative image model to create 3D models of historical landscapes and people, such as beach sand, ocean views, or even the user's family.
[1170] Step 11:
[1171] The server generates audio content, using a speech generation model to realistically recreate past conversations and environmental sounds (e.g., the sound of waves or family conversations).
[1172] Step 12:
[1173] The server then transmits the generated visual and audio content to the device, including 4K resolution video files and high-quality audio files.
[1174] Step 13:
[1175] The device provides the received content to the user, who then puts on the VR goggles and begins the experience using high-quality speakers.
[1176] Step 14:
[1177] Users can relive past memories through virtual experiences, such as a beach scene unfolding before their eyes, with the sound of waves and family conversations realistically recreated.
[1178] Step 15:
[1179] After the experience, users can enter their feedback, sending their impressions of the experience and requests for improvements to the system through a dedicated feedback form.
[1180] Step 16:
[1181] The terminal transmits the user's feedback to the server.
[1182] Step 17:
[1183] The server analyzes the received feedback and uses it as data to improve the system. Based on the analysis results, the accuracy of the AI model is improved and reflected in the next experience generation.
[1184] In this way, the present invention allows users to relive their past memories vividly and in detail, providing a personalized virtual time travel experience with an emotion engine.
[1185] Example 2
[1186] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1187] Conventional virtual experience systems have difficulty integrating user-provided data with external information to generate a realistic experience. Furthermore, because personalization that takes user emotions into account is not performed, it is not possible to provide an optimal experience for each individual user. Therefore, a system is needed that analyzes the diverse data provided by users and integrates it with external information to provide an individually optimized virtual experience that reflects the user's emotions.
[1188] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a means for receiving past image data, text data, and audio data from the user, a means for analyzing this data and extracting metadata, and a means for acquiring related supplementary information from an external information source. This makes it possible to provide a realistic and personalized virtual experience while taking into account the user's emotions.
[1189] A "user" is a person who uses the virtual experience system.
[1190] "Image data" refers to digital data that contains visual information such as photographs and illustrations.
[1191] "Text data" refers to digital data that contains information in the form of sentences or characters.
[1192] "Audio data" refers to digital data that contains auditory information such as human voices and environmental sounds.
[1193] "Metadata" is data that represents information related to data, and is information that serves as the basis for analysis and integration.
[1194] "External sources" are resources that provide additional information, such as databases or APIs outside the system.
[1195] A "scenario" is the story or sequence of events that make up the overall picture of a virtual experience.
[1196] "Visual content" refers to visual information such as images and videos that are presented to a user.
[1197] "Audio content" refers to auditory information such as music and sound effects that is presented to the user.
[1198] "Feedback" refers to impressions and suggestions for improvement provided by users of the virtual experience system.
[1199] "Emotion" refers to information that represents the user's psychological state, such as joy, sadness, or nostalgia.
[1200] "Personalization" is the process of optimizing the content and presentation of an experience for each individual user.
[1201] To implement this invention, the user must provide the system with previously captured image data, text data, and audio data, which is then analyzed and integrated to generate a virtual experience. In addition, the invention incorporates an emotion engine that recognizes the user's emotions, allowing for a more personalized experience.
[1202] System configuration
[1203] This system is based on users, terminals, and servers, and is composed of the following main functional modules:
[1204] 1. Data Collection and Analysis Module
[1205] 2. Experience Generation Module
[1206] 3. Interaction Module
[1207] 4. Emotion Engine
[1208] Data Acquisition and Analysis Module
[1209] First, the user accesses the system using a terminal and logs in. The user then uploads image data, text data, and audio data related to past memories.
[1210] The terminal transmits the uploaded data to the server.
[1211] The server securely stores the received data and analyzes each piece of data using AI algorithms. For example, it recognizes landscapes and people from photo data and extracts the location and time of the photo. It also uses voice recognition technology to convert audio data into text. The specific software used is an "image recognition API" for image analysis and a "voice recognition API" for audio analysis.
[1212] External Data Collection
[1213] The server sends requests to external APIs to retrieve relevant information from public databases, such as weather data for a specific date and location, or news articles for that day. Specifically, it uses the "Weather API" for weather information and the "News API" for news articles.
[1214] This allows user-provided data and external information to be integrated consistently, laying the foundation for more realistic reconstructions of past events.
[1215] Experience Generation Module
[1216] The server combines the collected metadata with external information and generates a scenario for the virtual experience using a generative AI model. For example, it creates a scenario that includes a detailed description of a day on a family trip or a specific event. The AI model used is a "generative AI model (such as GPT-4)."
[1217] Based on the generated scenario, visual content (e.g., 3D models of old landscapes and people) and audio content (e.g., conversations and environmental sounds) are generated. Visual content is generated using "visual content generation software (e.g., 3D modeling software)," and audio content is generated using "audio editing software."
[1218] Emotion Engine
[1219] The server recognizes emotions using image data, audio data, and text data provided by the user. The emotion engine analyzes the user's emotional state (e.g., joy, sadness, nostalgia) extracted from the data. An "emotion analysis API" is used for emotion analysis.
[1220] This allows the virtual experience scenario to be adjusted according to the user's emotions, providing a more personalized experience.
[1221] Interaction Module
[1222] The server transmits the completed visual and audio content to the terminal.
[1223] The device provides users with an experience through VR goggles and speakers, allowing them to virtually relive scenes from a past family trip, for example, and hear the conversations and environmental sounds from that time.
[1224] Feedback collection
[1225] After the experience, users provide feedback via their device, including their impressions of the experience and requests for improvement.
[1226] The server analyzes the collected feedback and uses it to generate the next experience, improving the accuracy of the system and user satisfaction.
[1227] Specific examples
[1228] Example: Recreating family trip memories
[1229] 1. The user uploads image data, voice messages, and social media posts of tourist spots visited with their family to the system.
[1230] Examples: a photo taken at a place you visited on a summer day, an audio recording of you having a meal with your family.
[1231] 2. The server analyzes this data and retrieves the day's weather (e.g., sunny with occasional clouds, temperature 30°C) and local news articles (e.g., local events being held) from an external database.
[1232] 3. The server creates a scenario of a day in a specific location based on the retrieved metadata and complementary information, e.g., visiting tourist spots in the morning, having a family meal in the afternoon, and then going to a local event.
[1233] 4. The server uses an emotion engine to read emotions from the user's uploaded data. For example, it identifies the user's "joy" emotion from a photo.
[1234] 5. The server reflects the user's emotions in the generated scenario. For example, it creates a scenario that emphasizes the joy of a tourist spot or audio content that enhances the enjoyment of a meal.
[1235] 6. The server transmits the generated visual and audio content to the terminal.
[1236] 7. The device provides users with an experience through VR goggles, allowing them to experience the scenery of tourist spots right before their eyes and hear the sounds of past conversations and events in a realistic way.
[1237] 8. Users provide feedback after their experience, and the system uses that feedback to further improve the experience next time.
[1238] Prompt Sentence Examples
[1239] Below is an example of a prompt sentence to input to the generative AI model.
[1240] I've uploaded photos and audio recordings from past family trips. Please use these to generate a virtual experience. I'd particularly like to emphasize the emotion of joy. Include supplemental information like the weather for that day and local events.
[1241] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1242] Step 1:
[1243] The user accesses the system using a terminal and logs in by entering their ID and password on the login page.
[1244] Input: ID, password
[1245] Output: User credentials
[1246] Specific operation: The device sends the login information to the server, and the server verifies the authentication information and allows the user to log in. If authentication is successful, the user's dashboard is displayed.
[1247] Step 2:
[1248] Users select image data, text data, and audio data related to past memories from the dashboard and press the upload button to upload this data to the system.
[1249] Input: image data, text data, audio data
[1250] Output: Upload data
[1251] Specific operation: The terminal divides the selected data into packets and sends them to the server.
[1252] Step 3:
[1253] The terminal transmits the uploaded data to the server.
[1254] Input: Upload data
[1255] Output: Transmitted data
[1256] Specific operation: The terminal converts the data into the appropriate format for each type (image, audio, text) and sends it to the server.
[1257] Step 4:
[1258] The server will securely store the received data in the specified directory.
[1259] Input: Send data
[1260] Output: Saved data
[1261] Specific operation: The server changes the file name and saves it in the appropriate folder depending on the type of data.
[1262] Step 5:
[1263] The server analyzes the data using AI algorithms such as image recognition, voice recognition, and text analysis.
[1264] Input: Saved data
[1265] Output: Metadata
[1266] Specific operation: Extracts scenery and people from photo data and converts audio data into text. Specifically, it uses "image recognition API" and "voice recognition API." Analyzed metadata includes location, time, and characters.
[1267] Step 6:
[1268] The server sends external API requests to retrieve relevant information from designated public sources.
[1269] Input: Metadata
[1270] Output: Complementary information
[1271] Specific operation: Compose a query for the information you want to obtain (weather, news articles, etc.) and access external sources. For example, use the "Weather API" for weather information and the "News API" for news articles.
[1272] Step 7:
[1273] The server integrates the collected metadata with external information and generates a virtual experience scenario using a generative AI model.
[1274] Input: Metadata, complementary information
[1275] Output: Scenario data
[1276] Specific operation: Automatically generates a storyboard based on the specified date and location based on user-provided data and external information. A generative AI model (such as GPT-4) is used as the generative AI model.
[1277] Step 8:
[1278] The server creates visual and audio content based on the generated scenario.
[1279] Input: Scenario data
[1280] Output: Visual content, audio content
[1281] What it does: Visual content is generated using visual content generation software (e.g., 3D modeling software) and audio content is generated using audio editing software, which generates the detailed elements of the virtual experience.
[1282] Step 9:
[1283] The server uses an emotion engine to analyze emotions from the image data, audio data, and text data provided by the user.
[1284] Input: image data, audio data, text data
[1285] Output: Emotion data
[1286] Specific operation: Emotion analysis uses the "Emotion Analysis API" to identify the user's psychological state. For example, emotions such as "joy" and "sadness" can be extracted from a photo.
[1287] Step 10:
[1288] Based on the analysis results, the server adjusts the experience scenario to match the user's emotions.
[1289] Input: Emotion data, scenario data
[1290] Output: personalized scenario data
[1291] Specific behavior: Adding music and effects that correspond to emotions and optimizing the content of the scenario for each individual user, providing an experience that emphasizes specific emotions.
[1292] Step 11:
[1293] The server transmits the completed visual and audio content to the terminal.
[1294] Input: Visual content, audio content
[1295] Output: Send content
[1296] Specific operation: The server acts as a file server and provides streaming or download links to devices.
[1297] Step 12:
[1298] The device provides users with an experience using VR goggles and speakers.
[1299] Input: Submit content
[1300] Output: Experience data
[1301] How it works: Users can enjoy a virtual experience through a dedicated VR app, and past experiences are realistically recreated for the user.
[1302] Step 13:
[1303] After the experience, users provide feedback via the device.
[1304] Input: Experience data
[1305] Output: Feedback data
[1306] What to do: Enter your thoughts and suggestions for improvement in the feedback form and submit it. Use the text fields and rating sliders.
[1307] Step 14:
[1308] The server analyzes the collected feedback using machine learning models and reflects it in improvements to the system.
[1309] Input: Feedback data
[1310] Output: Improvement data
[1311] Specific actions: Perform natural language processing of feedback to discover new areas for improvement. Specific software used includes "natural language processing software."
[1312] (Application example 2)
[1313] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1314] In today's world, there is a growing demand for new experiences that vividly recreate past memories. However, existing technologies simply display or play back past photographs and audio recordings, making it difficult to integrate this information and provide a deeply personalized virtual experience. Furthermore, they are unable to tailor the experience to the user's emotions, creating a need for a system that allows users to relive the emotions of past memories. The present invention aims to solve these problems and provide a virtual experience system that realistically recreates past experiences while recognizing the user's emotions.
[1315] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for receiving the user's past photos, documents, and voice recordings, analyzing this data, and extracting metadata; means for acquiring related complementary information from an external database; means for generating a scenario for a virtual time travel experience based on the analysis and complementary information; means for generating visual and audio content based on the generated scenario; and means for recognizing the user's emotions and adjusting the experience scenario and content accordingly. This enables the user to recreate realistic past experiences personalized according to their emotions.
[1316] "Means for recognizing a user's emotions and adjusting the experience scenario and content accordingly" refers to technology that uses artificial intelligence to analyze a user's emotional state from photos, documents, audio recordings, etc. provided by the user, and personalizes the content of the virtual experience based on those emotions.
[1317] "Generative AI models" are artificial intelligence algorithms and frameworks that perform natural language processing and image generation, examples of which include GPT-4 and DALL-E.
[1318] "Metadata" is data that provides information about the original data, such as the date and location when a photo was taken, and the contents of audio recordings.
[1319] "External databases" refer to data resources and databases that are publicly available on the Internet, including weather information and past newspaper articles.
[1320] A "virtual time travel experience scenario" is a scenario generated by artificial intelligence based on the user's past records and related information, and refers to a script or plan of action that allows the user to recreate past experiences in virtual reality.
[1321] "Visual and audio content" refers to the visual and audio elements that provide the user with a virtual experience, such as generated 3D models, ambient sounds, and dialogue.
[1322] "Means for collecting, analyzing, and reflecting feedback in generating the next experience" refers to a system that records the impressions and opinions provided by users after an experience, analyzes them using artificial intelligence, and improves the next experience scenario.
[1323] "Complementary information obtained from external databases" is additional information collected to complement the data provided by the user, examples of which include weather information and event information.
[1324] 1. System Overview
[1325] This system is based on the user, terminal, and server, and consists of the following main functional modules that analyze and integrate the user's past photos, documents, and audio recordings to generate and provide a virtual time travel experience.
[1326] 1.1 Data Collection and Analysis Module
[1327] The server receives photos, documents, and audio recordings uploaded by users using their devices. It analyzes this data and extracts metadata about location, time, and content. It uses Python and TensorFlow to apply AI algorithms to recognize scenes and people in photos and convert audio data into text.
[1328] 1.2 External Data Collection Module
[1329] The server retrieves relevant complementary information from external databases, such as the day's weather information or newspaper articles from public databases on the Internet, using a REST API to retrieve this external data.
[1330] 1.3 Experience Generation Module
[1331] The server combines the collected metadata with external information and generates a virtual time travel scenario using a generative AI model (e.g., GPT-4, DALL-E). Based on the generated scenario, visual and audio content such as 3D models and environmental sounds are generated. This process is performed using Unity or Unreal Engine.
[1332] 1.4 Emotion Engine
[1333] The server recognizes emotions from the data uploaded by the user. The emotion engine uses a natural language processing (NLP) library to analyze the user's emotional state and incorporate it into the scenario. For example, it identifies the emotion of joy from a photo and incorporates it into the scene.
[1334] 1.5 Interaction Module
[1335] The server then sends the completed visual and audio content to the terminal, providing the user with a virtual experience, which can be enjoyed through VR goggles or smart glasses.
[1336] 1.6 Feedback Collection Module
[1337] The server collects and analyzes feedback from users, including impressions of the experience and requests for improvement. The analyzed feedback is used to generate the next experience.
[1338] 2. Specific Examples
[1339] Example: Recreating family vacation memories
[1340] 1. The user uploads photos and voice messages of tourist spots that they have visited with their family to the system.
[1341] Example: Photos of tourist spots you visited on a summer day and a recording of the conversations you had at the time.
[1342] 2. The server analyzes this data and retrieves information about the day's weather and local events from an external database.
[1343] 3. The server generates a family trip experience scenario based on the acquired metadata and supplementary information, including, for example, scenes of play at tourist spots and events of the day.
[1344] 4. The server uses an emotion engine to read emotions from the user's uploaded data. For example, it identifies the user's emotion of "joy" from photos of tourist spots.
[1345] 5. The server reflects the user's emotions in the generated scenario, for example, creating a scenario that emphasizes joy or audio content that enhances enjoyment.
[1346] 6. The server transmits the generated visual and audio content to the terminal.
[1347] 7. The device provides users with an experience through VR goggles, allowing them to experience the scenery of tourist spots right before their eyes and hear past conversations and environmental sounds in a realistic way.
[1348] 8. Users provide feedback after their experience, and the server uses that feedback to further improve the experience next time.
[1349] 3. Examples of prompts
[1350] 1. To generate a scenario that recreates family vacation memories, the following prompt sentence is input into the generative AI model:
[1351] "Generate an experience scenario of a past family trip based on photos and audio recordings of tourist spots visited by the family. The scenario should include the flow of a day enjoyed at tourist spots visited on a summer day. In particular, the photos convey the emotion of 'joy,' so please reflect this emotion in the scenario."
[1352] As described above, the present invention realizes a system that realistically reproduces a user's past experiences and provides a personalized virtual time travel experience that responds to emotions.
[1353] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1354] Step 1:
[1355] Users log in to the system using a terminal and upload photos, documents, and audio recordings related to past memories. This data is input and sent from the terminal to the server. Specifically, users open the upload screen on their device, such as a smartphone or PC, select the target data using the file selection button, and press the upload button.
[1356] Step 2:
[1357] The server receives data sent by the user and stores it securely. The received data is the input, and storing it internally on the server is the output. Specifically, the server's file storage system stores the data and adds a metadata record to the database.
[1358] Step 3:
[1359] The server analyzes the stored data, recognizes scenes and people in the photos, and converts audio data into text. This is done using Python and TensorFlow. Photo and audio data are the input, and extracted metadata is the output. Specifically, the image recognition algorithm identifies scenes and people in the photos and generates associated tags. At the same time, the speech recognition algorithm converts audio into text.
[1360] Step 4:
[1361] The server sends a request to the public database's API to obtain related information from the external database. For example, to obtain weather information or event information for a specified date and time. The input is location and date information related to the user's input data, and the output is weather information and event information obtained based on that. Specifically, the server sends an HTTP request to the external API, parses the JSON response, and extracts the required information.
[1362] Step 5:
[1363] The server uses a generative AI model (e.g., GPT-4, DALL-E) to generate a scenario for a virtual time travel experience based on the analyzed metadata and the acquired external information. The input is the metadata and external information, and the output is a virtual experience scenario. Specifically, the server sends a prompt to the AI model to generate the scenario. This prompt includes the user's data and related information.
[1364] Step 6:
[1365] The server generates visual and audio content based on the scenario. For example, it creates 3D models and environmental sounds. The input is the generated scenario, and the output is the visual and audio content. Specifically, it uses Unity or Unreal Engine to model the visual content and the audio engine to generate the environmental sounds.
[1366] Step 7:
[1367] The server reads emotions from the user's uploaded data, analyzes the emotional state using an emotion engine, and reflects it in the generated scenario. The input is the user's photo and voice data, and the output is a scenario that reflects the emotions. Specifically, it uses an NLP library to perform emotion analysis and add emotional expressions to the scenario text.
[1368] Step 8:
[1369] The server sends the completed visual and audio content to the device and provides it to the user. The input is the generated content, and the output is the user's experience. Specifically, the content is streamed to the user's device and displayed and played to the user through the device's VR goggles or smart glasses.
[1370] Step 9:
[1371] After the experience, the user provides feedback through the device. The input is their impressions of the experience and requests for improvement, and the output is feedback data. Specifically, the user enters their comments in the feedback form on the device and presses the send button.
[1372] Step 10:
[1373] The server analyzes the collected feedback and reflects it in the generation of the next experience. The input is feedback data, and the output is an improved experience scenario. Specifically, the server analyzes the feedback data with an analysis tool, obtains new insights, and reflects them in the next prompt.
[1374] Through the above processing steps, the user can recreate a realistic past experience that is personalized according to their emotions.
[1375] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1376] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1377] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1378] [Fourth embodiment]
[1379] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1380] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1381] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1382] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1383] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1384] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1385] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1386] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1387] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1388] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1389] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1390] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1391] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1392] To implement the present invention, a user must provide historical photographs, documents, and audio recordings to the system, which then analyzes and integrates this data to generate a virtual time travel experience. The following describes in detail the mode for implementing the invention.
[1393] 1. System Overview
[1394] This system is based on users, terminals, and servers, and is composed of the following major functional modules:
[1395] 1. Data Collection and Analysis Module
[1396] 2. Experience Generation Module
[1397] 3. Interaction Module
[1398] 2. Program Processing
[1399] Data Acquisition and Analysis Module
[1400] First, users access the system using a terminal and create an account. After logging in, they upload photos, documents, and audio recordings (e.g., photos from a family trip or audio recordings of conversations) related to past memories.
[1401] The terminal transmits the uploaded data to the server.
[1402] The server securely stores the received data and analyzes it using AI algorithms. For example, it can recognize landscapes and people from photo data and extract the location and time of the photo. It also uses voice recognition technology to convert audio data into text.
[1403] External Data Collection
[1404] The server also retrieves relevant complementary information from external databases (e.g., weather databases, news archives), collecting data about the day's weather and important events.
[1405] This allows user-provided data and external information to be integrated consistently, laying the foundation for more realistic reconstructions of past events.
[1406] Experience Generation Module
[1407] The server combines the collected metadata with external information and uses AI models to generate virtual time-travel scenarios, such as a day in the life of a family vacation or a detailed description of specific events.
[1408] Based on the generated scenario, visual content (e.g., 3D models of old landscapes and people) and audio content (e.g., conversations and environmental sounds) are generated.
[1409] Interaction Module
[1410] The server transmits the completed visual and audio content to the terminal.
[1411] The device provides users with an experience through VR goggles and speakers, allowing them to virtually relive scenes from a past family trip, for example, and hear the conversations and environmental sounds from that time.
[1412] After the experience, users provide feedback via the device, including their impressions of the experience and requests for improvement.
[1413] The server analyzes the collected feedback and uses it to generate the next experience, improving the accuracy of the system and user satisfaction.
[1414] Specific examples
[1415] Example: Recreating family trip memories
[1416] 1. The user uploads photos, voice messages, and social media posts from tourist spots they have visited with their family to the system.
[1417] Examples: A photo taken at the beach on a summer day, or an audio recording of a family barbecue.
[1418] 2. The server analyzes this data and retrieves the day's weather (e.g., sunny with occasional clouds, temperature 30°C) and local newspaper articles (e.g., that a local festival was held) from an external database.
[1419] 3. The server creates a scenario of a day at the beach based on the retrieved metadata and complementary information. For example, spending the morning at the beach, having a family barbecue in the afternoon, and then going to a local festival.
[1420] 4. Based on the generated scenario, the server generates visual and audio content that realistically recreates the beach scenery and festival scenes.
[1421] 5. The device provides users with an experience through VR goggles, allowing them to experience the beach scenery right before their eyes and hear the conversations from past barbecues and the sounds of festivals in a realistic way.
[1422] 6. Users provide feedback after their experience, and the system uses that feedback to further improve the experience next time.
[1423] The present invention allows users to re-experience past memories in vivid detail, providing a meaningful virtual time travel experience for each individual user.
[1424] The processing flow will be explained below.
[1425] Step 1:
[1426] Users access the system using a terminal and log in. They select and upload past photos, documents, and audio recordings.
[1427] Step 2:
[1428] The device sends the data uploaded by the user to the server, including photo files, audio files, text files, etc.
[1429] Step 3:
[1430] The server stores the received data in a secure database and begins analyzing each piece of data, using AI algorithms to extract landscape and person metadata (e.g., date and time of photo capture, location), and convert the audio data into text.
[1431] Step 4:
[1432] The server makes requests to external APIs to retrieve relevant information from public databases, for example, weather data for a particular date and location, or newspaper articles for that day.
[1433] Step 5:
[1434] The server synthesizes the collected and analyzed data and generates scenarios for virtual time travel experiences, using AI models to create a coherent timeline based on the user's experiences. For example, it generates a scenario where the user arrives at the beach at 10:00 AM and has lunch at 1:00 PM.
[1435] Step 6:
[1436] The server generates visual content based on the scenario. It uses image generation models to create 3D models of past landscapes and people. Specifically, it realistically recreates beach sand, ocean views, and the user's family.
[1437] Step 7:
[1438] The server also generates audio content based on the scenario, using a speech generation model to realistically recreate past conversations and environmental sounds, providing an audio experience that includes, for example, the sounds of waves, wind, and conversations with family members.
[1439] Step 8:
[1440] The server then sends the finished visual and audio content to the device, which includes 4K resolution video and high-quality audio files to the user.
[1441] Step 9:
[1442] The device provides users with a virtual experience using VR goggles and high-quality speakers. Users can experience visual content through the VR goggles they wear and listen to audio content through the speakers.
[1443] Step 10:
[1444] After the experience, users can enter their feedback, such as their impressions of the experience or requests for improvements, into a dedicated form.
[1445] Step 11:
[1446] The terminal transmits the user's feedback to the server.
[1447] Step 12:
[1448] The server analyzes the received feedback and uses it as data to improve the system, improving the accuracy of the AI model and reflecting it in the next experience generation.
[1449] The above are the specific processing steps for allowing the user to vividly re-experience past memories and events.
[1450] Example 1
[1451] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1452] Previous technologies for recreating past memories lacked the accuracy and consistency of information analysis, making it difficult for users to vividly relive events that they actually experienced. Furthermore, integration with external information was insufficient, resulting in an inconsistent reproduction of past events. Furthermore, there was a lack of a way to effectively utilize user feedback to improve the experience.
[1453] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1454] In this invention, the server includes means for receiving past multimedia data from a user, means for analyzing the data and extracting metadata, means for obtaining related supplemental information from an external database, means for generating a scenario for a virtual time travel experience based on the analysis and supplemental information, means for providing prompts to a generative AI model to generate details of the scenario, means for generating visual and audio content based on the generated scenario, means for providing the generated content to the user, and means for collecting, analyzing, and reflecting user feedback in the generation of the next experience. This allows the user to vividly relive past memories and enjoy a realistic virtual time travel experience while maintaining consistency with external information. Furthermore, by utilizing feedback, the experience can be continuously improved.
[1455] A "user" is an individual or group who uses the system to relive past memories.
[1456] "Multimedia data" refers to digital data in various forms, such as photographs, documents, and audio recordings.
[1457] The "server" is a computer system that performs the core processing of the system, such as analyzing received data, obtaining supplementary information, generating scenarios, and creating visual and audio content.
[1458] "Metadata" refers to information that accompanies the main data, such as a photograph or audio data (e.g., the location where the photograph was taken, the time, recognized people or scenery, etc.).
[1459] An "external database" is a database located outside the system that is a source of information that provides supplementary information such as weather data and news data.
[1460] "Virtual time travel experience" refers to a virtual reality experience provided by a system that allows users to re-experience past events and environments through visual and audio experiences.
[1461] A "scenario" refers to a description or story that includes the specific sequence and details of past events that a user re-experiences.
[1462] "Generative AI model" refers to artificial intelligence technology for generating scenario details and descriptions based on prompt text.
[1463] A "prompt" is an input document that provides instructions to a generative AI model and provides a starting point for generating scenarios and explanatory text.
[1464] "Visual Content" refers to the visual elements, such as footage, images, and 3D models, provided in the virtual time travel experience.
[1465] "Audio Content" refers to the audio elements, such as music, sound effects, and dialogue, that are provided in a virtual time travel experience.
[1466] "Feedback" refers to information such as evaluations, impressions, and requests for improvement provided by users after experiencing a product.
[1467] The system of the present invention aims to provide a virtual time travel experience that allows users to vividly relive past memories. This system is mainly composed of three entities: a user, a terminal, and a server. Specific embodiments for implementing this system are described in detail below.
[1468] 1. Data Collection and Analysis
[1469] Users log in to the system via a dedicated application or website. After logging in, they upload multimedia data such as photos, documents, and audio recordings related to past memories via their device. The device automatically sends the uploaded data to the server, which stores it in secure storage. The server then uses AI algorithms to analyze the photo data, identifying scenes and people, and extracting metadata such as the location and time of the photo. Similarly, audio data is converted into text using voice recognition technology, and the content is analyzed.
[1470] 2. Acquiring and integrating external data
[1471] The server queries external databases (e.g., weather databases, news archives) based on the provided metadata to obtain complementary information. This allows for the collection of complementary information that is consistent with the user's multimedia data, creating a more realistic experience.
[1472] 3. Generation of Virtual Time Travel Scenario
[1473] The server combines the acquired metadata and complementary information to generate a virtual time travel scenario. This process utilizes a generative AI model, specifically using prompts such as:
[1474] "May 15, 2022, 3pm, on a beach in Tokyo"
[1475] Based on this prompt, the AI model generates a detailed scenario, such as "It's 3 p.m. and the beach is crowded with tourists, the temperature is 25 degrees, you can feel the sea breeze, and you can hear the laughter of family members all around."
[1476] 4. Visual and audio content generation
[1477] The server creates visual and audio content based on the generated scenario. The visual content is generated using a 3D graphics engine, and the audio content is reproduced realistically using voice synthesis technology.
[1478] 5. Providing a user experience
[1479] The device then provides the generated visual and audio content to the user through VR goggles and speakers, allowing the user to vividly and realistically re-experience past events, such as the beach scenery, conversations, and environmental sounds of a past family trip.
[1480] 6. Feedback Collection and Analysis
[1481] After the experience, the user provides feedback via their device. This feedback includes impressions of the experience and requests for improvements. The server collects and analyzes this feedback to help improve the system. By reflecting this feedback in the generation of the next experience, it becomes possible to increase user satisfaction.
[1482] In this way, the system of the present invention provides a virtual time travel experience that allows users to vividly relive past memories.
[1483] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1484] Step 1:
[1485] To log in to the system, users access a dedicated application or website. They enter their login information and complete authentication. Then, they upload multimedia data such as photos, documents, and audio recordings related to past memories. The input is various digital files selected by the user, and the output is data packets sent to the terminal.
[1486] Step 2:
[1487] The terminal sends the received multimedia data to the server. The terminal uses a data transfer protocol to send the files uploaded by the user to the server and add them to a queue for processing. The input is the file uploaded to the terminal, and the output is the data sent to the server.
[1488] Step 3:
[1489] The server stores the data in secure storage as soon as it receives it. Next, it uses AI algorithms to identify landscapes and people in the photo data and extract metadata such as the location and time of the photo. The audio data is converted to text using speech recognition technology and the content is analyzed. The input is the data sent from the device, and the output is the analyzed metadata and the converted audio data.
[1490] Step 4:
[1491] The server queries external databases based on the extracted metadata to obtain related complementary information. Specifically, it sends requests via API to weather databases, news archives, etc. to obtain the necessary information. The input is metadata (e.g., date and time of shooting, location), and the output is complementary information obtained from the external database.
[1492] Step 5:
[1493] The server combines internal data with external supplementary information and provides prompts to the generative AI model, which then generates a virtual time travel scenario. The prompts specify specific dates, times, and locations, and the AI model generates detailed scenarios based on them. The input is the prompts combined with metadata and supplementary information, and the output is the generated scenario.
[1494] For example, enter the prompt text as "May 15, 2022, 3:00 PM, on a beach in Tokyo."
[1495] Step 6:
[1496] The server creates visual and audio content based on the generated scenario. The visual content is realistically reproduced using a 3D graphics engine, and the audio content is realistically reproduced using speech synthesis technology. The input is the generated scenario, and the output is the visual and audio content provided to the user.
[1497] Step 7:
[1498] The terminal provides the user with visual and audio content sent from the server. Specifically, it uses VR goggles and speakers to make the experience feel realistic. The input is the content sent from the server, and the output is the virtual time travel experience provided to the user.
[1499] Step 8:
[1500] After completing the virtual time travel experience, the user provides feedback. The feedback includes impressions of the experience and suggestions for improvement, and is sent to the server via the terminal. The input is the user's impressions and suggestions for improvement, and the output is the feedback data sent to the server.
[1501] Step 9:
[1502] The server analyzes the received feedback and reflects it in the next experience generation, thereby improving user satisfaction. The input is the feedback data, and the output is the analysis results that will help improve the system.
[1503] (Application example 1)
[1504] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1505] Previous time travel experience systems have had limitations in the realism and detail of the visual and audio content when recreating past memories based on user-provided data. Furthermore, there has been a lack of methods to improve the consistency and accuracy of the generated content. This has resulted in a lack of realism and detailed reproduction in the virtual time travel experience that users can feel.
[1506] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1507] In this invention, the server includes a means for generating content to reproduce past memories provided by the user in more detail and more realistically, a means for creating optimal prompt sentences to be input into a generative AI model and generating content based on the prompt sentences, and a means for generating visual and audio data to reproduce a specific date and place in the past based on data provided by the user, thereby enabling the user to virtually re-experience past memories with high accuracy and realism.
[1508] A "user" is an individual who provides data such as photographs, documents, and audio recordings to use the system to recreate past memories as a virtual time travel experience.
[1509] "Photos" are image data taken by the user in the past, and serve as materials for generating visual content in the virtual time travel experience.
[1510] A "document" is text data created by a user in the past, and is used as auxiliary information in generating a scenario for virtual time travel.
[1511] "Audio recordings" are audio data recorded by users in the past, and are used to generate audio content for virtual time travel experiences.
[1512] "Metadata" is information extracted from photographs, documents, and audio recordings, and is additional information such as date, time, place, and people that is necessary to generate a scenario for a virtual time travel experience.
[1513] An "external database" is an external data source that provides additional information to complement the user's past memories, such as weather data or news articles.
[1514] The term "virtual time travel experience" refers to a user reliving past memories using virtual reality or augmented reality technology.
[1515] An "AI model" is an algorithm that uses machine learning techniques to analyze and integrate user-provided data to generate a virtual time travel experience.
[1516] A "prompt" is a text sentence input into a generative AI model that specifically instructs the scenario and visual and audio content of the virtual time travel experience.
[1517] "Visual content" means the visual representations, such as images and 3D models, that a user can view in the virtual time travel experience.
[1518] "Audio content" refers to the acoustic representation of conversations, environmental sounds, and other sounds that a user can hear during a virtual time travel experience.
[1519] "Feedback" refers to reaction data such as impressions and requests for improvement provided by a user after completing the virtual time travel experience.
[1520] "Content generation means" means the technical means for generating visual and audio content using AI models based on data provided by the user and complementary information from external databases.
[1521] The present invention is a system that analyzes past photographs, documents, and audio recordings provided by a user to generate a virtual time travel experience. This system is mainly composed of a user, a terminal, and a server, and is realized by the following processing steps.
[1522] 1. System Overview
[1523] The system consists of a data collection and analysis module, an experience generation module, and an interaction module.
[1524] 1.1 Data Collection and Analysis Module
[1525] Users upload photos, documents, and audio recordings related to past memories to the system. The device then sends the uploaded data to the server, which then analyzes it using AI algorithms. Scenes and people are recognized from the photos, and the location and time of the photo are extracted. Audio data is also converted into text using voice recognition technology.
[1526] 1.2 External Data Collection
[1527] The server retrieves relevant complementary information from external databases (e.g., weather databases, news archives), allowing the user-provided data to be integrated with external information to more realistically recreate past events.
[1528] 1.3 Experience Generation Module
[1529] The server integrates the collected metadata with external information and uses an AI model to generate a virtual time travel scenario, including a detailed description of a day in the family trip and specific events, and generates visual content (e.g., 3D models of historical landscapes and people) and audio content (e.g., conversations and environmental sounds).
[1530] 1.4 Interaction Modules
[1531] The server then sends the generated visual and audio content to the device, providing the experience to the user through VR goggles and speakers. The user can virtually relive scenes from a past family trip and hear the conversations and surrounding environmental sounds. After the experience, the user provides feedback, which the server analyzes and reflects in the generation of the next experience.
[1532] 2. Hardware and Software Used
[1533] The system uses the following hardware and software:
[1534] Hardware: Personal computer, VR goggles, smartphone, speakers
[1535] Software: Google Cloud Vision API (image analysis), Google Cloud Speech-to-Text API (voice recognition), cloud storage service
[1536] 3. Examples of concrete examples and prompts
[1537] Examples:
[1538] Users upload photos, voice messages, and social media posts from tourist spots they've visited with their family to the system. The system analyzes this data and retrieves the day's weather and newspaper articles from an external database. Based on the retrieved metadata and supplementary information, the system creates a scenario of a day at the beach and generates visual and audio content that realistically recreates beach scenes and festival scenes.
[1539] Example prompt sentence:
[1540] "Using user-provided data, recreate a family beach trip on August 15, 2005."
[1541] This system allows users to virtually re-experience past memories with high accuracy and realism.
[1542] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1543] Step 1:
[1544] Users access the system using a terminal and upload photos, documents, and audio recordings related to past memories. These data are used as input to generate a virtual time travel experience. The input for this step is the data provided by the user (photos, documents, audio recordings), and the output is this data sent to the server.
[1545] Step 2:
[1546] The device sends the data uploaded by the user to the server, which receives and securely stores this data. The input of this step is the data sent from the device, and the output is the data stored on the server.
[1547] Step 3:
[1548] The server analyzes the stored data and extracts metadata using AI algorithms. For example, it recognizes landscapes and people in photos and extracts the location and time of the photo, and converts audio data into text using voice recognition technology. The input for this step is the data stored on the server, and the output is the extracted metadata.
[1549] Step 4:
[1550] The server retrieves relevant complementary information from external databases (e.g., weather databases, news archives), providing complementary information that is consistent with the user's data. The input of this step is the extracted metadata, and the output is the retrieved complementary information.
[1551] Step 5:
[1552] The server integrates the collected metadata with external information and uses an AI model to generate a virtual time travel scenario. This scenario includes detailed descriptions of the day and specific events. The input of this step is the integrated metadata and complementary information, and the output is the generated scenario.
[1553] Step 6:
[1554] The server generates visual and audio content based on the generated scenario. The visual content includes 3D models of historical landscapes and people, and the audio content includes conversations and environmental sounds. The input of this step is the generated scenario, and the output is the visual and audio content.
[1555] Step 7:
[1556] The server transmits the generated visual and audio content to the terminal, where the user can experience virtual time travel using VR goggles and speakers. The input of this step is the generated visual and audio content, and the output is the provision of content to the user.
[1557] Step 8:
[1558] After the user finishes the virtual time travel experience, they provide feedback about the experience. The device sends this feedback to the server, which analyzes it. The input of this step is the user's feedback, and the output is the analyzed feedback data.
[1559] Step 9:
[1560] The server then refines the system based on the analyzed feedback, which then incorporates new knowledge into the next experience generation, improving the system's accuracy and user satisfaction. The input of this step is the analyzed feedback data, and the output is an improved system.
[1561] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1562] To implement the present invention, a user must provide past photographs, documents, and audio recordings to the system, which then analyzes and integrates this data to generate a virtual time travel experience. Furthermore, the present invention incorporates an emotion engine that recognizes the user's emotions, further personalizing the user's experience. The following describes in detail the modes for implementing the invention.
[1563] 1. System Overview
[1564] This system is based on users, terminals, and servers, and is composed of the following major functional modules:
[1565] 1. Data Collection and Analysis Module
[1566] 2. Experience Generation Module
[1567] 3. Interaction Module
[1568] 4. Emotion Engine
[1569] 2. Program Processing
[1570] Data Acquisition and Analysis Module
[1571] First, a user accesses the system using a terminal and logs in. The user uploads photos, documents, and audio recordings (e.g., photos from a family trip or audio of a conversation) related to past memories.
[1572] The terminal transmits the uploaded data to the server.
[1573] The server securely stores the received data and analyzes it using AI algorithms. For example, it can recognize landscapes and people from photo data and extract the location and time of the photo. It also uses voice recognition technology to convert audio data into text.
[1574] External Data Collection
[1575] The server makes requests to external APIs to retrieve relevant information from public databases, for example, weather data for a particular date and location, or newspaper articles for that day.
[1576] This allows user-provided data and external information to be integrated consistently, laying the foundation for more realistic reconstructions of past events.
[1577] Experience Generation Module
[1578] The server combines the collected metadata with external information and uses AI models to generate virtual time-travel scenarios, such as a day in the life of a family vacation or a detailed description of specific events.
[1579] Based on the generated scenario, visual content (e.g., 3D models of old landscapes and people) and audio content (e.g., conversations and environmental sounds) are generated.
[1580] Emotion Engine
[1581] The server recognizes emotions using photos, voice recordings, and text data provided by the user. The emotion engine analyzes the user's emotional state (e.g., joy, sadness, nostalgia) extracted from the data.
[1582] This allows the scenario of the virtual time travel experience to be adjusted according to the user's emotions, providing a more personalized experience.
[1583] Interaction Module
[1584] The server transmits the completed visual and audio content to the terminal.
[1585] The device provides users with an experience through VR goggles and speakers, allowing them to virtually relive scenes from a past family trip, for example, and hear the conversations and environmental sounds from that time.
[1586] Feedback collection
[1587] After the experience, users provide feedback via the device, including their impressions of the experience and requests for improvement.
[1588] The server analyzes the collected feedback and uses it to generate the next experience, improving the accuracy of the system and user satisfaction.
[1589] Specific examples
[1590] Example: Recreating family trip memories
[1591] 1. The user uploads photos, voice messages, and social media posts from tourist spots they have visited with their family to the system.
[1592] Examples: A photo taken at the beach on a summer day, or an audio recording of a family barbecue.
[1593] 2. The server analyzes this data and retrieves the day's weather (e.g., sunny with occasional clouds, temperature 30°C) and local newspaper articles (e.g., that a local festival was held) from an external database.
[1594] 3. The server creates a scenario of a day at the beach based on the retrieved metadata and complementary information. For example, spending the morning at the beach, having a family barbecue in the afternoon, and then going to a local festival.
[1595] 4. The server uses an emotion engine to read emotions from the user's uploaded data. Example: Identifying the user's emotion of "joy" from a photo of them at the beach.
[1596] 5. The server reflects the user's emotions in the generated scenario, for example, creating a scenario that emphasizes the joy of being at the beach or creating audio content that enhances the fun of having a barbecue.
[1597] 6. The server transmits the generated visual and audio content to the terminal.
[1598] 7. The device provides users with an experience through VR goggles, allowing them to experience the beach scenery right before their eyes and hear the conversations from past barbecues and the sounds of festivals in a realistic way.
[1599] 8. Users provide feedback after their experience, and the system uses that feedback to further improve the experience next time.
[1600] The present invention allows users to re-experience past memories in vivid detail, and the emotion engine provides a more personalized virtual time travel experience.
[1601] The processing flow will be explained below.
[1602] Step 1:
[1603] Users access the system using a terminal, create an account and log in. Users select past photos, documents and audio recordings and upload them to the system.
[1604] Step 2:
[1605] The device sends the data uploaded by the user to the server, including photo files, audio files, and text files.
[1606] Step 3:
[1607] The server stores the received data in a secure database, after which it begins analyzing each piece of data using AI algorithms.
[1608] Step 4:
[1609] The server analyzes the photo data and extracts metadata such as the landscape and people. For example, it uses image recognition technology to identify the location and date of the photo.
[1610] Step 5:
[1611] The server analyzes the audio data, converts the audio content into text, and uses voice recognition technology to identify the content of the conversation and who is speaking, saving it as metadata.
[1612] Step 6:
[1613] The server analyzes the text data provided by the user and applies emotion recognition algorithms to extract the user's emotional state (e.g., joy, sadness, nostalgia).
[1614] Step 7:
[1615] The server accesses external databases to gather additional information relevant to a particular date and location, such as weather data or newspaper articles for that day.
[1616] Step 8:
[1617] The server combines the collected metadata with external information to generate a scenario for a virtual time-travel experience, using an AI model to create a timeline with specific actions, such as "playing at the beach in the morning and having a family barbecue in the afternoon."
[1618] Step 9:
[1619] The server uses an emotion engine to reflect the user's emotions in the generated scenario. For example, it generates vivid and bright visuals to emphasize scenes in which the user felt "joy."
[1620] Step 10:
[1621] The server generates visual content, using a generative image model to create 3D models of historical landscapes and people, such as beach sand, ocean views, or even the user's family.
[1622] Step 11:
[1623] The server generates audio content, using a speech generation model to realistically recreate past conversations and environmental sounds (e.g., the sound of waves or family conversations).
[1624] Step 12:
[1625] The server then transmits the generated visual and audio content to the device, including 4K resolution video files and high-quality audio files.
[1626] Step 13:
[1627] The device provides the received content to the user, who then puts on the VR goggles and begins the experience using high-quality speakers.
[1628] Step 14:
[1629] Users can relive past memories through virtual experiences, such as a beach scene unfolding before their eyes, with the sound of waves and family conversations realistically recreated.
[1630] Step 15:
[1631] After the experience, users can enter their feedback, sending their impressions of the experience and requests for improvements to the system through a dedicated feedback form.
[1632] Step 16:
[1633] The terminal transmits the user's feedback to the server.
[1634] Step 17:
[1635] The server analyzes the received feedback and uses it as data to improve the system. Based on the analysis results, the accuracy of the AI model is improved and reflected in the next experience generation.
[1636] In this way, the present invention allows users to relive their past memories vividly and in detail, providing a personalized virtual time travel experience with an emotion engine.
[1637] Example 2
[1638] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1639] Conventional virtual experience systems have difficulty integrating user-provided data with external information to generate a realistic experience. Furthermore, because personalization that takes user emotions into account is not performed, it is not possible to provide an optimal experience for each individual user. Therefore, a system is needed that analyzes the diverse data provided by users and integrates it with external information to provide an individually optimized virtual experience that reflects the user's emotions.
[1640] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a means for receiving past image data, text data, and audio data from the user, a means for analyzing this data and extracting metadata, and a means for acquiring related supplementary information from an external information source. This makes it possible to provide a realistic and personalized virtual experience while taking into account the user's emotions.
[1641] A "user" is a person who uses the virtual experience system.
[1642] "Image data" refers to digital data that contains visual information such as photographs and illustrations.
[1643] "Text data" refers to digital data that contains information in the form of sentences or characters.
[1644] "Audio data" refers to digital data that contains auditory information such as human voices and environmental sounds.
[1645] "Metadata" is data that represents information related to data, and is information that serves as the basis for analysis and integration.
[1646] "External sources" are resources that provide additional information, such as databases or APIs outside the system.
[1647] A "scenario" is the story or sequence of events that make up the overall picture of a virtual experience.
[1648] "Visual content" refers to visual information such as images and videos that are presented to a user.
[1649] "Audio content" refers to auditory information such as music and sound effects that is presented to the user.
[1650] "Feedback" refers to impressions and suggestions for improvement provided by users of the virtual experience system.
[1651] "Emotion" refers to information that represents the user's psychological state, such as joy, sadness, or nostalgia.
[1652] "Personalization" is the process of optimizing the content and presentation of an experience for each individual user.
[1653] To implement this invention, the user must provide the system with previously captured image data, text data, and audio data, which is then analyzed and integrated to generate a virtual experience. In addition, the invention incorporates an emotion engine that recognizes the user's emotions, allowing for a more personalized experience.
[1654] System configuration
[1655] This system is based on users, terminals, and servers, and is composed of the following main functional modules:
[1656] 1. Data Collection and Analysis Module
[1657] 2. Experience Generation Module
[1658] 3. Interaction Module
[1659] 4. Emotion Engine
[1660] Data Acquisition and Analysis Module
[1661] First, the user accesses the system using a terminal and logs in. The user then uploads image data, text data, and audio data related to past memories.
[1662] The terminal transmits the uploaded data to the server.
[1663] The server securely stores the received data and analyzes each piece of data using AI algorithms. For example, it recognizes landscapes and people from photo data and extracts the location and time of the photo. It also uses voice recognition technology to convert audio data into text. The specific software used is an "image recognition API" for image analysis and a "voice recognition API" for audio analysis.
[1664] External Data Collection
[1665] The server sends requests to external APIs to retrieve relevant information from public databases, such as weather data for a specific date and location, or news articles for that day. Specifically, it uses the "Weather API" for weather information and the "News API" for news articles.
[1666] This allows user-provided data and external information to be integrated consistently, laying the foundation for more realistic reconstructions of past events.
[1667] Experience Generation Module
[1668] The server combines the collected metadata with external information and generates a scenario for the virtual experience using a generative AI model. For example, it creates a scenario that includes a detailed description of a day on a family trip or a specific event. The AI model used is a "generative AI model (such as GPT-4)."
[1669] Based on the generated scenario, visual content (e.g., 3D models of old landscapes and people) and audio content (e.g., conversations and environmental sounds) are generated. Visual content is generated using "visual content generation software (e.g., 3D modeling software)," and audio content is generated using "audio editing software."
[1670] Emotion Engine
[1671] The server recognizes emotions using image data, audio data, and text data provided by the user. The emotion engine analyzes the user's emotional state (e.g., joy, sadness, nostalgia) extracted from the data. An "emotion analysis API" is used for emotion analysis.
[1672] This allows the virtual experience scenario to be adjusted according to the user's emotions, providing a more personalized experience.
[1673] Interaction Module
[1674] The server transmits the completed visual and audio content to the terminal.
[1675] The device provides users with an experience through VR goggles and speakers, allowing them to virtually relive scenes from a past family trip, for example, and hear the conversations and environmental sounds from that time.
[1676] Feedback collection
[1677] After the experience, users provide feedback via their device, including their impressions of the experience and requests for improvement.
[1678] The server analyzes the collected feedback and uses it to generate the next experience, improving the accuracy of the system and user satisfaction.
[1679] Specific examples
[1680] Example: Recreating family trip memories
[1681] 1. The user uploads image data, voice messages, and social media posts of tourist spots visited with their family to the system.
[1682] Examples: a photo taken at a place you visited on a summer day, an audio recording of you having a meal with your family.
[1683] 2. The server analyzes this data and retrieves the day's weather (e.g., sunny with occasional clouds, temperature 30°C) and local news articles (e.g., local events being held) from an external database.
[1684] 3. The server creates a scenario of a day in a specific location based on the retrieved metadata and complementary information, e.g., visiting tourist spots in the morning, having a family meal in the afternoon, and then going to a local event.
[1685] 4. The server uses an emotion engine to read emotions from the user's uploaded data. For example, it identifies the user's "joy" emotion from a photo.
[1686] 5. The server reflects the user's emotions in the generated scenario. For example, it creates a scenario that emphasizes the joy of a tourist spot or audio content that enhances the enjoyment of a meal.
[1687] 6. The server transmits the generated visual and audio content to the terminal.
[1688] 7. The device provides users with an experience through VR goggles, allowing them to experience the scenery of tourist spots right before their eyes and hear the sounds of past conversations and events in a realistic way.
[1689] 8. Users provide feedback after their experience, and the system uses that feedback to further improve the experience next time.
[1690] Prompt Sentence Examples
[1691] Below is an example of a prompt sentence to input to the generative AI model.
[1692] I've uploaded photos and audio recordings from past family trips. Please use these to generate a virtual experience. I'd particularly like to emphasize the emotion of joy. Include supplemental information like the weather for that day and local events.
[1693] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1694] Step 1:
[1695] The user accesses the system using a terminal and logs in by entering their ID and password on the login page.
[1696] Input: ID, password
[1697] Output: User credentials
[1698] Specific operation: The device sends the login information to the server, and the server verifies the authentication information and allows the user to log in. If authentication is successful, the user's dashboard is displayed.
[1699] Step 2:
[1700] Users select image data, text data, and audio data related to past memories from the dashboard and press the upload button to upload this data to the system.
[1701] Input: image data, text data, audio data
[1702] Output: Upload data
[1703] Specific operation: The terminal divides the selected data into packets and sends them to the server.
[1704] Step 3:
[1705] The terminal transmits the uploaded data to the server.
[1706] Input: Upload data
[1707] Output: Transmitted data
[1708] Specific operation: The terminal converts the data into the appropriate format for each type (image, audio, text) and sends it to the server.
[1709] Step 4:
[1710] The server will securely store the received data in the specified directory.
[1711] Input: Send data
[1712] Output: Saved data
[1713] Specific operation: The server changes the file name and saves it in the appropriate folder depending on the type of data.
[1714] Step 5:
[1715] The server analyzes the data using AI algorithms such as image recognition, voice recognition, and text analysis.
[1716] Input: Saved data
[1717] Output: Metadata
[1718] Specific operation: Extracts scenery and people from photo data and converts audio data into text. Specifically, it uses "image recognition API" and "voice recognition API." Analyzed metadata includes location, time, and characters.
[1719] Step 6:
[1720] The server sends external API requests to retrieve relevant information from designated public sources.
[1721] Input: Metadata
[1722] Output: Complementary information
[1723] Specific operation: Compose a query for the information you want to obtain (weather, news articles, etc.) and access external sources. For example, use the "Weather API" for weather information and the "News API" for news articles.
[1724] Step 7:
[1725] The server integrates the collected metadata with external information and generates a virtual experience scenario using a generative AI model.
[1726] Input: Metadata, complementary information
[1727] Output: Scenario data
[1728] Specific operation: Automatically generates a storyboard based on the specified date and location based on user-provided data and external information. A generative AI model (such as GPT-4) is used as the generative AI model.
[1729] Step 8:
[1730] The server creates visual and audio content based on the generated scenario.
[1731] Input: Scenario data
[1732] Output: Visual content, audio content
[1733] What it does: Visual content is generated using visual content generation software (e.g., 3D modeling software) and audio content is generated using audio editing software, which generates the detailed elements of the virtual experience.
[1734] Step 9:
[1735] The server uses an emotion engine to analyze emotions from the image data, audio data, and text data provided by the user.
[1736] Input: image data, audio data, text data
[1737] Output: Emotion data
[1738] Specific operation: Emotion analysis uses the "Emotion Analysis API" to identify the user's psychological state. For example, emotions such as "joy" and "sadness" can be extracted from a photo.
[1739] Step 10:
[1740] Based on the analysis results, the server adjusts the experience scenario to match the user's emotions.
[1741] Input: Emotion data, scenario data
[1742] Output: personalized scenario data
[1743] Specific behavior: Adding music and effects that correspond to emotions and optimizing the content of the scenario for each individual user, providing an experience that emphasizes specific emotions.
[1744] Step 11:
[1745] The server transmits the completed visual and audio content to the terminal.
[1746] Input: Visual content, audio content
[1747] Output: Send content
[1748] Specific operation: The server acts as a file server and provides streaming or download links to devices.
[1749] Step 12:
[1750] The device provides users with an experience using VR goggles and speakers.
[1751] Input: Submit content
[1752] Output: Experience data
[1753] How it works: Users can enjoy a virtual experience through a dedicated VR app, and past experiences are realistically recreated for the user.
[1754] Step 13:
[1755] After the experience, users provide feedback via the device.
[1756] Input: Experience data
[1757] Output: Feedback data
[1758] What to do: Enter your thoughts and suggestions for improvement in the feedback form and submit it. Use the text fields and rating sliders.
[1759] Step 14:
[1760] The server analyzes the collected feedback using machine learning models and reflects it in improvements to the system.
[1761] Input: Feedback data
[1762] Output: Improvement data
[1763] Specific actions: Perform natural language processing of feedback to discover new areas for improvement. Specific software used includes "natural language processing software."
[1764] (Application example 2)
[1765] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1766] In today's world, there is a growing demand for new experiences that vividly recreate past memories. However, existing technologies simply display or play back past photographs and audio recordings, making it difficult to integrate this information and provide a deeply personalized virtual experience. Furthermore, they are unable to tailor the experience to the user's emotions, creating a need for a system that allows users to relive the emotions of past memories. The present invention aims to solve these problems and provide a virtual experience system that realistically recreates past experiences while recognizing the user's emotions.
[1767] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for receiving the user's past photos, documents, and voice recordings, analyzing this data, and extracting metadata; means for acquiring related complementary information from an external database; means for generating a scenario for a virtual time travel experience based on the analysis and complementary information; means for generating visual and audio content based on the generated scenario; and means for recognizing the user's emotions and adjusting the experience scenario and content accordingly. This enables the user to recreate realistic past experiences personalized according to their emotions.
[1768] "Means for recognizing a user's emotions and adjusting the experience scenario and content accordingly" refers to technology that uses artificial intelligence to analyze a user's emotional state from photos, documents, audio recordings, etc. provided by the user, and personalizes the content of the virtual experience based on those emotions.
[1769] "Generative AI models" are artificial intelligence algorithms and frameworks that perform natural language processing and image generation, examples of which include GPT-4 and DALL-E.
[1770] "Metadata" is data that provides information about the original data, such as the date and location when a photo was taken, and the contents of audio recordings.
[1771] "External databases" refer to data resources and databases that are publicly available on the Internet, including weather information and past newspaper articles.
[1772] A "virtual time travel experience scenario" is a scenario generated by artificial intelligence based on the user's past records and related information, and refers to a script or plan of action that allows the user to recreate past experiences in virtual reality.
[1773] "Visual and audio content" refers to the visual and audio elements that provide the user with a virtual experience, such as generated 3D models, ambient sounds, and dialogue.
[1774] "Means for collecting, analyzing, and reflecting feedback in generating the next experience" refers to a system that records the impressions and opinions provided by users after an experience, analyzes them using artificial intelligence, and improves the next experience scenario.
[1775] "Complementary information obtained from external databases" is additional information collected to complement the data provided by the user, examples of which include weather information and event information.
[1776] 1. System Overview
[1777] This system is based on the user, terminal, and server, and consists of the following main functional modules that analyze and integrate the user's past photos, documents, and audio recordings to generate and provide a virtual time travel experience.
[1778] 1.1 Data Collection and Analysis Module
[1779] The server receives photos, documents, and audio recordings uploaded by users using their devices. It analyzes this data and extracts metadata about location, time, and content. It uses Python and TensorFlow to apply AI algorithms to recognize scenes and people in photos and convert audio data into text.
[1780] 1.2 External Data Collection Module
[1781] The server retrieves relevant complementary information from external databases, such as the day's weather information or newspaper articles from public databases on the Internet, using a REST API to retrieve this external data.
[1782] 1.3 Experience Generation Module
[1783] The server combines the collected metadata with external information and generates a virtual time travel scenario using a generative AI model (e.g., GPT-4, DALL-E). Based on the generated scenario, visual and audio content such as 3D models and environmental sounds are generated. This process is performed using Unity or Unreal Engine.
[1784] 1.4 Emotion Engine
[1785] The server recognizes emotions from the data uploaded by the user. The emotion engine uses a natural language processing (NLP) library to analyze the user's emotional state and incorporate it into the scenario. For example, it identifies the emotion of joy from a photo and incorporates it into the scene.
[1786] 1.5 Interaction Module
[1787] The server then sends the completed visual and audio content to the terminal, providing the user with a virtual experience, which can be enjoyed through VR goggles or smart glasses.
[1788] 1.6 Feedback Collection Module
[1789] The server collects and analyzes feedback from users, including impressions of the experience and requests for improvement. The analyzed feedback is used to generate the next experience.
[1790] 2. Specific Examples
[1791] Example: Recreating family vacation memories
[1792] 1. The user uploads photos and voice messages of tourist spots that they have visited with their family to the system.
[1793] Example: Photos of tourist spots you visited on a summer day and a recording of the conversations you had at the time.
[1794] 2. The server analyzes this data and retrieves information about the day's weather and local events from an external database.
[1795] 3. The server generates a family trip experience scenario based on the acquired metadata and supplementary information, including, for example, scenes of play at tourist spots and events of the day.
[1796] 4. The server uses an emotion engine to read emotions from the user's uploaded data. For example, it identifies the user's emotion of "joy" from photos of tourist spots.
[1797] 5. The server reflects the user's emotions in the generated scenario, for example, creating a scenario that emphasizes joy or audio content that enhances enjoyment.
[1798] 6. The server transmits the generated visual and audio content to the terminal.
[1799] 7. The device provides users with an experience through VR goggles, allowing them to experience the scenery of tourist spots right before their eyes and hear past conversations and environmental sounds in a realistic way.
[1800] 8. Users provide feedback after their experience, and the server uses that feedback to further improve the experience next time.
[1801] 3. Examples of prompts
[1802] 1. To generate a scenario that recreates family vacation memories, the following prompt sentence is input into the generative AI model:
[1803] "Generate an experience scenario of a past family trip based on photos and audio recordings of tourist spots visited by the family. The scenario should include the flow of a day enjoyed at tourist spots visited on a summer day. In particular, the photos convey the emotion of 'joy,' so please reflect this emotion in the scenario."
[1804] As described above, the present invention realizes a system that realistically reproduces a user's past experiences and provides a personalized virtual time travel experience that responds to emotions.
[1805] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1806] Step 1:
[1807] Users log in to the system using a terminal and upload photos, documents, and audio recordings related to past memories. This data is input and sent from the terminal to the server. Specifically, users open the upload screen on their device, such as a smartphone or PC, select the target data using the file selection button, and press the upload button.
[1808] Step 2:
[1809] The server receives data sent by the user and stores it securely. The received data is the input, and storing it internally on the server is the output. Specifically, the server's file storage system stores the data and adds a metadata record to the database.
[1810] Step 3:
[1811] The server analyzes the stored data, recognizes scenes and people in the photos, and converts audio data into text. This is done using Python and TensorFlow. Photo and audio data are the input, and extracted metadata is the output. Specifically, the image recognition algorithm identifies scenes and people in the photos and generates associated tags. At the same time, the speech recognition algorithm converts audio into text.
[1812] Step 4:
[1813] The server sends a request to the public database's API to obtain related information from the external database. For example, to obtain weather information or event information for a specified date and time. The input is location and date information related to the user's input data, and the output is weather information and event information obtained based on that. Specifically, the server sends an HTTP request to the external API, parses the JSON response, and extracts the required information.
[1814] Step 5:
[1815] The server uses a generative AI model (e.g., GPT-4, DALL-E) to generate a scenario for a virtual time travel experience based on the analyzed metadata and the acquired external information. The input is the metadata and external information, and the output is a virtual experience scenario. Specifically, the server sends a prompt to the AI model to generate the scenario. This prompt includes the user's data and related information.
[1816] Step 6:
[1817] The server generates visual and audio content based on the scenario. For example, it creates 3D models and environmental sounds. The input is the generated scenario, and the output is the visual and audio content. Specifically, it uses Unity or Unreal Engine to model the visual content and the audio engine to generate the environmental sounds.
[1818] Step 7:
[1819] The server reads emotions from the user's uploaded data, analyzes the emotional state using an emotion engine, and reflects it in the generated scenario. The input is the user's photo and voice data, and the output is a scenario that reflects the emotions. Specifically, it uses an NLP library to perform emotion analysis and add emotional expressions to the scenario text.
[1820] Step 8:
[1821] The server sends the completed visual and audio content to the device and provides it to the user. The input is the generated content, and the output is the user's experience. Specifically, the content is streamed to the user's device and displayed and played to the user through the device's VR goggles or smart glasses.
[1822] Step 9:
[1823] After the experience, the user provides feedback through the device. The input is their impressions of the experience and requests for improvement, and the output is feedback data. Specifically, the user enters their comments in the feedback form on the device and presses the send button.
[1824] Step 10:
[1825] The server analyzes the collected feedback and reflects it in the generation of the next experience. The input is feedback data, and the output is an improved experience scenario. Specifically, the server analyzes the feedback data with an analysis tool, obtains new insights, and reflects them in the next prompt.
[1826] Through the above processing steps, the user can recreate a realistic past experience that is personalized according to their emotions.
[1827] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1828] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1829] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1830] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1831] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1832] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1833] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1834] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1835] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1836] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1837] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1838] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1839] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1840] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1841] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1842] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1843] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1844] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources...
Claims
1. means for receiving historical photographs, documents, and voice recordings from a user; a means for analyzing this data and extracting metadata; means for retrieving relevant complementary information from external databases; a means for generating a scenario for a virtual time travel experience based on the analysis and complementary information; a means for generating visual and audio content based on the generated scenario; means for providing the generated content to a user; A means of collecting and analyzing user feedback to inform the creation of the next experience; A system including:
2. The system of claim 1 further comprising: means for generating a visual image based on data received from a user to recreate a specific date and location in the past; and means for generating audio data based on audio information to recreate environmental sounds and conversations.
3. The system of claim 1 , further comprising means for adjusting the scenario of the virtual time travel experience to ensure consistency based on weather data and newspaper articles obtained from an external database.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A