system

The system addresses the challenge of reproducing personal memories by collecting and processing user data to generate high-quality video and interactive game content in a virtual reality environment, providing a realistic and cost-effective experience.

JP2026064794APending Publication Date: 2026-04-14SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-02
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

Existing systems face challenges in realistically reproducing personal memories in high-quality formats like texts, videos, and games, and integrating different types of data for a consistent experience, which is costly and difficult for ordinary users to achieve.

Method used

A system that collects user data such as text, audio, and images, performs natural language and image processing, automatically generates scenarios, and converts them into video and interactive game content within a virtual reality environment using template-based methods.

Benefits of technology

Enables users to relive and share their memories in a high-quality, realistic manner, making it cost-effective and accessible to general users.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026064794000001_ABST
    Figure 2026064794000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] Means for collecting one or more data from users, including text, audio, and images, Means for performing natural language processing and image processing to analyze collected data, A method for automatically generating scenarios based on analyzed data, A means for converting the generated scenario into video content and interactive game content, A means for providing the aforementioned video content and game content in a virtual reality environment, A system that includes this.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, the method including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In modern times, means for specifically reproducing precious personal memories and passing them on to future generations are limited. In particular, it is difficult to realistically reproduce memories in the form of texts, videos, games, etc. and share the experience with others. Also, usually, providing such an experience requires a high cost and is difficult for ordinary people to use. Furthermore, since the technology for integrating different types of data (text, voice, image, etc.) and converting them into a consistent experience is immature, it is difficult to automatically generate high-quality content. Therefore, there is a need for a system that solves these problems and can reproduce, store, and share personal memories with high quality and cost efficiency.

Means for Solving the Problems

[0005] The present invention provides a system that includes means for collecting one or more data from a user, such as text, audio, and images; means for performing natural language processing and image processing to analyze the collected data; means for automatically generating scenarios based on the analyzed data; means for converting the generated scenarios into video content and interactive game content; and means for providing the video content and game content in a virtual reality environment. This allows users to easily unify data in different formats, automatically generate high-quality content based on it, and reproduce it as a realistic experience. Furthermore, by using template-based scenario generation, this system can be made cost-effective and available to general users.

[0006] A "user" refers to an individual who provides data to collect and recreate their memories using the system.

[0007] "Collecting" refers to acquiring and centrally managing various types of data provided by users, such as text, audio, and images.

[0008] "Natural language processing" refers to the technology of analyzing collected text data and extracting semantically important keywords and phrases.

[0009] "Image analysis" refers to the technique of analyzing collected image data to extract key elements and scenes within the image.

[0010] A "scenario" refers to the structure of a story generated based on extracted data, and it is a story template that forms the basis of content such as videos and games.

[0011] "Automatic generation" refers to generating scenarios and content based on collected and analyzed data using specific templates and algorithms, without human intervention.

[0012] "Video content" refers to movies and videos that are visually recreated based on a generated scenario.

[0013] "Interactive game content" refers to game-style content that users can experience through actions and choices based on a generated scenario.

[0014] A "virtual reality environment" refers to technology that provides users with a virtual three-dimensional space, enabling a more realistic experience through devices such as headsets.

[0015] A "template" refers to a predefined format or structure used when generating a scenario, on which a specific story is constructed. [Brief explanation of the drawing]

[0016] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9]Shows an emotion map to which a plurality of emotions are mapped. [Figure 10] Shows an emotion map to which a plurality of emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Example 2 when an emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when an emotion engine is combined.

Mode for Carrying Out the Invention

[0017] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described according to the accompanying drawings.

[0018] First, the language used in the following description will be explained.

[0019] In the following embodiments, the labeled processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of a plurality of arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of a plurality of types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.

[0020] In the following embodiments, signed RAM (Random Access Memory) is a memory that temporarily stores information and is used as work memory by the processor.

[0021] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.

[0022] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0023] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0024] [First Embodiment]

[0025] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0026] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0027] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0028] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0029] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0030] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0031] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0032] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0033] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0034] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0035] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0036] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0037] The system of the present invention concretely recreates the user's memories and converts them into text, images, or interactive game formats, and is implemented as follows.

[0038] Collection of user input

[0039] First, users input data about their memories using a device. The device includes a dedicated application, allowing users to write their memories in a text input field. Users can also provide audio data by speaking using the voice recording function. Furthermore, users can upload related photos and images. This data is collected centrally and sent to the next processing step.

[0040] Data analysis and transformation

[0041] The collected data is analyzed by a server. The server analyzes the text entered by the user using natural language processing techniques to extract important keywords and phrases. Next, the server performs speech recognition techniques to transcribe the audio data and convert it into text data. The server also analyzes image data and uses image recognition techniques to extract key elements and scenes within the images.

[0042] Scenario generation

[0043] Based on the collected and analyzed data, the server automatically generates scenarios. The server employs a template-based scenario generation method, which efficiently creates stories based on user data. For example, if the memory is of a family trip, keywords such as "trip," "family," and "specific place name" are applied to a template, and a story is built based on that.

[0044] Generation of video and game content

[0045] Once a scenario is generated, the server converts it into video content and interactive game content. For video content, a storyboard is created based on the scenario, and short films or videos are generated accordingly. For game content, a game is created that includes interactive elements that users can experience through actions and choices.

[0046] Providing an experience

[0047] The generated video and game content is delivered through the device in a virtual reality (VR) environment. The device works in conjunction with a VR headset to stream the content, providing users with a realistic experience. This allows users to relive their memories in a new way.

[0048] Specific example

[0049] As a concrete example, consider a scenario where a user inputs "memories of a family trip to Hokkaido when they were 10 years old" into the system. The user uploads text containing "Hokkaido trip," an audio recording of what they said at the time, and photos taken during that trip. The server analyzes this data, extracting keywords such as "family," "trip," and "Hokkaido," and obtaining information from the photos such as "snowy landscape" and "family group photo." Based on this, the server generates a scenario themed around "a family trip to Hokkaido," and then creates a short film and an interactive game based on that scenario.

[0050] Users can wear a VR headset and watch movies featuring snowy landscapes of Hokkaido with their families, or play games where they collect items during their travels. This system allows users to relive their cherished memories in high quality and realism, and to save and pass them on to future generations.

[0051] The following describes the processing flow.

[0052] Step 1:

[0053] Users launch the application using their devices and input data related to their memories. Specifically, they write episodes in text fields, record audio data using the voice recording function, and upload photos and images from that time. This data is collected centrally.

[0054] Step 2:

[0055] The device sends the collected data to the server. The server tokenizes the received text data using natural language processing techniques and extracts important keywords and phrases. For example, keywords such as "family," "travel," and "Hokkaido" are extracted.

[0056] Step 3:

[0057] The server analyzes the audio data and transcribes it using speech recognition technology. The recorded audio of someone talking about their trip to Hokkaido is converted into text data. This text data is also analyzed using the aforementioned natural language processing technology.

[0058] Step 4:

[0059] The server analyzes image data and uses image recognition technology to extract key elements and scenes. For example, information such as "snowy landscape" or "family photo" can be obtained from a photograph. This allows for a clear understanding of the specific content of the image.

[0060] Step 5:

[0061] The server generates scenarios based on the collected and analyzed data. Templates are used in this process, for example, to generate a scenario such as "I traveled to Hokkaido with my family and enjoyed the snowy scenery."

[0062] Step 6:

[0063] The server uses the generated scenario to create video content. Specifically, it creates storyboards and then renders them as short films or videos. The rendered video is saved on the server.

[0064] Step 7:

[0065] The server simultaneously creates interactive game content based on the generated scenario. This game includes activities such as the user collecting items while traveling. The generated game content is also saved on the server.

[0066] Step 8:

[0067] The user wears a VR headset using a device. The device plays video content streamed from a server, providing the user with a realistic visual experience. By watching short films, the user can relive their memories with a strong sense of presence.

[0068] Step 9:

[0069] The device runs game content in a VR environment, providing users with an interactive experience. Users can explore a virtual Hokkaido and enjoy a fun time as if they were traveling with their family.

[0070] Through the steps described above, the system of the present invention enables users to transform their memories into a high-quality, realistic form, and to save and share them.

[0071] (Example 1)

[0072] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0073] Existing systems designed to recreate users' memories and deliver them in text, video, and interactive game formats face several challenges in terms of efficiency and quality. For example, incomplete data collection or inaccurate data analysis can lead to a low degree of memory recreation. Furthermore, a lack of quality in the generated content and the realism of the user experience are also problematic. Therefore, there is a need for a system that allows users to relive their memories realistically and in a high-quality manner.

[0074] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0075] In this invention, the server includes means for collecting one or more data from a user, such as text, voice, and images; means for performing natural language processing and image processing to analyze the collected data; means for automatically generating scenarios based on the analyzed data; means for converting the generated scenarios into video content and interactive game content; means for providing the video content and game content in a virtual reality environment; means for centrally managing and transmitting the collected data to the server; means for analyzing the data using natural language processing technology, speech recognition technology, and image recognition technology; and means for creating storyboards using a template-based method and generating video content and game content based on them. This enables users to relive their memories in high quality and realism, save them, and pass them on to future generations.

[0076] A "user" is an individual or group that uses this system to recreate their own memories.

[0077] A "terminal" is a device used by a user to input data and access a system, and includes smartphones, personal computers, and other similar devices.

[0078] "Data" refers to information in various forms, including text, audio, images, and other information entered by the user.

[0079] "Means of collection" refers to a device or software that has the function of centrally collecting data entered by users and transmitting it to a server.

[0080] "Means of analysis" refers to the process of analyzing collected data using natural language processing, speech recognition, and image recognition technologies.

[0081] "Natural language processing" is a technology that analyzes text data and extracts important keywords and phrases.

[0082] "Image analysis" is a technique that analyzes image data to extract key elements and scenes.

[0083] "Methods for automatically generating scenarios" refer to the process of automatically creating the structure of a story or narrative using a template-based method based on analyzed data.

[0084] "Video content" refers to visual media such as movies and videos that are produced based on a generated scenario.

[0085] "Interactive game content" refers to a type of game where users perform actions and make choices, recreating memories through the experience.

[0086] A "virtual reality environment" is a technology that allows users to experience a virtually recreated environment by wearing devices such as VR headsets.

[0087] "A means of centralized management" refers to the function of a system that unifies and manages collected data and efficiently transmits it to a server.

[0088] "Natural language processing technology" is a technique for analyzing text data and extracting information based on that analysis.

[0089] "Speech recognition technology" is a technology that converts speech data into text data and then analyzes it.

[0090] "Image recognition technology" is a technique that analyzes image data and extracts key elements or scenes from it.

[0091] A "template-based approach" is a method that efficiently generates scenarios by applying data based on a predefined template.

[0092] A "storyboard" is a visual representation of the scene structure of a video or story.

[0093] The system of this invention concretely recreates the user's memories and provides them in the form of text, images, or interactive games. Its detailed configuration and processing are described below.

[0094] Hardware and software configuration

[0095] User input

[0096] The user launches a dedicated application using their device (smartphone or personal computer). The application includes a text input field, voice recording function, and image upload function.

[0097] Data collection

[0098] The device centrally collects text data entered by the user, voice data, and uploaded image data, and sends this data to the server.

[0099] Data analysis and transformation

[0100] Natural Language Processing and Speech Recognition

[0101] The server analyzes the collected text data using natural language processing technology (e.g., Google® Cloud Natural Language API) to extract important keywords and phrases. Additionally, the collected audio data is transcribed using speech recognition technology (e.g., Google Speech-to-Text) and further analyzed.

[0102] Image Recognition

[0103] The server analyzes the image data using image recognition technology (e.g., Google Cloud Vision API) to extract key elements and scenes.

[0104] Scenario generation

[0105] The server automatically generates scenarios using a template-based methodology based on the analyzed data. These scenarios enable the efficient creation of stories based on user data.

[0106] Specific example

[0107] For example, if a user enters "memories of a family trip to Hokkaido when they were 10 years old" and uploads photos from the trip and audio recordings of them talking about their memories, the server will extract keywords such as "family," "trip," and "Hokkaido," as well as information such as "snowy scenery" and "family group photo." Then, a scenario themed around "a family trip to Hokkaido" will be generated.

[0108] Generation of video and interactive game content

[0109] Video content

[0110] The server creates storyboards based on the generated scenarios, and then uses these to produce short films or videos (for example, using Adobe Premiere Pro).

[0111] Interactive game content

[0112] The server develops an interactive game that users can experience through actions and choices (for example, using Unity).

[0113] Providing the final experience

[0114] Provided in a VR environment

[0115] The generated video and interactive game content is delivered through the device in a virtual reality (VR) environment. Users wear a VR headset (e.g., Oculus Rift) and experience the content streamed from the device.

[0116] Specific examples and prompt statements

[0117] The user inputs "memories of a family trip to Hokkaido when they were 10 years old" into the system, providing audio and images. The server analyzes the collected data, generates scenarios, and creates video and game content based on them. For example, the following prompt sentences are input into the generating AI model:

[0118] Example of a prompt:

[0119] "Generate memorable scenarios based on the places and events the user has experienced during their travels."

[0120] "Create a Hokkaido travel scenario using the following keywords: family, travel, snowscape."

[0121] In this way, users can relive their precious memories in high quality and with realism, and save them to pass on to future generations.

[0122] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0123] Step 1:

[0124] The user launches a dedicated application using their device.

[0125] Specifically, the user downloads the application and logs in.

[0126] Input: User account information

[0127] Output: Application launched, user authentication successful

[0128] Step 2:

[0129] The user enters data related to their memories.

[0130] Specifically, the user writes their memories in a text input field, records their memories verbally using the voice recording function, and uploads related photos and images.

[0131] Input: Text data, audio data, image data

[0132] Output: Collected memory data

[0133] Step 3:

[0134] The device collects and centrally manages user input data (text, voice, images). Next, the device sends the collected data to a server.

[0135] Specifically, the data is organized within the device and then uploaded to the server via the network.

[0136] Input: User-submitted memory data

[0137] Output: Dataset stored on the server

[0138] Step 4:

[0139] The server analyzes the collected text data using natural language processing techniques.

[0140] In terms of specific operations, the server calls a natural language processing API to extract important keywords and phrases from the text data.

[0141] Input: Text data

[0142] Output: Extracted keywords and phrases

[0143] Step 5:

[0144] The server transcribes the audio data using speech recognition technology and converts it into text data. Next, this text data is analyzed using natural language processing technology.

[0145] Specifically, the server utilizes a speech recognition API to convert speech into text, which is then automatically analyzed.

[0146] Input: Audio data

[0147] Output: Analyzed transcript text data

[0148] Step 6:

[0149] The server analyzes the image data using image recognition technology.

[0150] In terms of specific operations, the server calls an image recognition API to extract key elements and scenes from the image data.

[0151] Input: Image data

[0152] Output: Extracted key elements and scenes

[0153] Step 7:

[0154] The server automatically generates scenarios based on the analyzed data.

[0155] Specifically, a scenario generation algorithm is applied to generate stories using a template-based approach.

[0156] Input: Information from extracted text, audio, and images

[0157] Output: Generated scenario

[0158] Step 8:

[0159] The server creates video content based on the generated scenario.

[0160] Specifically, this involves creating storyboards according to a scenario and then producing short films or videos based on those storyboards. Video editing tools are used for this purpose.

[0161] Input: Generated scenario

[0162] Output: Short film or video content

[0163] Step 9:

[0164] The server develops interactive games based on generated scenarios.

[0165] Specifically, this involves using a game development platform to create interactive games that users can experience through actions and choices.

[0166] Input: Generated scenario

[0167] Output: Interactive game content

[0168] Step 10:

[0169] The device provides generated video and game content in a VR environment.

[0170] Specifically, the device works in conjunction with a VR headset to stream content and deliver it to the user.

[0171] Input: Video content, game content

[0172] Output: User experience in a VR environment

[0173] Through these steps, this system enables users to relive, save, and pass on their precious memories to future generations in high quality and with realism.

[0174] (Application Example 1)

[0175] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0176] Traditionally, there have been limited means to concretely recreate users' memories, making it difficult to preserve individual memories as narratives or images. Furthermore, while methods existed to collect and analyze user memories, there was no way to recreate them in a physical customer experience environment. As a result, users lacked opportunities to easily re-experience their memories in a high-quality format.

[0177] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0178] In this invention, the server includes means for collecting one or more data from a user, such as text, voice, and images; means for performing natural language processing and image processing to analyze the collected data; means for automatically generating scenarios based on the analyzed data; means for converting the generated scenarios into video content and interactive game content; and means for providing the video content and game content in a physical customer experience environment. This enables users to re-experience their cherished memories in high quality in a physical environment such as a retail store.

[0179] A "user" is an individual or group that uses the system to provide data about memories and receives the recreated content of those memories.

[0180] "Text" refers to data entered by the user, representing the user's memories and information as written text.

[0181] "Voice" refers to data provided by users through speech, which can be converted into text data using speech recognition technology, representing the user's memories and information.

[0182] "Images" refer to photographs and drawings uploaded by users, and are data that includes visual records of memories.

[0183] "Natural language processing" is a technology that analyzes text and audio data collected from users to extract important keywords and phrases.

[0184] "Image analysis" is a technology that analyzes image data collected from users to extract key elements and scenes within the image.

[0185] A "scenario" is a story or narrative generated based on collected and analyzed data, and serves as the basis for video content and game content.

[0186] A "generative AI model" is a type of artificial intelligence that automatically generates stories based on user data, and in particular, it is a model that uses natural language generation technology.

[0187] "Video content" refers to digital content in the form of videos or movies created based on a generated scenario.

[0188] "Interactive game content" refers to digital games based on generated scenarios that users can experience through actions and choices.

[0189] A "physical customer experience environment" refers to a physical store or other physical space where users experience reproduced content.

[0190] This invention is a system that concretely recreates users' memories and provides that recreated content in a physical customer experience environment. The details are described below.

[0191] Collection of user input

[0192] First, users use a device to input data about their memories. The device includes a dedicated application, allowing users to write their memories in a text input field. Users can also provide audio data by speaking using the voice recording function. Furthermore, users can upload related photos and images. This data is collected centrally and sent to a server.

[0193] Data analysis

[0194] The server analyzes text data entered by users using natural language processing techniques to extract important keywords and phrases. Specifically, the text data is analyzed using natural language processing. The server also performs speech recognition techniques to transcribe audio data and convert it into text data. Furthermore, the server analyzes image data and uses image recognition techniques to extract key elements and scenes within the images.

[0195] Scenario generation

[0196] Based on the collected and analyzed data, the server automatically generates scenarios. This uses a generative AI model. The server employs a template-based scenario generation method, for example, by applying keywords such as "travel," "family," and "specific place names" to templates and building stories based on them.

[0197] Generation of video and game content

[0198] After the scenario is generated, the server converts it into video content and interactive game content. For video content, short films and videos are generated based on the scenario. For interactive game content, a game is created that includes elements that users can experience through actions and choices.

[0199] Providing an experience

[0200] The generated video and game content will be delivered in a physical customer experience environment accessible to users. For example, users can experience the content on the spot using tablet devices or robots in a physical store.

[0201] Specific example

[0202] As a concrete example, consider a scenario where a user inputs "memories of a family trip to Hokkaido" into the system. The user uploads text such as "Hokkaido trip," an audio recording of what they said at the time, and photos taken during that trip. The server analyzes this data, extracting keywords such as "family," "trip," and "Hokkaido," and obtaining information from the photos such as "snowy landscape" and "family group photo." Based on this, the server generates a scenario themed around "a family trip to Hokkaido," and then creates a short film and an interactive game based on that scenario.

[0203] Users can use terminals in physical stores to watch movies featuring family members enjoying the snowy landscapes of Hokkaido, and participate in interactive games where they can collect items during their trip.

[0204] Example of a prompt

[0205] For example, when a user is recounting a memory, the following might be used as a prompt:

[0206] "I remember a trip I took to Hokkaido with my family. We enjoyed a warm hot spring bath while admiring the beautiful snowy scenery."

[0207] Based on this prompt, the generative AI model can generate a scenario and provide the user with high-quality content that recreates their memories.

[0208] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0209] Step 1:

[0210] Users use their devices to input one or more data related to their memories: text, audio, images, and so on. This includes typing into text fields, recording audio using a microphone, and uploading photos and images. The entered data is sent to the server.

[0211] Input: Text, audio, images

[0212] Output: Data to send to the server

[0213] Step 2:

[0214] The server analyzes the received text data using natural language processing techniques to extract important keywords and phrases. Specifically, it tokenizes the text data, performs morphological analysis, and extracts meaningful words and phrases.

[0215] Input: Text data

[0216] Output: Extracted keywords and phrases

[0217] Step 3:

[0218] The server performs speech recognition technology to transcribe the audio data and convert it into text data. The converted text data is similarly analyzed using natural language processing technology to extract keywords and phrases.

[0219] Input: Audio data

[0220] Output: Text data, extracted keywords and phrases

[0221] Step 4:

[0222] The server analyzes the received image data and uses image recognition techniques to extract key elements and scenes from the image. Image analysis employs methods such as object detection and scene classification.

[0223] Input: Image data

[0224] Output: Extracted key elements and scenes

[0225] Step 5:

[0226] The server automatically generates scenarios using a generative AI model based on extracted keywords, phrases, key elements, and scenes. The generative AI model generates a story based on the input prompt sentences.

[0227] Input: Keywords, phrases, key elements or scenes

[0228] Output: Generated scenario

[0229] Step 6:

[0230] The server converts the generated scenarios into video content and interactive game content. For video content, short films and videos are generated based on the scenarios, while interactive games include elements that users can experience through actions and choices.

[0231] Input: Generated scenario

[0232] Output: Video content, interactive game content

[0233] Step 7:

[0234] The generated video and game content will be provided to users to experience on the spot using terminals and robots within physical stores.

[0235] Input: Video content, interactive game content

[0236] Output: Providing experiences in a physical customer experience environment

[0237] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0238] The system of the present invention concretely recreates the user's memories and converts them into text, images, and interactive game formats. Furthermore, by combining this with an emotion engine that recognizes the user's emotions, it provides a more immersive experience. The following are specific embodiments of this system.

[0239] Collection of user input

[0240] First, the user launches the application using their device and enters data about their memories. Users can write episodes in a text field. They can also provide audio data by speaking about their memories using the voice recording function. Furthermore, they can upload photos and images from that time. This data is collected centrally and sent to the next processing step.

[0241] Data analysis and transformation

[0242] The collected data is analyzed by a server. The server analyzes the text entered by the user using natural language processing technology to extract important keywords and phrases. Next, the server analyzes the audio data and transcribes it using speech recognition technology. The server also analyzes the image data and extracts the main elements and scenes within the images.

[0243] Recognition of emotions

[0244] Next, the server uses an emotion engine to recognize the user's emotions based on the collected audio and text data. From the audio data, it analyzes voice tone, pitch, and volume, and from the text data, it analyzes context and word choice to detect what emotions the user is feeling.

[0245] Scenario generation

[0246] Based on the collected and analyzed data and the output of the emotion engine, the server automatically generates scenarios. Specifically, it employs a template-based scenario generation method to create stories that reflect emotional elements. For example, if a user expresses the emotion of "happiness," a scenario reflecting that emotion will be created.

[0247] Generation of video and game content

[0248] Based on the generated scenario, the server produces video content and interactive game content. For video content, storyboards are created based on the scenario and then rendered as short films or videos. Character expressions and attitudes are also adjusted based on the output of the emotion engine. In game content, interactive elements are created in which the scenario changes according to the user's emotions.

[0249] Providing an experience

[0250] The generated video and game content is delivered through the device in a virtual reality (VR) environment. The device works in conjunction with a VR headset to stream the content, providing users with a realistic experience. This allows users to relive their memories, including emotional elements.

[0251] Specific example

[0252] As a concrete example, consider a scenario where a user inputs "memories of a family trip to Hokkaido when they were 10 years old" into the system. The user uploads text containing "trip to Hokkaido," an audio recording of their conversation at the time, and photos taken during that trip. The server analyzes this data, extracting keywords such as "family," "trip," and "Hokkaido," and obtaining information from the photos such as "snowy landscape" and "family group photo." Furthermore, using an emotion engine, the system detects whether the user was feeling "happy" while recounting this memory.

[0253] Based on this, the server generates a scenario themed around a "family trip to Hokkaido," reflecting the emotion of "happiness." For example, the scenario will include many scenes where the whole family is smiling and enjoying themselves. A short film and an interactive game are then created based on this scenario.

[0254] The user wears a VR headset and watches a short film of a family enjoying a snowball fight in a park. In the game, the user can also experience a task of collecting items with their family, all while smiling. This system allows users to experience their cherished memories in a high-quality, realistic form, including the emotions they evoke.

[0255] The following describes the processing flow.

[0256] Step 1:

[0257] Users launch the application using their devices and input data related to their memories. Specifically, they write episodes in text fields, record audio data using the voice recording function, and upload related photos and images. This data is centrally collected on the device and sent to the next processing step.

[0258] Step 2:

[0259] The device sends the collected data to the server. The server receives the text data, tokenizes it using natural language processing techniques, and extracts important keywords and phrases. For example, keywords such as "family," "travel," and "Hokkaido" are extracted.

[0260] Step 3:

[0261] The server analyzes the received audio data. It uses speech recognition technology to convert the audio into text, and then analyzes the resulting text using natural language processing technology to extract important information.

[0262] Step 4:

[0263] The server analyzes the image data. Using image recognition technology, it extracts key elements and scenes from the image. For example, information such as "snowy landscape" or "family photo" can be obtained from a photograph.

[0264] Step 5:

[0265] The server uses an emotion engine to recognize the user's emotions based on collected audio and text data. From the audio data, it analyzes voice tone, pitch, and volume, and from the text data, it analyzes context and word choice to detect what emotions the user is feeling.

[0266] Step 6:

[0267] The server generates scenarios based on the analyzed data and the output of the emotion engine. Using a template-based scenario generation method, it constructs a story that reflects the emotions expressed by the user. For example, if the user expresses the emotion of "happiness," a scenario reflecting that emotion will be created.

[0268] Step 7:

[0269] The server produces video content using the generated scenario. It creates storyboards based on the scenario and then renders them as short films or videos. Character expressions and attitudes are also adjusted based on the output of the emotion engine.

[0270] Step 8:

[0271] The server creates interactive game content based on the generated scenario. The game is designed to include interactive elements where the scenario changes in response to the user's emotions.

[0272] Step 9:

[0273] The user wears a VR headset using a device. The device plays video content streamed from a server, providing the user with a realistic visual experience. By watching short films, the user can relive their memories with a strong sense of presence.

[0274] Step 10:

[0275] The device runs game content in a VR environment, providing users with an interactive experience. Users can explore a virtual Hokkaido and enjoy a fun time as if they were traveling with their family.

[0276] Through the steps described above, the system of the present invention enables users to transform their memories into a high-quality, realistic form, and to save and share them. Furthermore, the introduction of an emotion engine allows for the provision of an immersive experience that reflects the user's emotions.

[0277] (Example 2)

[0278] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0279] Modern users have a need to relive their memories in a richer way, but this is difficult to achieve with current technology. Existing systems are particularly insufficient when a concrete recreation, including the emotions associated with those memories, is required. Furthermore, there is the challenge of effectively analyzing user input data and converting it into interactive content that reflects those emotions.

[0280] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for collecting one or more data from the user, such as text, voice, and images; means for performing natural language processing and image processing to analyze the collected data; means for recognizing the user's emotions based on the collected voice and text data; means for automatically generating a scenario based on the analyzed data; means for converting the generated scenario into video content and interactive game content; and means for providing the video content and game content in a virtual reality environment. This makes it possible for the user to re-experience their memories in a high-quality, realistic way, including their emotions.

[0281] A "user" is the individual who uses this system to input and re-experience memories.

[0282] "Text" refers to the written data entered by the user.

[0283] "Voice" refers to the data of sounds input by the user speaking.

[0284] "Image" refers to visual data such as photos or pictures taken by the user.

[0285] "Data collection means" refers to the means for collecting any one or a plurality of data of text, voice, and image from the user.

[0286] "Natural language analysis" refers to the technology of analyzing text data to extract meaningful information or keywords.

[0287] "Image analysis" refers to the technology of analyzing image data to extract main elements or scenes.

[0288] "Speech recognition technology" refers to the technology for converting voice data into text data.

[0289] "Emotion recognition means" refers to the means for recognizing the user's emotion based on the collected voice data and text data.

[0290] "Scenario automatic generation means" refers to the means for automatically generating a scenario based on the analyzed data.

[0291] "Template" refers to the fixed format or pattern used in scenario automatic generation.

[0292] "Video content conversion means" refers to the means for converting the generated scenario into video content.

[0293] "Interactive game content conversion means" refers to the means for converting the generated scenario into interactive game content.

[0294] "Virtual reality environment provision means" refers to means for providing the aforementioned video content and game content in a virtual reality environment.

[0295] This invention relates to a system that concretely recreates a user's memories and converts them into text, video, or interactive game formats. Furthermore, this system incorporates an emotion engine that recognizes the user's emotions to provide a more immersive experience.

[0296] First, the user launches the application using their device. The user utilizes the device's built-in text input, voice recording, and image upload functions to input data related to their memories. Specifically, the user can write episodes in the text field, provide audio data using the voice recording function, and upload photos or images from that time. This data is then sent from the device to the server.

[0297] Next, the server analyzes the collected data. For text data, it uses Python's natural language processing libraries (such as NLTK and spaCy) to analyze it and extract important keywords and phrases. For audio data, it uses speech recognition technology (such as the Google Speech-to-Text API) to transcribe it. For image data, it uses image analysis libraries (such as OpenCV) to extract key elements and scenes.

[0298] Based on these analysis results, the server uses an emotion engine to recognize the user's emotions. The server analyzes voice tone, pitch, and volume from audio data, and context and word choice from text data to detect what emotions the user is experiencing. This emotion recognition uses IBM Watson® Tone Analyzer and Microsoft® Azure® Text Analytics API.

[0299] Based on the analyzed data and the results of emotion recognition, the server automatically generates scenarios using a template-based scenario generation method. For example, if the emotion engine detects that the user is feeling "happy," a scenario reflecting that emotion will be generated.

[0300] Based on the generated scenario, the server produces video content and interactive game content. The video content is rendered as short films or videos using video production tools (such as Adobe Premiere Pro and Final Cut Pro). Character expressions and attitudes are also adjusted based on the output of the emotion engine. The game content is created using a game engine (such as Unity or Unreal Engine) to produce a game that includes interactive elements where the scenario changes in response to emotions.

[0301] The final generated video and game content is delivered through the device in a virtual reality (VR) environment. The device works in conjunction with VR headsets (such as Oculus Rift and HTC Vive) to provide users with a realistic experience. This allows users to re-experience their memories, including emotional elements.

[0302] To give a concrete example, if a user inputs "memories of a family trip to Hokkaido when they were 10 years old" into the system, the user would type "trip to Hokkaido" into their terminal and upload audio recordings and photos from that time. The server analyzes this data, extracting keywords such as "family," "trip," and "Hokkaido," and obtaining information from the photos such as "snowy landscape" and "family group photo." Using an emotion engine, it detects the "happy" emotion the user felt while recounting this memory. Based on this, the server generates a scenario themed around "a family trip to Hokkaido" that reflects the "happy" emotion. A short film and an interactive game are created based on the scenario, and the user can experience them through a VR headset.

[0303] Examples of prompt sentences to be input into the generative AI model are as follows:

[0304] "Please write about the episode that made you happy about your memories of Hokkaido when you were 10 years old and traveled with your family. Also, please add photos and audio. Based on this data, generate scenarios with an emotion reflection system and produce video content and games."

[0305] The above is the form for implementing the invention. With this system, users can convert their precious memories into a high-quality and realistic form including emotions and re-experience them.

[0306] The flow of the specific process in Example 2 will be described using FIG. 13.

[0307] Step 1: Collection of user input

[0308] The user starts the application using the terminal and inputs data related to memories. The specific operations are as follows. The user writes an episode in the text field (input: text data), provides voice data using the voice recording function (input: voice data), and uploads photos or images (input: image data). These data are uniformly transmitted from the terminal to the server (output: transmission of integrated data).

[0309] Step 2: Analysis of data

[0310] The server receives the transmitted data and begins analysis. First, it uses Python's natural language processing libraries (e.g., NLTK, spaCy) on the text data to extract important keywords and phrases (input: text data, output: important keywords). Next, it uses speech recognition technology such as the Google Speech-to-Text API on the audio data to transcribe it (input: audio data, output: text data). Finally, it uses image analysis libraries such as OpenCV on the image data to extract key elements and scenes (input: image data, output: important elements).

[0311] Step 3: Recognizing Emotions

[0312] The server recognizes emotions based on the analyzed data. For audio data, it analyzes voice tone, pitch, and volume (input: audio data) to identify the user's emotions (output: emotion data). For text data, it uses IBM Watson Tone Analyzer or Microsoft Azure Text Analytics API to analyze context and word choice to identify emotions (input: text data, output: emotion data).

[0313] Step 4: Scenario Generation

[0314] The server automatically generates template-based scenarios based on the analysis data and emotion recognition results. Specifically, it uses a scenario generation algorithm to create stories that reflect emotional elements (input: analysis data and emotion data, output: scenario). For example, if the user indicates the emotion of "happiness," a scenario reflecting that emotion will be automatically generated.

[0315] Step 5: Generating video and game content

[0316] The server produces video content and interactive game content based on the generated scenario. For video content, it uses video production tools (e.g., Adobe Premiere Pro, Final Cut Pro) to render short films and videos based on the scenario (input: scenario, output: video content). For game content, it uses a game engine (e.g., Unity, Unreal Engine) to create interactive games that change in response to emotions (input: scenario, output: game content).

[0317] Step 6: Providing the experience

[0318] The device delivers generated video and game content in a virtual reality (VR) environment. Users wear a VR headset (e.g., Oculus Rift, HTC Vive) and experience this content (input: video and game content, output: VR experience). This allows users to re-experience their memories in a high-quality, realistic form, including the emotions they evoke.

[0319] (Application Example 2)

[0320] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0321] Traditionally, systems designed to recreate users' memories have simply displayed collected data, making it difficult to reproduce users' emotions and subtle nuances. Furthermore, it has been challenging to deliver individual user memories in real time and provide an immersive experience. Therefore, there is a need for a system that can achieve a higher level of realism and individual personalization.

[0322] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for collecting one or more data from the user, such as text, voice, and images; means for performing natural language processing and image processing to analyze the collected data; means for automatically generating a scenario based on the analyzed data; means for converting the generated scenario into video content and interactive game content; and means for streaming the scenario-based content in real time via a terminal such as a smartphone. This makes it possible to reproduce the user's emotions and memories with a high sense of realism and to provide a real-time, individually personalized experience.

[0323] A "user" refers to an individual who uses the system to input and analyze their own memories.

[0324] "Text" refers to the written data entered by the user.

[0325] "Audio" refers to audio data that records what the user has said.

[0326] "Images" refer to visual data such as photos and drawings uploaded by users.

[0327] "Natural language processing" refers to the process of analyzing collected text and audio data to extract keywords and context.

[0328] "Image analysis" refers to the process of analyzing collected image data and extracting key elements and scenes from the image.

[0329] "Automatic scenario generation" refers to a method of automatically creating scenarios that reflect user emotions using templates, based on analyzed data.

[0330] "Video content" refers to visual content rendered based on a generated scenario.

[0331] "Interactive game content" refers to game-style content that progresses through interaction with the user based on a generated scenario.

[0332] A "virtual reality environment" refers to technology that allows users to experience things in a virtual space as if they were reality.

[0333] "Smartphones and other devices" refers to portable information devices used by users to access memory-recreation streaming services.

[0334] "Streaming distribution" refers to the technology that provides generated video and game content to users in real time via a network.

[0335] "Sentiment analysis" refers to the process of analyzing the tone, pitch, and other characteristics of text and audio data collected from users to estimate their emotions.

[0336] The system of this invention reproduces user-inputted memories in real time and delivers them as personalized content that takes emotions into account. This system analyzes data collected from users, generates scenarios, and provides video content and interactive game content based on those scenarios.

[0337] Hardware and software to be used

[0338] 1. Smartphone: The primary device used by users to input text, voice, and image data and receive the results.

[0339] 2. Server / Cloud: A centralized processing system for data analysis, sentiment recognition, scenario generation, and content rendering.

[0340] 3. Natural Language Processing (NLP) libraries (e.g., SpaCy, NLTK): These libraries analyze the collected text data and extract important keywords and phrases.

[0341] 4. Speech recognition library (e.g., Google Cloud Speech-to-Text): Converts audio data into text data and transcribes the audio into text.

[0342] 5. Emotion recognition engine (e.g., IBM Watson Tone Analyzer): Analyzes the emotions in text and audio data.

[0343] 6. Video rendering engine (e.g., Unity): Generates video content and interactive game content based on the scenario.

[0344] Data processing and data calculation workflow

[0345] 1. Collecting user input:

[0346] Users input text, voice, and images using their smartphones. This data is sent to a server and collected centrally.

[0347] 2. Data Analysis:

[0348] The server uses natural language processing technology to analyze text data and extract important keywords and phrases. It also uses speech recognition technology to transcribe audio data and image analysis technology to analyze key elements of image data.

[0349] 3. Recognition of emotions:

[0350] The server uses an emotion recognition engine to analyze the user's emotions from text and voice data. This identifies the tone, pitch, volume, and context of the emotion.

[0351] 4. Scenario generation:

[0352] The server generates emotionally reflective scenarios based on the analyzed data and emotion recognition results. A template-based scenario generation method is employed to create personalized stories.

[0353] 5. Content generation:

[0354] The server creates video content and interactive game content based on the generated scenario. The generated content reflects emotional elements, and the characters' expressions and attitudes are adjusted accordingly.

[0355] 6. Providing experiences:

[0356] Ultimately, the generated video content and interactive game content are streamed in real time via devices such as smartphones. By integrating VR headsets and other devices, a more immersive experience can be provided.

[0357] Specific example

[0358] For example, if a user inputs "memories of a family trip to Hokkaido when they were 10 years old" in text, audio, or image format, the server analyzes this and extracts keywords such as "family," "trip," and "Hokkaido." In addition, it uses an emotion recognition engine to detect that the user is feeling "happy." Based on this information, the server generates a scenario themed around "a family trip to Hokkaido" that reflects the user's emotions.

[0359] Example of a prompt

[0360] Theme: Family trip to Hokkaido

[0361] Emotion: Happy

[0362] Keywords: family, travel, snowscape

[0363] This system allows users to recreate their memories, including the emotions they evoke, and experience them with a high degree of realism.

[0364] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0365] Step 1:

[0366] The user launches the application using their smartphone and inputs one or more text, audio, or image data related to their memories. For audio input, they use the voice recording function and upload photos or images taken at that time. This data is sent from the device to the server. The input data includes detailed memories (e.g., "Memories of a family trip to Hokkaido when I was 10 years old").

[0367] Step 2:

[0368] The server analyzes the collected text data using natural language processing techniques (e.g., SpaCy, NLTK) to extract important keywords and phrases. Specifically, it extracts keywords such as "family," "travel," and "Hokkaido," and understands the context of the text data. The input is text data, and the output is keywords and contextual information.

[0369] Step 3:

[0370] The server uses speech recognition technology (e.g., Google Cloud Speech-to-Text) to convert audio data into text data and perform transcription. Specifically, it converts what the user says into text format and performs similar natural language processing. The input is audio data, and the output is transcribed text data.

[0371] Step 4:

[0372] The server uses image analysis technology to analyze uploaded image data and extract key elements and scenes. Specifically, it obtains information such as "snowy landscape" or "family group photo" from the image. The input is image data, and the output is information about key elements and scenes.

[0373] Step 5:

[0374] The server analyzes the user's emotions using an emotion recognition engine (e.g., IBM Watson Tone Analyzer) based on the parsed text and audio data. Specifically, it analyzes the context of the text, the tone of the voice, and the pitch to detect whether the user is experiencing emotions such as "happy." The input is the parsed text and audio data, and the output is emotion data.

[0375] Step 6:

[0376] The server automatically generates a story using a template-based scenario generation method, based on extracted keywords, key elements, scene information, and emotion data. Specifically, it creates a scenario themed around a "family trip to Hokkaido" that reflects the emotion of "happiness." The input consists of keywords, key elements, scene information, and emotion data, and the output is the generated scenario.

[0377] Step 7:

[0378] The server produces video content and interactive game content based on the generated scenario. Specifically, it uses a video rendering engine such as Unity to create short films and interactive games while adjusting character expressions and attitudes based on emotional data. The input is the generated scenario, and the output is the video content and game content.

[0379] Step 8:

[0380] The system allows users to stream generated video content and interactive game content in real time using devices such as smartphones. Specifically, it works in conjunction with VR headsets to provide users with a high level of immersion. The input is video content and game content, and the output is the user's experience.

[0381] Through these steps, users can transform their memories into a high-quality, realistic form that includes emotions, and experience them.

[0382] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0383] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0384] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0385] [Second Embodiment]

[0386] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0387] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0388] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0389] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0390] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0391] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0392] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0393] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0394] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0395] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0396] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0397] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0398] The system of the present invention concretely recreates the user's memories and converts them into text, images, or interactive game formats, and is implemented as follows.

[0399] Collection of user input

[0400] First, users input data about their memories using a device. The device includes a dedicated application, allowing users to write their memories in a text input field. Users can also provide audio data by speaking using the voice recording function. Furthermore, users can upload related photos and images. This data is collected centrally and sent to the next processing step.

[0401] Data analysis and transformation

[0402] The collected data is analyzed by a server. The server analyzes the text entered by the user using natural language processing techniques to extract important keywords and phrases. Next, the server performs speech recognition techniques to transcribe the audio data and convert it into text data. The server also analyzes image data and uses image recognition techniques to extract key elements and scenes within the images.

[0403] Scenario generation

[0404] Based on the collected and analyzed data, the server automatically generates scenarios. The server employs a template-based scenario generation method, which efficiently creates stories based on user data. For example, if the memory is of a family trip, keywords such as "trip," "family," and "specific place name" are applied to a template, and a story is built based on that.

[0405] Generation of video and game content

[0406] Once a scenario is generated, the server converts it into video content and interactive game content. For video content, a storyboard is created based on the scenario, and short films or videos are generated accordingly. For game content, a game is created that includes interactive elements that users can experience through actions and choices.

[0407] Providing an experience

[0408] The generated video and game content is delivered through the device in a virtual reality (VR) environment. The device works in conjunction with a VR headset to stream the content, providing users with a realistic experience. This allows users to relive their memories in a new way.

[0409] Specific example

[0410] As a concrete example, consider a scenario where a user inputs "memories of a family trip to Hokkaido when they were 10 years old" into the system. The user uploads text containing "Hokkaido trip," an audio recording of what they said at the time, and photos taken during that trip. The server analyzes this data, extracting keywords such as "family," "trip," and "Hokkaido," and obtaining information from the photos such as "snowy landscape" and "family group photo." Based on this, the server generates a scenario themed around "a family trip to Hokkaido," and then creates a short film and an interactive game based on that scenario.

[0411] Users can wear a VR headset and watch movies featuring snowy landscapes of Hokkaido with their families, or play games where they collect items during their travels. This system allows users to relive their cherished memories in high quality and realism, and to save and pass them on to future generations.

[0412] The following describes the processing flow.

[0413] Step 1:

[0414] Users launch the application using their devices and input data related to their memories. Specifically, they write episodes in text fields, record audio data using the voice recording function, and upload photos and images from that time. This data is collected centrally.

[0415] Step 2:

[0416] The device sends the collected data to the server. The server tokenizes the received text data using natural language processing techniques and extracts important keywords and phrases. For example, keywords such as "family," "travel," and "Hokkaido" are extracted.

[0417] Step 3:

[0418] The server analyzes the audio data and transcribes it using speech recognition technology. The recorded audio of someone talking about their trip to Hokkaido is converted into text data. This text data is also analyzed using the aforementioned natural language processing technology.

[0419] Step 4:

[0420] The server analyzes image data and uses image recognition technology to extract key elements and scenes. For example, information such as "snowy landscape" or "family photo" can be obtained from a photograph. This allows for a clear understanding of the specific content of the image.

[0421] Step 5:

[0422] The server generates scenarios based on the collected and analyzed data. Templates are used in this process, for example, to generate a scenario such as "I traveled to Hokkaido with my family and enjoyed the snowy scenery."

[0423] Step 6:

[0424] The server uses the generated scenario to create video content. Specifically, it creates storyboards and then renders them as short films or videos. The rendered video is saved on the server.

[0425] Step 7:

[0426] The server simultaneously creates interactive game content based on the generated scenario. This game includes activities such as the user collecting items while traveling. The generated game content is also saved on the server.

[0427] Step 8:

[0428] The user wears a VR headset using a device. The device plays video content streamed from a server, providing the user with a realistic visual experience. By watching short films, the user can relive their memories with a strong sense of presence.

[0429] Step 9:

[0430] The device runs game content in a VR environment, providing users with an interactive experience. Users can explore a virtual Hokkaido and enjoy a fun time as if they were traveling with their family.

[0431] Through the steps described above, the system of the present invention enables users to transform their memories into a high-quality, realistic form, and to save and share them.

[0432] (Example 1)

[0433] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0434] Existing systems designed to recreate users' memories and deliver them in text, video, and interactive game formats face several challenges in terms of efficiency and quality. For example, incomplete data collection or inaccurate data analysis can lead to a low degree of memory recreation. Furthermore, a lack of quality in the generated content and the realism of the user experience are also problematic. Therefore, there is a need for a system that allows users to relive their memories realistically and in a high-quality manner.

[0435] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0436] In this invention, the server includes means for collecting one or more data from a user, such as text, voice, and images; means for performing natural language processing and image processing to analyze the collected data; means for automatically generating scenarios based on the analyzed data; means for converting the generated scenarios into video content and interactive game content; means for providing the video content and game content in a virtual reality environment; means for centrally managing and transmitting the collected data to the server; means for analyzing the data using natural language processing technology, speech recognition technology, and image recognition technology; and means for creating storyboards using a template-based method and generating video content and game content based on them. This enables users to relive their memories in high quality and realism, save them, and pass them on to future generations.

[0437] A "user" is an individual or group that uses this system to recreate their own memories.

[0438] A "terminal" is a device used by a user to input data and access a system, and includes smartphones, personal computers, and other similar devices.

[0439] "Data" refers to information in various forms, including text, audio, images, and other information entered by the user.

[0440] "Means of collection" refers to a device or software that has the function of centrally collecting data entered by users and transmitting it to a server.

[0441] "Means of analysis" refers to the process of analyzing collected data using natural language processing, speech recognition, and image recognition technologies.

[0442] "Natural language processing" is a technology that analyzes text data and extracts important keywords and phrases.

[0443] "Image analysis" is a technique that analyzes image data to extract key elements and scenes.

[0444] "Methods for automatically generating scenarios" refer to the process of automatically creating the structure of a story or narrative using a template-based method based on analyzed data.

[0445] "Video content" refers to visual media such as movies and videos that are produced based on a generated scenario.

[0446] "Interactive game content" refers to a type of game where users perform actions and make choices, recreating memories through the experience.

[0447] A "virtual reality environment" is a technology that allows users to experience a virtually recreated environment by wearing devices such as VR headsets.

[0448] "A means of centralized management" refers to the function of a system that unifies and manages collected data and efficiently transmits it to a server.

[0449] "Natural language processing technology" is a technique for analyzing text data and extracting information based on that analysis.

[0450] "Speech recognition technology" is a technology that converts speech data into text data and then analyzes it.

[0451] "Image recognition technology" is a technique that analyzes image data and extracts key elements or scenes from it.

[0452] A "template-based approach" is a method that efficiently generates scenarios by applying data based on a predefined template.

[0453] A "storyboard" is a visual representation of the scene structure of a video or story.

[0454] The system of this invention concretely recreates the user's memories and provides them in the form of text, images, or interactive games. Its detailed configuration and processing are described below.

[0455] Hardware and software configuration

[0456] User input

[0457] The user launches a dedicated application using their device (smartphone or personal computer). The application includes a text input field, voice recording function, and image upload function.

[0458] Data collection

[0459] The device centrally collects text data entered by the user, voice data, and uploaded image data, and sends this data to the server.

[0460] Data analysis and transformation

[0461] Natural Language Processing and Speech Recognition

[0462] The server analyzes the collected text data using natural language processing technology (e.g., Google Cloud Natural Language API) to extract important keywords and phrases. Additionally, the collected audio data is transcribed using speech recognition technology (e.g., Google Speech-to-Text) and further analyzed.

[0463] Image Recognition

[0464] The server analyzes the image data using image recognition technology (e.g., Google Cloud Vision API) to extract key elements and scenes.

[0465] Scenario generation

[0466] The server automatically generates scenarios using a template-based methodology based on the analyzed data. These scenarios enable the efficient creation of stories based on user data.

[0467] Specific example

[0468] For example, if a user enters "memories of a family trip to Hokkaido when they were 10 years old" and uploads photos from the trip and audio recordings of them talking about their memories, the server will extract keywords such as "family," "trip," and "Hokkaido," as well as information such as "snowy scenery" and "family group photo." Then, a scenario themed around "a family trip to Hokkaido" will be generated.

[0469] Generation of video and interactive game content

[0470] Video content

[0471] The server creates storyboards based on the generated scenarios, and then uses these to produce short films or videos (for example, using Adobe Premiere Pro).

[0472] Interactive game content

[0473] The server develops an interactive game that users can experience through actions and choices (for example, using Unity).

[0474] Providing the final experience

[0475] Provided in a VR environment

[0476] The generated video and interactive game content is delivered through the device in a virtual reality (VR) environment. Users wear a VR headset (e.g., Oculus Rift) and experience the content streamed from the device.

[0477] Specific examples and prompt statements

[0478] The user inputs "memories of a family trip to Hokkaido when they were 10 years old" into the system, providing audio and images. The server analyzes the collected data, generates scenarios, and creates video and game content based on them. For example, the following prompt sentences are input into the generating AI model:

[0479] Example of a prompt:

[0480] "Generate memorable scenarios based on the places and events the user has experienced during their travels."

[0481] "Create a Hokkaido travel scenario using the following keywords: family, travel, snowscape."

[0482] In this way, users can relive their precious memories in high quality and with realism, and save them to pass on to future generations.

[0483] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0484] Step 1:

[0485] The user launches a dedicated application using their device.

[0486] Specifically, the user downloads the application and logs in.

[0487] Input: User account information

[0488] Output: Application launched, user authentication successful

[0489] Step 2:

[0490] The user enters data related to their memories.

[0491] Specifically, the user writes their memories in a text input field, records their memories verbally using the voice recording function, and uploads related photos and images.

[0492] Input: Text data, audio data, image data

[0493] Output: Collected memory data

[0494] Step 3:

[0495] The device collects and centrally manages user input data (text, voice, images). Next, the device sends the collected data to a server.

[0496] Specifically, the data is organized within the device and then uploaded to the server via the network.

[0497] Input: User-submitted memory data

[0498] Output: Dataset stored on the server

[0499] Step 4:

[0500] The server analyzes the collected text data using natural language processing techniques.

[0501] In terms of specific operations, the server calls a natural language processing API to extract important keywords and phrases from the text data.

[0502] Input: Text data

[0503] Output: Extracted keywords and phrases

[0504] Step 5:

[0505] The server transcribes the audio data using speech recognition technology and converts it into text data. Next, this text data is analyzed using natural language processing technology.

[0506] Specifically, the server utilizes a speech recognition API to convert speech into text, which is then automatically analyzed.

[0507] Input: Audio data

[0508] Output: Analyzed transcript text data

[0509] Step 6:

[0510] The server analyzes the image data using image recognition technology.

[0511] In terms of specific operations, the server calls an image recognition API to extract key elements and scenes from the image data.

[0512] Input: Image data

[0513] Output: Extracted key elements and scenes

[0514] Step 7:

[0515] The server automatically generates scenarios based on the analyzed data.

[0516] Specifically, a scenario generation algorithm is applied to generate stories using a template-based approach.

[0517] Input: Information from extracted text, audio, and images

[0518] Output: Generated scenario

[0519] Step 8:

[0520] The server creates video content based on the generated scenario.

[0521] Specifically, this involves creating storyboards according to a scenario and then producing short films or videos based on those storyboards. Video editing tools are used for this purpose.

[0522] Input: Generated scenario

[0523] Output: Short film or video content

[0524] Step 9:

[0525] The server develops interactive games based on generated scenarios.

[0526] Specifically, this involves using a game development platform to create interactive games that users can experience through actions and choices.

[0527] Input: Generated scenario

[0528] Output: Interactive game content

[0529] Step 10:

[0530] The device provides generated video and game content in a VR environment.

[0531] Specifically, the device works in conjunction with a VR headset to stream content and deliver it to the user.

[0532] Input: Video content, game content

[0533] Output: User experience in a VR environment

[0534] Through these steps, this system enables users to relive, save, and pass on their precious memories to future generations in high quality and with realism.

[0535] (Application Example 1)

[0536] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0537] Traditionally, there have been limited means to concretely recreate users' memories, making it difficult to preserve individual memories as narratives or images. Furthermore, while methods existed to collect and analyze user memories, there was no way to recreate them in a physical customer experience environment. As a result, users lacked opportunities to easily re-experience their memories in a high-quality format.

[0538] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0539] In this invention, the server includes means for collecting one or more data from a user, such as text, voice, and images; means for performing natural language processing and image processing to analyze the collected data; means for automatically generating scenarios based on the analyzed data; means for converting the generated scenarios into video content and interactive game content; and means for providing the video content and game content in a physical customer experience environment. This enables users to re-experience their cherished memories in high quality in a physical environment such as a retail store.

[0540] A "user" is an individual or group that uses the system to provide data about memories and receives the recreated content of those memories.

[0541] "Text" refers to data entered by the user, representing the user's memories and information as written text.

[0542] "Voice" refers to data provided by users through speech, which can be converted into text data using speech recognition technology, representing the user's memories and information.

[0543] "Images" refer to photographs and drawings uploaded by users, and are data that includes visual records of memories.

[0544] "Natural language processing" is a technology that analyzes text and audio data collected from users to extract important keywords and phrases.

[0545] "Image analysis" is a technology that analyzes image data collected from users to extract key elements and scenes within the image.

[0546] A "scenario" is a story or narrative generated based on collected and analyzed data, and serves as the basis for video content and game content.

[0547] A "generative AI model" is a type of artificial intelligence that automatically generates stories based on user data, and in particular, it is a model that uses natural language generation technology.

[0548] "Video content" refers to digital content in the form of videos or movies created based on a generated scenario.

[0549] "Interactive game content" refers to digital games based on generated scenarios that users can experience through actions and choices.

[0550] A "physical customer experience environment" refers to a physical store or other physical space where users experience reproduced content.

[0551] This invention is a system that concretely recreates users' memories and provides that recreated content in a physical customer experience environment. The details are described below.

[0552] Collection of user input

[0553] First, users use a device to input data about their memories. The device includes a dedicated application, allowing users to write their memories in a text input field. Users can also provide audio data by speaking using the voice recording function. Furthermore, users can upload related photos and images. This data is collected centrally and sent to a server.

[0554] Data analysis

[0555] The server analyzes text data entered by users using natural language processing techniques to extract important keywords and phrases. Specifically, the text data is analyzed using natural language processing. The server also performs speech recognition techniques to transcribe audio data and convert it into text data. Furthermore, the server analyzes image data and uses image recognition techniques to extract key elements and scenes within the images.

[0556] Scenario generation

[0557] Based on the collected and analyzed data, the server automatically generates scenarios. This uses a generative AI model. The server employs a template-based scenario generation method, for example, by applying keywords such as "travel," "family," and "specific place names" to templates and building stories based on them.

[0558] Generation of video and game content

[0559] After the scenario is generated, the server converts it into video content and interactive game content. For video content, short films and videos are generated based on the scenario. For interactive game content, a game is created that includes elements that users can experience through actions and choices.

[0560] Providing an experience

[0561] The generated video and game content will be delivered in a physical customer experience environment accessible to users. For example, users can experience the content on the spot using tablet devices or robots in a physical store.

[0562] Specific example

[0563] As a concrete example, consider a scenario where a user inputs "memories of a family trip to Hokkaido" into the system. The user uploads text such as "Hokkaido trip," an audio recording of what they said at the time, and photos taken during that trip. The server analyzes this data, extracting keywords such as "family," "trip," and "Hokkaido," and obtaining information from the photos such as "snowy landscape" and "family group photo." Based on this, the server generates a scenario themed around "a family trip to Hokkaido," and then creates a short film and an interactive game based on that scenario.

[0564] Users can use terminals in physical stores to watch movies featuring family members enjoying the snowy landscapes of Hokkaido, and participate in interactive games where they can collect items during their trip.

[0565] Example of a prompt

[0566] For example, when a user is recounting a memory, the following might be used as a prompt:

[0567] "I remember a trip I took to Hokkaido with my family. We enjoyed a warm hot spring bath while admiring the beautiful snowy scenery."

[0568] Based on this prompt, the generative AI model can generate a scenario and provide the user with high-quality content that recreates their memories.

[0569] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0570] Step 1:

[0571] Users use their devices to input one or more data related to their memories: text, audio, images, and so on. This includes typing into text fields, recording audio using a microphone, and uploading photos and images. The entered data is sent to the server.

[0572] Input: Text, audio, images

[0573] Output: Data to send to the server

[0574] Step 2:

[0575] The server analyzes the received text data using natural language processing techniques to extract important keywords and phrases. Specifically, it tokenizes the text data, performs morphological analysis, and extracts meaningful words and phrases.

[0576] Input: Text data

[0577] Output: Extracted keywords and phrases

[0578] Step 3:

[0579] The server performs speech recognition technology to transcribe the audio data and convert it into text data. The converted text data is similarly analyzed using natural language processing technology to extract keywords and phrases.

[0580] Input: Audio data

[0581] Output: Text data, extracted keywords and phrases

[0582] Step 4:

[0583] The server analyzes the received image data and uses image recognition techniques to extract key elements and scenes from the image. Image analysis employs methods such as object detection and scene classification.

[0584] Input: Image data

[0585] Output: Extracted key elements and scenes

[0586] Step 5:

[0587] The server automatically generates scenarios using a generative AI model based on extracted keywords, phrases, key elements, and scenes. The generative AI model generates a story based on the input prompt sentences.

[0588] Input: Keywords, phrases, key elements or scenes

[0589] Output: Generated scenario

[0590] Step 6:

[0591] The server converts the generated scenarios into video content and interactive game content. For video content, short films and videos are generated based on the scenarios, while interactive games include elements that users can experience through actions and choices.

[0592] Input: Generated scenario

[0593] Output: Video content, interactive game content

[0594] Step 7:

[0595] The generated video and game content will be provided to users to experience on the spot using terminals and robots within physical stores.

[0596] Input: Video content, interactive game content

[0597] Output: Providing experiences in a physical customer experience environment

[0598] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0599] The system of the present invention concretely recreates the user's memories and converts them into text, images, and interactive game formats. Furthermore, by combining this with an emotion engine that recognizes the user's emotions, it provides a more immersive experience. The following are specific embodiments of this system.

[0600] Collection of user input

[0601] First, the user launches the application using their device and enters data about their memories. Users can write episodes in a text field. They can also provide audio data by speaking about their memories using the voice recording function. Furthermore, they can upload photos and images from that time. This data is collected centrally and sent to the next processing step.

[0602] Data analysis and transformation

[0603] The collected data is analyzed by a server. The server analyzes the text entered by the user using natural language processing technology to extract important keywords and phrases. Next, the server analyzes the audio data and transcribes it using speech recognition technology. The server also analyzes the image data and extracts the main elements and scenes within the images.

[0604] Recognition of emotions

[0605] Next, the server uses an emotion engine to recognize the user's emotions based on the collected audio and text data. From the audio data, it analyzes voice tone, pitch, and volume, and from the text data, it analyzes context and word choice to detect what emotions the user is feeling.

[0606] Scenario generation

[0607] Based on the collected and analyzed data and the output of the emotion engine, the server automatically generates scenarios. Specifically, it employs a template-based scenario generation method to create stories that reflect emotional elements. For example, if a user expresses the emotion of "happiness," a scenario reflecting that emotion will be created.

[0608] Generation of video and game content

[0609] Based on the generated scenario, the server produces video content and interactive game content. For video content, storyboards are created based on the scenario and then rendered as short films or videos. Character expressions and attitudes are also adjusted based on the output of the emotion engine. In game content, interactive elements are created in which the scenario changes according to the user's emotions.

[0610] Providing an experience

[0611] The generated video and game content is delivered through the device in a virtual reality (VR) environment. The device works in conjunction with a VR headset to stream the content, providing users with a realistic experience. This allows users to relive their memories, including emotional elements.

[0612] Specific example

[0613] As a concrete example, consider a scenario where a user inputs "memories of a family trip to Hokkaido when they were 10 years old" into the system. The user uploads text containing "trip to Hokkaido," an audio recording of their conversation at the time, and photos taken during that trip. The server analyzes this data, extracting keywords such as "family," "trip," and "Hokkaido," and obtaining information from the photos such as "snowy landscape" and "family group photo." Furthermore, using an emotion engine, the system detects whether the user was feeling "happy" while recounting this memory.

[0614] Based on this, the server generates a scenario themed around a "family trip to Hokkaido," reflecting the emotion of "happiness." For example, the scenario will include many scenes where the whole family is smiling and enjoying themselves. A short film and an interactive game are then created based on this scenario.

[0615] The user wears a VR headset and watches a short film of a family enjoying a snowball fight in a park. In the game, the user can also experience a task of collecting items with their family, all while smiling. This system allows users to experience their cherished memories in a high-quality, realistic form, including the emotions they evoke.

[0616] The following describes the processing flow.

[0617] Step 1:

[0618] Users launch the application using their devices and input data related to their memories. Specifically, they write episodes in text fields, record audio data using the voice recording function, and upload related photos and images. This data is centrally collected on the device and sent to the next processing step.

[0619] Step 2:

[0620] The device sends the collected data to the server. The server receives the text data, tokenizes it using natural language processing techniques, and extracts important keywords and phrases. For example, keywords such as "family," "travel," and "Hokkaido" are extracted.

[0621] Step 3:

[0622] The server analyzes the received audio data. It uses speech recognition technology to convert the audio into text, and then analyzes the resulting text using natural language processing technology to extract important information.

[0623] Step 4:

[0624] The server analyzes the image data. Using image recognition technology, it extracts key elements and scenes from the image. For example, information such as "snowy landscape" or "family photo" can be obtained from a photograph.

[0625] Step 5:

[0626] The server uses an emotion engine to recognize the user's emotions based on collected audio and text data. From the audio data, it analyzes voice tone, pitch, and volume, and from the text data, it analyzes context and word choice to detect what emotions the user is feeling.

[0627] Step 6:

[0628] The server generates scenarios based on the analyzed data and the output of the emotion engine. Using a template-based scenario generation method, it constructs a story that reflects the emotions expressed by the user. For example, if the user expresses the emotion of "happiness," a scenario reflecting that emotion will be created.

[0629] Step 7:

[0630] The server produces video content using the generated scenario. It creates storyboards based on the scenario and then renders them as short films or videos. Character expressions and attitudes are also adjusted based on the output of the emotion engine.

[0631] Step 8:

[0632] The server creates interactive game content based on the generated scenario. The game is designed to include interactive elements where the scenario changes in response to the user's emotions.

[0633] Step 9:

[0634] The user wears a VR headset using a device. The device plays video content streamed from a server, providing the user with a realistic visual experience. By watching short films, the user can relive their memories with a strong sense of presence.

[0635] Step 10:

[0636] The device runs game content in a VR environment, providing users with an interactive experience. Users can explore a virtual Hokkaido and enjoy a fun time as if they were traveling with their family.

[0637] Through the steps described above, the system of the present invention enables users to transform their memories into a high-quality, realistic form, and to save and share them. Furthermore, the introduction of an emotion engine allows for the provision of an immersive experience that reflects the user's emotions.

[0638] (Example 2)

[0639] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0640] Modern users have a need to relive their memories in a richer way, but this is difficult to achieve with current technology. Existing systems are particularly insufficient when a concrete recreation, including the emotions associated with those memories, is required. Furthermore, there is the challenge of effectively analyzing user input data and converting it into interactive content that reflects those emotions.

[0641] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for collecting one or more data from the user, such as text, voice, and images; means for performing natural language processing and image processing to analyze the collected data; means for recognizing the user's emotions based on the collected voice and text data; means for automatically generating a scenario based on the analyzed data; means for converting the generated scenario into video content and interactive game content; and means for providing the video content and game content in a virtual reality environment. This makes it possible for the user to re-experience their memories in a high-quality, realistic way, including their emotions.

[0642] A "user" is the individual who uses this system to input and re-experience memories.

[0643] "Text" refers to the written data entered by the user.

[0644] "Voice" refers to the sound data that is input when a user speaks.

[0645] "Images" refer to visual data such as photographs and drawings taken by users.

[0646] "Data collection means" refers to means of collecting one or more of the following data from users: text, audio, images, etc.

[0647] "Natural language processing" refers to the technology of analyzing text data to extract meaningful information and keywords.

[0648] "Image analysis" refers to the technique of analyzing image data to extract key elements and scenes.

[0649] "Speech recognition technology" refers to technology used to convert speech data into text data.

[0650] "Emotion recognition means" refers to means for recognizing a user's emotions based on collected audio and text data.

[0651] "Automatic scenario generation method" refers to a method for automatically generating scenarios based on analyzed data.

[0652] A "template" refers to a fixed format or pattern used in automatic scenario generation.

[0653] "Video content conversion means" refers to the means for converting a generated scenario into video content.

[0654] "Interactive game content conversion means" refers to a means for converting a generated scenario into interactive game content.

[0655] "Virtual reality environment provision means" refers to means for providing the aforementioned video content and game content in a virtual reality environment.

[0656] This invention relates to a system that concretely recreates a user's memories and converts them into text, video, or interactive game formats. Furthermore, this system incorporates an emotion engine that recognizes the user's emotions to provide a more immersive experience.

[0657] First, the user launches the application using their device. The user utilizes the device's built-in text input, voice recording, and image upload functions to input data related to their memories. Specifically, the user can write episodes in the text field, provide audio data using the voice recording function, and upload photos or images from that time. This data is then sent from the device to the server.

[0658] Next, the server analyzes the collected data. For text data, it uses Python's natural language processing libraries (such as NLTK and spaCy) to analyze it and extract important keywords and phrases. For audio data, it uses speech recognition technology (such as the Google Speech-to-Text API) to transcribe it. For image data, it uses image analysis libraries (such as OpenCV) to extract key elements and scenes.

[0659] Based on these analysis results, the server uses an emotion engine to recognize the user's emotions. The server analyzes voice tone, pitch, and volume from audio data, and context and word choice from text data to detect what emotions the user is experiencing. IBM Watson Tone Analyzer and Microsoft Azure Text Analytics API are used for this emotion recognition.

[0660] Based on the analyzed data and the results of emotion recognition, the server automatically generates scenarios using a template-based scenario generation method. For example, if the emotion engine detects that the user is feeling "happy," a scenario reflecting that emotion will be generated.

[0661] Based on the generated scenario, the server produces video content and interactive game content. The video content is rendered as short films or videos using video production tools (such as Adobe Premiere Pro and Final Cut Pro). Character expressions and attitudes are also adjusted based on the output of the emotion engine. The game content is created using a game engine (such as Unity or Unreal Engine) to produce a game that includes interactive elements where the scenario changes in response to emotions.

[0662] The final generated video and game content is delivered through the device in a virtual reality (VR) environment. The device works in conjunction with VR headsets (such as Oculus Rift and HTC Vive) to provide users with a realistic experience. This allows users to re-experience their memories, including emotional elements.

[0663] To give a concrete example, if a user inputs "memories of a family trip to Hokkaido when they were 10 years old" into the system, the user would type "trip to Hokkaido" into their terminal and upload audio recordings and photos from that time. The server analyzes this data, extracting keywords such as "family," "trip," and "Hokkaido," and obtaining information from the photos such as "snowy landscape" and "family group photo." Using an emotion engine, it detects the "happy" emotion the user felt while recounting this memory. Based on this, the server generates a scenario themed around "a family trip to Hokkaido" that reflects the "happy" emotion. A short film and an interactive game are created based on the scenario, and the user can experience them through a VR headset.

[0664] Examples of prompts to input into a generative AI model are as follows:

[0665] "Please write about a happy memory from a family trip to Hokkaido when you were 10 years old. Please include photos and audio recordings. Based on this data, please generate a scenario using an emotion-reflecting system and create video content or a game."

[0666] The above describes the form for carrying out the invention. This system allows users to transform their cherished memories into a high-quality, realistic form, including emotions, and re-experience them.

[0667] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0668] Step 1: Collecting User Input

[0669] The user launches the application using their device and inputs data related to their memories. The specific actions are as follows: The user writes episodes in a text field (input: text data), provides audio data using the voice recording function (input: audio data), and uploads photos and images (input: image data). This data is sent centrally from the device to the server (output: transmission of integrated data).

[0670] Step 2: Analysis of the month

[0671] The server receives the transmitted data and begins analysis. First, it uses Python's natural language processing libraries (e.g., NLTK, spaCy) on the text data to extract important keywords and phrases (input: text data, output: important keywords). Next, it uses speech recognition technology such as the Google Speech-to-Text API on the audio data to transcribe it (input: audio data, output: text data). Finally, it uses image analysis libraries such as OpenCV on the image data to extract key elements and scenes (input: image data, output: important elements).

[0672] Step 3: Recognizing Emotions

[0673] The server recognizes emotions based on the analyzed data. For audio data, it analyzes voice tone, pitch, and volume (input: audio data) to identify the user's emotions (output: emotion data). For text data, it uses IBM Watson Tone Analyzer or Microsoft Azure Text Analytics API to analyze context and word choice to identify emotions (input: text data, output: emotion data).

[0674] Step 4: Scenario Generation

[0675] The server automatically generates template-based scenarios based on the analysis data and emotion recognition results. Specifically, it uses a scenario generation algorithm to create stories that reflect emotional elements (input: analysis data and emotion data, output: scenario). For example, if the user indicates the emotion of "happiness," a scenario reflecting that emotion will be automatically generated.

[0676] Step 5: Generating video and game content

[0677] The server produces video content and interactive game content based on the generated scenario. For video content, it uses video production tools (e.g., Adobe Premiere Pro, Final Cut Pro) to render short films and videos based on the scenario (input: scenario, output: video content). For game content, it uses a game engine (e.g., Unity, Unreal Engine) to create interactive games that change in response to emotions (input: scenario, output: game content).

[0678] Step 6: Providing the experience

[0679] The device delivers generated video and game content in a virtual reality (VR) environment. Users wear a VR headset (e.g., Oculus Rift, HTC Vive) and experience this content (input: video and game content, output: VR experience). This allows users to re-experience their memories in a high-quality, realistic form, including the emotions they evoke.

[0680] (Application Example 2)

[0681] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0682] Traditionally, systems designed to recreate users' memories have simply displayed collected data, making it difficult to reproduce users' emotions and subtle nuances. Furthermore, it has been challenging to deliver individual user memories in real time and provide an immersive experience. Therefore, there is a need for a system that can achieve a higher level of realism and individual personalization.

[0683] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for collecting one or more data from the user, such as text, voice, and images; means for performing natural language processing and image processing to analyze the collected data; means for automatically generating a scenario based on the analyzed data; means for converting the generated scenario into video content and interactive game content; and means for streaming the scenario-based content in real time via a terminal such as a smartphone. This makes it possible to reproduce the user's emotions and memories with a high sense of realism and to provide a real-time, individually personalized experience.

[0684] A "user" refers to an individual who uses the system to input and analyze their own memories.

[0685] "Text" refers to the written data entered by the user.

[0686] "Audio" refers to audio data that records what the user has said.

[0687] "Images" refer to visual data such as photos and drawings uploaded by users.

[0688] "Natural language processing" refers to the process of analyzing collected text and audio data to extract keywords and context.

[0689] "Image analysis" refers to the process of analyzing collected image data and extracting key elements and scenes from the image.

[0690] "Automatic scenario generation" refers to a method of automatically creating scenarios that reflect user emotions using templates, based on analyzed data.

[0691] "Video content" refers to visual content rendered based on a generated scenario.

[0692] "Interactive game content" refers to game-style content that progresses through interaction with the user based on a generated scenario.

[0693] A "virtual reality environment" refers to technology that allows users to experience things in a virtual space as if they were reality.

[0694] "Smartphones and other devices" refers to portable information devices used by users to access memory-recreation streaming services.

[0695] "Streaming distribution" refers to the technology that provides generated video and game content to users in real time via a network.

[0696] "Sentiment analysis" refers to the process of analyzing the tone, pitch, and other characteristics of text and audio data collected from users to estimate their emotions.

[0697] The system of this invention reproduces user-inputted memories in real time and delivers them as personalized content that takes emotions into account. This system analyzes data collected from users, generates scenarios, and provides video content and interactive game content based on those scenarios.

[0698] Hardware and software to be used

[0699] 1. Smartphone: The primary device used by users to input text, voice, and image data and receive the results.

[0700] 2. Server / Cloud: A centralized processing system for data analysis, sentiment recognition, scenario generation, and content rendering.

[0701] 3. Natural Language Processing (NLP) libraries (e.g., SpaCy, NLTK): These libraries analyze the collected text data and extract important keywords and phrases.

[0702] 4. Speech recognition library (e.g., Google Cloud Speech-to-Text): Converts audio data into text data and transcribes the audio into text.

[0703] 5. Emotion recognition engine (e.g., IBM Watson Tone Analyzer): Analyzes the emotions in text and audio data.

[0704] 6. Video rendering engine (e.g., Unity): Generates video content and interactive game content based on the scenario.

[0705] Data processing and data calculation workflow

[0706] 1. Collecting user input:

[0707] Users input text, voice, and images using their smartphones. This data is sent to a server and collected centrally.

[0708] 2. Data Analysis:

[0709] The server uses natural language processing technology to analyze text data and extract important keywords and phrases. It also uses speech recognition technology to transcribe audio data and image analysis technology to analyze key elements of image data.

[0710] 3. Recognition of emotions:

[0711] The server uses an emotion recognition engine to analyze the user's emotions from text and voice data. This identifies the tone, pitch, volume, and context of the emotion.

[0712] 4. Scenario generation:

[0713] The server generates emotionally reflective scenarios based on the analyzed data and emotion recognition results. A template-based scenario generation method is employed to create personalized stories.

[0714] 5. Content generation:

[0715] The server creates video content and interactive game content based on the generated scenario. The generated content reflects emotional elements, and the characters' expressions and attitudes are adjusted accordingly.

[0716] 6. Providing experiences:

[0717] Ultimately, the generated video content and interactive game content are streamed in real time via devices such as smartphones. By integrating VR headsets and other devices, a more immersive experience can be provided.

[0718] Specific example

[0719] For example, if a user inputs "memories of a family trip to Hokkaido when they were 10 years old" in text, audio, or image format, the server analyzes this and extracts keywords such as "family," "trip," and "Hokkaido." In addition, it uses an emotion recognition engine to detect that the user is feeling "happy." Based on this information, the server generates a scenario themed around "a family trip to Hokkaido" that reflects the user's emotions.

[0720] Example of a prompt

[0721] Theme: Family trip to Hokkaido

[0722] Emotion: Happy

[0723] Keywords: family, travel, snowscape

[0724] This system allows users to recreate their memories, including the emotions they evoke, and experience them with a high degree of realism.

[0725] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0726] Step 1:

[0727] The user launches the application using their smartphone and inputs one or more text, audio, or image data related to their memories. For audio input, they use the voice recording function and upload photos or images taken at that time. This data is sent from the device to the server. The input data includes detailed memories (e.g., "Memories of a family trip to Hokkaido when I was 10 years old").

[0728] Step 2:

[0729] The server analyzes the collected text data using natural language processing techniques (e.g., SpaCy, NLTK) to extract important keywords and phrases. Specifically, it extracts keywords such as "family," "travel," and "Hokkaido," and understands the context of the text data. The input is text data, and the output is keywords and contextual information.

[0730] Step 3:

[0731] The server uses speech recognition technology (e.g., Google Cloud Speech-to-Text) to convert audio data into text data and perform transcription. Specifically, it converts what the user says into text format and performs similar natural language processing. The input is audio data, and the output is transcribed text data.

[0732] Step 4:

[0733] The server uses image analysis technology to analyze uploaded image data and extract key elements and scenes. Specifically, it obtains information such as "snowy landscape" or "family group photo" from the image. The input is image data, and the output is information about key elements and scenes.

[0734] Step 5:

[0735] The server analyzes the user's emotions using an emotion recognition engine (e.g., IBM Watson Tone Analyzer) based on the parsed text and audio data. Specifically, it analyzes the context of the text, the tone of the voice, and the pitch to detect whether the user is experiencing emotions such as "happy." The input is the parsed text and audio data, and the output is emotion data.

[0736] Step 6:

[0737] The server automatically generates a story using a template-based scenario generation method, based on extracted keywords, key elements, scene information, and emotion data. Specifically, it creates a scenario themed around a "family trip to Hokkaido" that reflects the emotion of "happiness." The input consists of keywords, key elements, scene information, and emotion data, and the output is the generated scenario.

[0738] Step 7:

[0739] The server produces video content and interactive game content based on the generated scenario. Specifically, it uses a video rendering engine such as Unity to create short films and interactive games while adjusting character expressions and attitudes based on emotional data. The input is the generated scenario, and the output is the video content and game content.

[0740] Step 8:

[0741] The system allows users to stream generated video content and interactive game content in real time using devices such as smartphones. Specifically, it works in conjunction with VR headsets to provide users with a high level of immersion. The input is video content and game content, and the output is the user's experience.

[0742] Through these steps, users can transform their memories into a high-quality, realistic form that includes emotions, and experience them.

[0743] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0744] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0745] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0746] [Third Embodiment]

[0747] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0748] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0749] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0750] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0751] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0752] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0753] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0754] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0755] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0756] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0757] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0758] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0759] The system of the present invention concretely recreates the user's memories and converts them into text, images, or interactive game formats, and is implemented as follows.

[0760] Collection of user input

[0761] First, users input data about their memories using a device. The device includes a dedicated application, allowing users to write their memories in a text input field. Users can also provide audio data by speaking using the voice recording function. Furthermore, users can upload related photos and images. This data is collected centrally and sent to the next processing step.

[0762] Data analysis and transformation

[0763] The collected data is analyzed by a server. The server analyzes the text entered by the user using natural language processing techniques to extract important keywords and phrases. Next, the server performs speech recognition techniques to transcribe the audio data and convert it into text data. The server also analyzes image data and uses image recognition techniques to extract key elements and scenes within the images.

[0764] Scenario generation

[0765] Based on the collected and analyzed data, the server automatically generates scenarios. The server employs a template-based scenario generation method, which efficiently creates stories based on user data. For example, if the memory is of a family trip, keywords such as "trip," "family," and "specific place name" are applied to a template, and a story is built based on that.

[0766] Generation of video and game content

[0767] Once a scenario is generated, the server converts it into video content and interactive game content. For video content, a storyboard is created based on the scenario, and short films or videos are generated accordingly. For game content, a game is created that includes interactive elements that users can experience through actions and choices.

[0768] Providing an experience

[0769] The generated video and game content is delivered through the device in a virtual reality (VR) environment. The device works in conjunction with a VR headset to stream the content, providing users with a realistic experience. This allows users to relive their memories in a new way.

[0770] Specific example

[0771] As a concrete example, consider a scenario where a user inputs "memories of a family trip to Hokkaido when they were 10 years old" into the system. The user uploads text containing "Hokkaido trip," an audio recording of what they said at the time, and photos taken during that trip. The server analyzes this data, extracting keywords such as "family," "trip," and "Hokkaido," and obtaining information from the photos such as "snowy landscape" and "family group photo." Based on this, the server generates a scenario themed around "a family trip to Hokkaido," and then creates a short film and an interactive game based on that scenario.

[0772] Users can wear a VR headset and watch movies featuring snowy landscapes of Hokkaido with their families, or play games where they collect items during their travels. This system allows users to relive their cherished memories in high quality and realism, and to save and pass them on to future generations.

[0773] The following describes the processing flow.

[0774] Step 1:

[0775] Users launch the application using their devices and input data related to their memories. Specifically, they write episodes in text fields, record audio data using the voice recording function, and upload photos and images from that time. This data is collected centrally.

[0776] Step 2:

[0777] The device sends the collected data to the server. The server tokenizes the received text data using natural language processing techniques and extracts important keywords and phrases. For example, keywords such as "family," "travel," and "Hokkaido" are extracted.

[0778] Step 3:

[0779] The server analyzes the audio data and transcribes it using speech recognition technology. The recorded audio of someone talking about their trip to Hokkaido is converted into text data. This text data is also analyzed using the aforementioned natural language processing technology.

[0780] Step 4:

[0781] The server analyzes image data and uses image recognition technology to extract key elements and scenes. For example, information such as "snowy landscape" or "family photo" can be obtained from a photograph. This allows for a clear understanding of the specific content of the image.

[0782] Step 5:

[0783] The server generates scenarios based on the collected and analyzed data. Templates are used in this process, for example, to generate a scenario such as "I traveled to Hokkaido with my family and enjoyed the snowy scenery."

[0784] Step 6:

[0785] The server uses the generated scenario to create video content. Specifically, it creates storyboards and then renders them as short films or videos. The rendered video is saved on the server.

[0786] Step 7:

[0787] The server simultaneously creates interactive game content based on the generated scenario. This game includes activities such as the user collecting items while traveling. The generated game content is also saved on the server.

[0788] Step 8:

[0789] The user wears a VR headset using a device. The device plays video content streamed from a server, providing the user with a realistic visual experience. By watching short films, the user can relive their memories with a strong sense of presence.

[0790] Step 9:

[0791] The device runs game content in a VR environment, providing users with an interactive experience. Users can explore a virtual Hokkaido and enjoy a fun time as if they were traveling with their family.

[0792] Through the steps described above, the system of the present invention enables users to transform their memories into a high-quality, realistic form, and to save and share them.

[0793] (Example 1)

[0794] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0795] Existing systems designed to recreate users' memories and deliver them in text, video, and interactive game formats face several challenges in terms of efficiency and quality. For example, incomplete data collection or inaccurate data analysis can lead to a low degree of memory recreation. Furthermore, a lack of quality in the generated content and the realism of the user experience are also problematic. Therefore, there is a need for a system that allows users to relive their memories realistically and in a high-quality manner.

[0796] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0797] In this invention, the server includes means for collecting one or more data from a user, such as text, voice, and images; means for performing natural language processing and image processing to analyze the collected data; means for automatically generating scenarios based on the analyzed data; means for converting the generated scenarios into video content and interactive game content; means for providing the video content and game content in a virtual reality environment; means for centrally managing and transmitting the collected data to the server; means for analyzing the data using natural language processing technology, speech recognition technology, and image recognition technology; and means for creating storyboards using a template-based method and generating video content and game content based on them. This enables users to relive their memories in high quality and realism, save them, and pass them on to future generations.

[0798] A "user" is an individual or group that uses this system to recreate their own memories.

[0799] A "terminal" is a device used by a user to input data and access a system, and includes smartphones, personal computers, and other similar devices.

[0800] "Data" refers to information in various forms, including text, audio, images, and other information entered by the user.

[0801] "Means of collection" refers to a device or software that has the function of centrally collecting data entered by users and transmitting it to a server.

[0802] "Means of analysis" refers to the process of analyzing collected data using natural language processing, speech recognition, and image recognition technologies.

[0803] "Natural language processing" is a technology that analyzes text data and extracts important keywords and phrases.

[0804] "Image analysis" is a technique that analyzes image data to extract key elements and scenes.

[0805] "Methods for automatically generating scenarios" refer to the process of automatically creating the structure of a story or narrative using a template-based method based on analyzed data.

[0806] "Video content" refers to visual media such as movies and videos that are produced based on a generated scenario.

[0807] "Interactive game content" refers to a type of game where users perform actions and make choices, recreating memories through the experience.

[0808] A "virtual reality environment" is a technology that allows users to experience a virtually recreated environment by wearing devices such as VR headsets.

[0809] "A means of centralized management" refers to the function of a system that unifies and manages collected data and efficiently transmits it to a server.

[0810] "Natural language processing technology" is a technique for analyzing text data and extracting information based on that analysis.

[0811] "Speech recognition technology" is a technology that converts speech data into text data and then analyzes it.

[0812] "Image recognition technology" is a technique that analyzes image data and extracts key elements or scenes from it.

[0813] A "template-based approach" is a method that efficiently generates scenarios by applying data based on a predefined template.

[0814] A "storyboard" is a visual representation of the scene structure of a video or story.

[0815] The system of this invention concretely recreates the user's memories and provides them in the form of text, images, or interactive games. Its detailed configuration and processing are described below.

[0816] Hardware and software configuration

[0817] User input

[0818] The user launches a dedicated application using their device (smartphone or personal computer). The application includes a text input field, voice recording function, and image upload function.

[0819] Data collection

[0820] The device centrally collects text data entered by the user, voice data, and uploaded image data, and sends this data to the server.

[0821] Data analysis and transformation

[0822] Natural Language Processing and Speech Recognition

[0823] The server analyzes the collected text data using natural language processing technology (e.g., Google Cloud Natural Language API) to extract important keywords and phrases. Additionally, the collected audio data is transcribed using speech recognition technology (e.g., Google Speech-to-Text) and further analyzed.

[0824] Image Recognition

[0825] The server analyzes the image data using image recognition technology (e.g., Google Cloud Vision API) to extract key elements and scenes.

[0826] Scenario generation

[0827] The server automatically generates scenarios using a template-based methodology based on the analyzed data. These scenarios enable the efficient creation of stories based on user data.

[0828] Specific example

[0829] For example, if a user enters "memories of a family trip to Hokkaido when they were 10 years old" and uploads photos from the trip and audio recordings of them talking about their memories, the server will extract keywords such as "family," "trip," and "Hokkaido," as well as information such as "snowy scenery" and "family group photo." Then, a scenario themed around "a family trip to Hokkaido" will be generated.

[0830] Generation of video and interactive game content

[0831] Video content

[0832] The server creates storyboards based on the generated scenarios, and then uses these to produce short films or videos (for example, using Adobe Premiere Pro).

[0833] Interactive game content

[0834] The server develops an interactive game that users can experience through actions and choices (for example, using Unity).

[0835] Providing the final experience

[0836] Provided in a VR environment

[0837] The generated video and interactive game content is delivered through the device in a virtual reality (VR) environment. Users wear a VR headset (e.g., Oculus Rift) and experience the content streamed from the device.

[0838] Specific examples and prompt statements

[0839] The user inputs "memories of a family trip to Hokkaido when they were 10 years old" into the system, providing audio and images. The server analyzes the collected data, generates scenarios, and creates video and game content based on them. For example, the following prompt sentences are input into the generating AI model:

[0840] Example of a prompt:

[0841] "Generate memorable scenarios based on the places and events the user has experienced during their travels."

[0842] "Create a Hokkaido travel scenario using the following keywords: family, travel, snowscape."

[0843] In this way, users can relive their precious memories in high quality and with realism, and save them to pass on to future generations.

[0844] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0845] Step 1:

[0846] The user launches a dedicated application using their device.

[0847] Specifically, the user downloads the application and logs in.

[0848] Input: User account information

[0849] Output: Application launched, user authentication successful

[0850] Step 2:

[0851] The user enters data related to their memories.

[0852] Specifically, the user writes their memories in a text input field, records their memories verbally using the voice recording function, and uploads related photos and images.

[0853] Input: Text data, audio data, image data

[0854] Output: Collected memory data

[0855] Step 3:

[0856] The device collects and centrally manages user input data (text, voice, images). Next, the device sends the collected data to a server.

[0857] Specifically, the data is organized within the device and then uploaded to the server via the network.

[0858] Input: User-submitted memory data

[0859] Output: Dataset stored on the server

[0860] Step 4:

[0861] The server analyzes the collected text data using natural language processing techniques.

[0862] In terms of specific operations, the server calls a natural language processing API to extract important keywords and phrases from the text data.

[0863] Input: Text data

[0864] Output: Extracted keywords and phrases

[0865] Step 5:

[0866] The server transcribes the audio data using speech recognition technology and converts it into text data. Next, this text data is analyzed using natural language processing technology.

[0867] Specifically, the server utilizes a speech recognition API to convert speech into text, which is then automatically analyzed.

[0868] Input: Audio data

[0869] Output: Analyzed transcript text data

[0870] Step 6:

[0871] The server analyzes the image data using image recognition technology.

[0872] In terms of specific operations, the server calls an image recognition API to extract key elements and scenes from the image data.

[0873] Input: Image data

[0874] Output: Extracted key elements and scenes

[0875] Step 7:

[0876] The server automatically generates scenarios based on the analyzed data.

[0877] Specifically, a scenario generation algorithm is applied to generate stories using a template-based approach.

[0878] Input: Information from extracted text, audio, and images

[0879] Output: Generated scenario

[0880] Step 8:

[0881] The server creates video content based on the generated scenario.

[0882] Specifically, this involves creating storyboards according to a scenario and then producing short films or videos based on those storyboards. Video editing tools are used for this purpose.

[0883] Input: Generated scenario

[0884] Output: Short film or video content

[0885] Step 9:

[0886] The server develops interactive games based on generated scenarios.

[0887] Specifically, this involves using a game development platform to create interactive games that users can experience through actions and choices.

[0888] Input: Generated scenario

[0889] Output: Interactive game content

[0890] Step 10:

[0891] The device provides generated video and game content in a VR environment.

[0892] Specifically, the device works in conjunction with a VR headset to stream content and deliver it to the user.

[0893] Input: Video content, game content

[0894] Output: User experience in a VR environment

[0895] Through these steps, this system enables users to relive, save, and pass on their precious memories to future generations in high quality and with realism.

[0896] (Application Example 1)

[0897] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0898] Traditionally, there have been limited means to concretely recreate users' memories, making it difficult to preserve individual memories as narratives or images. Furthermore, while methods existed to collect and analyze user memories, there was no way to recreate them in a physical customer experience environment. As a result, users lacked opportunities to easily re-experience their memories in a high-quality format.

[0899] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0900] In this invention, the server includes means for collecting one or more data from a user, such as text, voice, and images; means for performing natural language processing and image processing to analyze the collected data; means for automatically generating scenarios based on the analyzed data; means for converting the generated scenarios into video content and interactive game content; and means for providing the video content and game content in a physical customer experience environment. This enables users to re-experience their cherished memories in high quality in a physical environment such as a retail store.

[0901] A "user" is an individual or group that uses the system to provide data about memories and receives the recreated content of those memories.

[0902] "Text" refers to data entered by the user, representing the user's memories and information as written text.

[0903] "Voice" refers to data provided by users through speech, which can be converted into text data using speech recognition technology, representing the user's memories and information.

[0904] "Images" refer to photographs and drawings uploaded by users, and are data that includes visual records of memories.

[0905] "Natural language processing" is a technology that analyzes text and audio data collected from users to extract important keywords and phrases.

[0906] "Image analysis" is a technology that analyzes image data collected from users to extract key elements and scenes within the image.

[0907] A "scenario" is a story or narrative generated based on collected and analyzed data, and serves as the basis for video content and game content.

[0908] A "generative AI model" is a type of artificial intelligence that automatically generates stories based on user data, and in particular, it is a model that uses natural language generation technology.

[0909] "Video content" refers to digital content in the form of videos or movies created based on a generated scenario.

[0910] "Interactive game content" refers to digital games based on generated scenarios that users can experience through actions and choices.

[0911] A "physical customer experience environment" refers to a physical store or other physical space where users experience reproduced content.

[0912] This invention is a system that concretely recreates users' memories and provides that recreated content in a physical customer experience environment. The details are described below.

[0913] Collection of user input

[0914] First, users use a device to input data about their memories. The device includes a dedicated application, allowing users to write their memories in a text input field. Users can also provide audio data by speaking using the voice recording function. Furthermore, users can upload related photos and images. This data is collected centrally and sent to a server.

[0915] Data analysis

[0916] The server analyzes text data entered by users using natural language processing techniques to extract important keywords and phrases. Specifically, the text data is analyzed using natural language processing. The server also performs speech recognition techniques to transcribe audio data and convert it into text data. Furthermore, the server analyzes image data and uses image recognition techniques to extract key elements and scenes within the images.

[0917] Scenario generation

[0918] Based on the collected and analyzed data, the server automatically generates scenarios. This uses a generative AI model. The server employs a template-based scenario generation method, for example, by applying keywords such as "travel," "family," and "specific place names" to templates and building stories based on them.

[0919] Generation of video and game content

[0920] After the scenario is generated, the server converts it into video content and interactive game content. For video content, short films and videos are generated based on the scenario. For interactive game content, a game is created that includes elements that users can experience through actions and choices.

[0921] Providing an experience

[0922] The generated video and game content will be delivered in a physical customer experience environment accessible to users. For example, users can experience the content on the spot using tablet devices or robots in a physical store.

[0923] Specific example

[0924] As a concrete example, consider a scenario where a user inputs "memories of a family trip to Hokkaido" into the system. The user uploads text such as "Hokkaido trip," an audio recording of what they said at the time, and photos taken during that trip. The server analyzes this data, extracting keywords such as "family," "trip," and "Hokkaido," and obtaining information from the photos such as "snowy landscape" and "family group photo." Based on this, the server generates a scenario themed around "a family trip to Hokkaido," and then creates a short film and an interactive game based on that scenario.

[0925] Users can use terminals in physical stores to watch movies featuring family members enjoying the snowy landscapes of Hokkaido, and participate in interactive games where they can collect items during their trip.

[0926] Example of a prompt

[0927] For example, when a user is recounting a memory, the following might be used as a prompt:

[0928] "I remember a trip I took to Hokkaido with my family. We enjoyed a warm hot spring bath while admiring the beautiful snowy scenery."

[0929] Based on this prompt, the generative AI model can generate a scenario and provide the user with high-quality content that recreates their memories.

[0930] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0931] Step 1:

[0932] Users use their devices to input one or more data related to their memories: text, audio, images, and so on. This includes typing into text fields, recording audio using a microphone, and uploading photos and images. The entered data is sent to the server.

[0933] Input: Text, audio, images

[0934] Output: Data to send to the server

[0935] Step 2:

[0936] The server analyzes the received text data using natural language processing techniques to extract important keywords and phrases. Specifically, it tokenizes the text data, performs morphological analysis, and extracts meaningful words and phrases.

[0937] Input: Text data

[0938] Output: Extracted keywords and phrases

[0939] Step 3:

[0940] The server performs speech recognition technology to transcribe the audio data and convert it into text data. The converted text data is similarly analyzed using natural language processing technology to extract keywords and phrases.

[0941] Input: Audio data

[0942] Output: Text data, extracted keywords and phrases

[0943] Step 4:

[0944] The server analyzes the received image data and uses image recognition techniques to extract key elements and scenes from the image. Image analysis employs methods such as object detection and scene classification.

[0945] Input: Image data

[0946] Output: Extracted key elements and scenes

[0947] Step 5:

[0948] The server automatically generates scenarios using a generative AI model based on extracted keywords, phrases, key elements, and scenes. The generative AI model generates a story based on the input prompt sentences.

[0949] Input: Keywords, phrases, key elements or scenes

[0950] Output: Generated scenario

[0951] Step 6:

[0952] The server converts the generated scenarios into video content and interactive game content. For video content, short films and videos are generated based on the scenarios, while interactive games include elements that users can experience through actions and choices.

[0953] Input: Generated scenario

[0954] Output: Video content, interactive game content

[0955] Step 7:

[0956] The generated video and game content will be provided to users to experience on the spot using terminals and robots within physical stores.

[0957] Input: Video content, interactive game content

[0958] Output: Providing experiences in a physical customer experience environment

[0959] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0960] The system of the present invention concretely recreates the user's memories and converts them into text, images, and interactive game formats. Furthermore, by combining this with an emotion engine that recognizes the user's emotions, it provides a more immersive experience. The following are specific embodiments of this system.

[0961] Collection of user input

[0962] First, the user launches the application using their device and enters data about their memories. Users can write episodes in a text field. They can also provide audio data by speaking about their memories using the voice recording function. Furthermore, they can upload photos and images from that time. This data is collected centrally and sent to the next processing step.

[0963] Data analysis and transformation

[0964] The collected data is analyzed by a server. The server analyzes the text entered by the user using natural language processing technology to extract important keywords and phrases. Next, the server analyzes the audio data and transcribes it using speech recognition technology. The server also analyzes the image data and extracts the main elements and scenes within the images.

[0965] Recognition of emotions

[0966] Next, the server uses an emotion engine to recognize the user's emotions based on the collected audio and text data. From the audio data, it analyzes voice tone, pitch, and volume, and from the text data, it analyzes context and word choice to detect what emotions the user is feeling.

[0967] Scenario generation

[0968] Based on the collected and analyzed data and the output of the emotion engine, the server automatically generates scenarios. Specifically, it employs a template-based scenario generation method to create stories that reflect emotional elements. For example, if a user expresses the emotion of "happiness," a scenario reflecting that emotion will be created.

[0969] Generation of video and game content

[0970] Based on the generated scenario, the server produces video content and interactive game content. For video content, storyboards are created based on the scenario and then rendered as short films or videos. Character expressions and attitudes are also adjusted based on the output of the emotion engine. In game content, interactive elements are created in which the scenario changes according to the user's emotions.

[0971] Providing an experience

[0972] The generated video and game content is delivered through the device in a virtual reality (VR) environment. The device works in conjunction with a VR headset to stream the content, providing users with a realistic experience. This allows users to relive their memories, including emotional elements.

[0973] Specific example

[0974] As a concrete example, consider a scenario where a user inputs "memories of a family trip to Hokkaido when they were 10 years old" into the system. The user uploads text containing "trip to Hokkaido," an audio recording of their conversation at the time, and photos taken during that trip. The server analyzes this data, extracting keywords such as "family," "trip," and "Hokkaido," and obtaining information from the photos such as "snowy landscape" and "family group photo." Furthermore, using an emotion engine, the system detects whether the user was feeling "happy" while recounting this memory.

[0975] Based on this, the server generates a scenario themed around a "family trip to Hokkaido," reflecting the emotion of "happiness." For example, the scenario will include many scenes where the whole family is smiling and enjoying themselves. A short film and an interactive game are then created based on this scenario.

[0976] The user wears a VR headset and watches a short film of a family enjoying a snowball fight in a park. In the game, the user can also experience a task of collecting items with their family, all while smiling. This system allows users to experience their cherished memories in a high-quality, realistic form, including the emotions they evoke.

[0977] The following describes the processing flow.

[0978] Step 1:

[0979] Users launch the application using their devices and input data related to their memories. Specifically, they write episodes in text fields, record audio data using the voice recording function, and upload related photos and images. This data is centrally collected on the device and sent to the next processing step.

[0980] Step 2:

[0981] The device sends the collected data to the server. The server receives the text data, tokenizes it using natural language processing techniques, and extracts important keywords and phrases. For example, keywords such as "family," "travel," and "Hokkaido" are extracted.

[0982] Step 3:

[0983] The server analyzes the received audio data. It uses speech recognition technology to convert the audio into text, and then analyzes the resulting text using natural language processing technology to extract important information.

[0984] Step 4:

[0985] The server analyzes the image data. Using image recognition technology, it extracts key elements and scenes from the image. For example, information such as "snowy landscape" or "family photo" can be obtained from a photograph.

[0986] Step 5:

[0987] The server uses an emotion engine to recognize the user's emotions based on collected audio and text data. From the audio data, it analyzes voice tone, pitch, and volume, and from the text data, it analyzes context and word choice to detect what emotions the user is feeling.

[0988] Step 6:

[0989] The server generates scenarios based on the analyzed data and the output of the emotion engine. Using a template-based scenario generation method, it constructs a story that reflects the emotions expressed by the user. For example, if the user expresses the emotion of "happiness," a scenario reflecting that emotion will be created.

[0990] Step 7:

[0991] The server produces video content using the generated scenario. It creates storyboards based on the scenario and then renders them as short films or videos. Character expressions and attitudes are also adjusted based on the output of the emotion engine.

[0992] Step 8:

[0993] The server creates interactive game content based on the generated scenario. The game is designed to include interactive elements where the scenario changes in response to the user's emotions.

[0994] Step 9:

[0995] The user wears a VR headset using a device. The device plays video content streamed from a server, providing the user with a realistic visual experience. By watching short films, the user can relive their memories with a strong sense of presence.

[0996] Step 10:

[0997] The device runs game content in a VR environment, providing users with an interactive experience. Users can explore a virtual Hokkaido and enjoy a fun time as if they were traveling with their family.

[0998] Through the steps described above, the system of the present invention enables users to transform their memories into a high-quality, realistic form, and to save and share them. Furthermore, the introduction of an emotion engine allows for the provision of an immersive experience that reflects the user's emotions.

[0999] (Example 2)

[1000] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[1001] Modern users have a need to relive their memories in a richer way, but this is difficult to achieve with current technology. Existing systems are particularly insufficient when a concrete recreation, including the emotions associated with those memories, is required. Furthermore, there is the challenge of effectively analyzing user input data and converting it into interactive content that reflects those emotions.

[1002] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for collecting one or more data from the user, such as text, voice, and images; means for performing natural language processing and image processing to analyze the collected data; means for recognizing the user's emotions based on the collected voice and text data; means for automatically generating a scenario based on the analyzed data; means for converting the generated scenario into video content and interactive game content; and means for providing the video content and game content in a virtual reality environment. This makes it possible for the user to re-experience their memories in a high-quality, realistic way, including their emotions.

[1003] A "user" is the individual who uses this system to input and re-experience memories.

[1004] "Text" refers to the written data entered by the user.

[1005] "Voice" refers to the sound data that is input when a user speaks.

[1006] "Images" refer to visual data such as photographs and drawings taken by users.

[1007] "Data collection means" refers to means of collecting one or more of the following data from users: text, audio, images, etc.

[1008] "Natural language processing" refers to the technology of analyzing text data to extract meaningful information and keywords.

[1009] "Image analysis" refers to the technique of analyzing image data to extract key elements and scenes.

[1010] "Speech recognition technology" refers to technology used to convert speech data into text data.

[1011] "Emotion recognition means" refers to means for recognizing a user's emotions based on collected audio and text data.

[1012] "Automatic scenario generation method" refers to a method for automatically generating scenarios based on analyzed data.

[1013] A "template" refers to a fixed format or pattern used in automatic scenario generation.

[1014] "Video content conversion means" refers to the means for converting a generated scenario into video content.

[1015] "Interactive game content conversion means" refers to a means for converting a generated scenario into interactive game content.

[1016] "Virtual reality environment provision means" refers to means for providing the aforementioned video content and game content in a virtual reality environment.

[1017] This invention relates to a system that concretely recreates a user's memories and converts them into text, video, or interactive game formats. Furthermore, this system incorporates an emotion engine that recognizes the user's emotions to provide a more immersive experience.

[1018] First, the user launches the application using their device. The user utilizes the device's built-in text input, voice recording, and image upload functions to input data related to their memories. Specifically, the user can write episodes in the text field, provide audio data using the voice recording function, and upload photos or images from that time. This data is then sent from the device to the server.

[1019] Next, the server analyzes the collected data. For text data, it uses Python's natural language processing libraries (such as NLTK and spaCy) to analyze it and extract important keywords and phrases. For audio data, it uses speech recognition technology (such as the Google Speech-to-Text API) to transcribe it. For image data, it uses image analysis libraries (such as OpenCV) to extract key elements and scenes.

[1020] Based on these analysis results, the server uses an emotion engine to recognize the user's emotions. The server analyzes voice tone, pitch, and volume from audio data, and context and word choice from text data to detect what emotions the user is experiencing. IBM Watson Tone Analyzer and Microsoft Azure Text Analytics API are used for this emotion recognition.

[1021] Based on the analyzed data and the results of emotion recognition, the server automatically generates scenarios using a template-based scenario generation method. For example, if the emotion engine detects that the user is feeling "happy," a scenario reflecting that emotion will be generated.

[1022] Based on the generated scenario, the server produces video content and interactive game content. The video content is rendered as short films or videos using video production tools (such as Adobe Premiere Pro and Final Cut Pro). Character expressions and attitudes are also adjusted based on the output of the emotion engine. The game content is created using a game engine (such as Unity or Unreal Engine) to produce a game that includes interactive elements where the scenario changes in response to emotions.

[1023] The final generated video and game content is delivered through the device in a virtual reality (VR) environment. The device works in conjunction with VR headsets (such as Oculus Rift and HTC Vive) to provide users with a realistic experience. This allows users to re-experience their memories, including emotional elements.

[1024] To give a concrete example, if a user inputs "memories of a family trip to Hokkaido when they were 10 years old" into the system, the user would type "trip to Hokkaido" into their terminal and upload audio recordings and photos from that time. The server analyzes this data, extracting keywords such as "family," "trip," and "Hokkaido," and obtaining information from the photos such as "snowy landscape" and "family group photo." Using an emotion engine, it detects the "happy" emotion the user felt while recounting this memory. Based on this, the server generates a scenario themed around "a family trip to Hokkaido" that reflects the "happy" emotion. A short film and an interactive game are created based on the scenario, and the user can experience them through a VR headset.

[1025] Examples of prompts to input into a generative AI model are as follows:

[1026] "Please write about a happy memory from a family trip to Hokkaido when you were 10 years old. Please include photos and audio recordings. Based on this data, please generate a scenario using an emotion-reflecting system and create video content or a game."

[1027] The above describes the form for carrying out the invention. This system allows users to transform their cherished memories into a high-quality, realistic form, including emotions, and re-experience them.

[1028] The flow of the specific processing in Example 2 will be explained using Figure 13.

[1029] Step 1: Collecting User Input

[1030] The user launches the application using their device and inputs data related to their memories. The specific actions are as follows: The user writes episodes in a text field (input: text data), provides audio data using the voice recording function (input: audio data), and uploads photos and images (input: image data). This data is sent centrally from the device to the server (output: transmission of integrated data).

[1031] Step 2: Analysis of the month

[1032] The server receives the transmitted data and begins analysis. First, it uses Python's natural language processing libraries (e.g., NLTK, spaCy) on the text data to extract important keywords and phrases (input: text data, output: important keywords). Next, it uses speech recognition technology such as the Google Speech-to-Text API on the audio data to transcribe it (input: audio data, output: text data). Finally, it uses image analysis libraries such as OpenCV on the image data to extract key elements and scenes (input: image data, output: important elements).

[1033] Step 3: Recognizing Emotions

[1034] The server recognizes emotions based on the analyzed data. For audio data, it analyzes voice tone, pitch, and volume (input: audio data) to identify the user's emotions (output: emotion data). For text data, it uses IBM Watson Tone Analyzer or Microsoft Azure Text Analytics API to analyze context and word choice to identify emotions (input: text data, output: emotion data).

[1035] Step 4: Scenario Generation

[1036] The server automatically generates template-based scenarios based on the analysis data and emotion recognition results. Specifically, it uses a scenario generation algorithm to create stories that reflect emotional elements (input: analysis data and emotion data, output: scenario). For example, if the user indicates the emotion of "happiness," a scenario reflecting that emotion will be automatically generated.

[1037] Step 5: Generating video and game content

[1038] The server produces video content and interactive game content based on the generated scenario. For video content, it uses video production tools (e.g., Adobe Premiere Pro, Final Cut Pro) to render short films and videos based on the scenario (input: scenario, output: video content). For game content, it uses a game engine (e.g., Unity, Unreal Engine) to create interactive games that change in response to emotions (input: scenario, output: game content).

[1039] Step 6: Providing the experience

[1040] The device delivers generated video and game content in a virtual reality (VR) environment. Users wear a VR headset (e.g., Oculus Rift, HTC Vive) and experience this content (input: video and game content, output: VR experience). This allows users to re-experience their memories in a high-quality, realistic form, including the emotions they evoke.

[1041] (Application Example 2)

[1042] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[1043] Traditionally, systems designed to recreate users' memories have simply displayed collected data, making it difficult to reproduce users' emotions and subtle nuances. Furthermore, it has been challenging to deliver individual user memories in real time and provide an immersive experience. Therefore, there is a need for a system that can achieve a higher level of realism and individual personalization.

[1044] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for collecting one or more data from the user, such as text, voice, and images; means for performing natural language processing and image processing to analyze the collected data; means for automatically generating a scenario based on the analyzed data; means for converting the generated scenario into video content and interactive game content; and means for streaming the scenario-based content in real time via a terminal such as a smartphone. This makes it possible to reproduce the user's emotions and memories with a high sense of realism and to provide a real-time, individually personalized experience.

[1045] A "user" refers to an individual who uses the system to input and analyze their own memories.

[1046] "Text" refers to the written data entered by the user.

[1047] "Audio" refers to audio data that records what the user has said.

[1048] "Images" refer to visual data such as photos and drawings uploaded by users.

[1049] "Natural language processing" refers to the process of analyzing collected text and audio data to extract keywords and context.

[1050] "Image analysis" refers to the process of analyzing collected image data and extracting key elements and scenes from the image.

[1051] "Automatic scenario generation" refers to a method of automatically creating scenarios that reflect user emotions using templates, based on analyzed data.

[1052] "Video content" refers to visual content rendered based on a generated scenario.

[1053] "Interactive game content" refers to game-style content that progresses through interaction with the user based on a generated scenario.

[1054] A "virtual reality environment" refers to technology that allows users to experience things in a virtual space as if they were reality.

[1055] "Smartphones and other devices" refers to portable information devices used by users to access memory-recreation streaming services.

[1056] "Streaming distribution" refers to the technology that provides generated video and game content to users in real time via a network.

[1057] "Sentiment analysis" refers to the process of analyzing the tone, pitch, and other characteristics of text and audio data collected from users to estimate their emotions.

[1058] The system of this invention reproduces user-inputted memories in real time and delivers them as personalized content that takes emotions into account. This system analyzes data collected from users, generates scenarios, and provides video content and interactive game content based on those scenarios.

[1059] Hardware and software to be used

[1060] 1. Smartphone: The primary device used by users to input text, voice, and image data and receive the results.

[1061] 2. Server / Cloud: A centralized processing system for data analysis, sentiment recognition, scenario generation, and content rendering.

[1062] 3. Natural Language Processing (NLP) libraries (e.g., SpaCy, NLTK): These libraries analyze the collected text data and extract important keywords and phrases.

[1063] 4. Speech recognition library (e.g., Google Cloud Speech-to-Text): Converts audio data into text data and transcribes the audio into text.

[1064] 5. Emotion recognition engine (e.g., IBM Watson Tone Analyzer): Analyzes the emotions in text and audio data.

[1065] 6. Video rendering engine (e.g., Unity): Generates video content and interactive game content based on the scenario.

[1066] Data processing and data calculation workflow

[1067] 1. Collecting user input:

[1068] Users input text, voice, and images using their smartphones. This data is sent to a server and collected centrally.

[1069] 2. Data Analysis:

[1070] The server uses natural language processing technology to analyze text data and extract important keywords and phrases. It also uses speech recognition technology to transcribe audio data and image analysis technology to analyze key elements of image data.

[1071] 3. Recognition of emotions:

[1072] The server uses an emotion recognition engine to analyze the user's emotions from text and voice data. This identifies the tone, pitch, volume, and context of the emotion.

[1073] 4. Scenario generation:

[1074] The server generates emotionally reflective scenarios based on the analyzed data and emotion recognition results. A template-based scenario generation method is employed to create personalized stories.

[1075] 5. Content generation:

[1076] The server creates video content and interactive game content based on the generated scenario. The generated content reflects emotional elements, and the characters' expressions and attitudes are adjusted accordingly.

[1077] 6. Providing experiences:

[1078] Ultimately, the generated video content and interactive game content are streamed in real time via devices such as smartphones. By integrating VR headsets and other devices, a more immersive experience can be provided.

[1079] Specific example

[1080] For example, if a user inputs "memories of a family trip to Hokkaido when they were 10 years old" in text, audio, or image format, the server analyzes this and extracts keywords such as "family," "trip," and "Hokkaido." In addition, it uses an emotion recognition engine to detect that the user is feeling "happy." Based on this information, the server generates a scenario themed around "a family trip to Hokkaido" that reflects the user's emotions.

[1081] Example of a prompt

[1082] Theme: Family trip to Hokkaido

[1083] Emotion: Happy

[1084] Keywords: family, travel, snowscape

[1085] This system allows users to recreate their memories, including the emotions they evoke, and experience them with a high degree of realism.

[1086] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[1087] Step 1:

[1088] The user launches the application using their smartphone and inputs one or more text, audio, or image data related to their memories. For audio input, they use the voice recording function and upload photos or images taken at that time. This data is sent from the device to the server. The input data includes detailed memories (e.g., "Memories of a family trip to Hokkaido when I was 10 years old").

[1089] Step 2:

[1090] The server analyzes the collected text data using natural language processing techniques (e.g., SpaCy, NLTK) to extract important keywords and phrases. Specifically, it extracts keywords such as "family," "travel," and "Hokkaido," and understands the context of the text data. The input is text data, and the output is keywords and contextual information.

[1091] Step 3:

[1092] The server uses speech recognition technology (e.g., Google Cloud Speech-to-Text) to convert audio data into text data and perform transcription. Specifically, it converts what the user says into text format and performs similar natural language processing. The input is audio data, and the output is transcribed text data.

[1093] Step 4:

[1094] The server uses image analysis technology to analyze uploaded image data and extract key elements and scenes. Specifically, it obtains information such as "snowy landscape" or "family group photo" from the image. The input is image data, and the output is information about key elements and scenes.

[1095] Step 5:

[1096] The server analyzes the user's emotions using an emotion recognition engine (e.g., IBM Watson Tone Analyzer) based on the parsed text and audio data. Specifically, it analyzes the context of the text, the tone of the voice, and the pitch to detect whether the user is experiencing emotions such as "happy." The input is the parsed text and audio data, and the output is emotion data.

[1097] Step 6:

[1098] The server automatically generates a story using a template-based scenario generation method, based on extracted keywords, key elements, scene information, and emotion data. Specifically, it creates a scenario themed around a "family trip to Hokkaido" that reflects the emotion of "happiness." The input consists of keywords, key elements, scene information, and emotion data, and the output is the generated scenario.

[1099] Step 7:

[1100] The server produces video content and interactive game content based on the generated scenario. Specifically, it uses a video rendering engine such as Unity to create short films and interactive games while adjusting character expressions and attitudes based on emotional data. The input is the generated scenario, and the output is the video content and game content.

[1101] Step 8:

[1102] The system allows users to stream generated video content and interactive game content in real time using devices such as smartphones. Specifically, it works in conjunction with VR headsets to provide users with a high level of immersion. The input is video content and game content, and the output is the user's experience.

[1103] Through these steps, users can transform their memories into a high-quality, realistic form that includes emotions, and experience them.

[1104] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[1105] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1106] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[1107] [Fourth Embodiment]

[1108] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[1109] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1110] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1111] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[1112] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[1113] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[1114] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[1115] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[1116] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[1117] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1118] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1119] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[1120] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1121] The system of the present invention concretely recreates the user's memories and converts them into text, images, or interactive game formats, and is implemented as follows.

[1122] Collection of user input

[1123] First, users input data about their memories using a device. The device includes a dedicated application, allowing users to write their memories in a text input field. Users can also provide audio data by speaking using the voice recording function. Furthermore, users can upload related photos and images. This data is collected centrally and sent to the next processing step.

[1124] Data analysis and transformation

[1125] The collected data is analyzed by a server. The server analyzes the text entered by the user using natural language processing techniques to extract important keywords and phrases. Next, the server performs speech recognition techniques to transcribe the audio data and convert it into text data. The server also analyzes image data and uses image recognition techniques to extract key elements and scenes within the images.

[1126] Scenario generation

[1127] Based on the collected and analyzed data, the server automatically generates scenarios. The server employs a template-based scenario generation method, which efficiently creates stories based on user data. For example, if the memory is of a family trip, keywords such as "trip," "family," and "specific place name" are applied to a template, and a story is built based on that.

[1128] Generation of video and game content

[1129] Once a scenario is generated, the server converts it into video content and interactive game content. For video content, a storyboard is created based on the scenario, and short films or videos are generated accordingly. For game content, a game is created that includes interactive elements that users can experience through actions and choices.

[1130] Providing an experience

[1131] The generated video and game content is delivered through the device in a virtual reality (VR) environment. The device works in conjunction with a VR headset to stream the content, providing users with a realistic experience. This allows users to relive their memories in a new way.

[1132] Specific example

[1133] As a concrete example, consider a scenario where a user inputs "memories of a family trip to Hokkaido when they were 10 years old" into the system. The user uploads text containing "Hokkaido trip," an audio recording of what they said at the time, and photos taken during that trip. The server analyzes this data, extracting keywords such as "family," "trip," and "Hokkaido," and obtaining information from the photos such as "snowy landscape" and "family group photo." Based on this, the server generates a scenario themed around "a family trip to Hokkaido," and then creates a short film and an interactive game based on that scenario.

[1134] Users can wear a VR headset and watch movies featuring snowy landscapes of Hokkaido with their families, or play games where they collect items during their travels. This system allows users to relive their cherished memories in high quality and realism, and to save and pass them on to future generations.

[1135] The following describes the processing flow.

[1136] Step 1:

[1137] Users launch the application using their devices and input data related to their memories. Specifically, they write episodes in text fields, record audio data using the voice recording function, and upload photos and images from that time. This data is collected centrally.

[1138] Step 2:

[1139] The device sends the collected data to the server. The server tokenizes the received text data using natural language processing techniques and extracts important keywords and phrases. For example, keywords such as "family," "travel," and "Hokkaido" are extracted.

[1140] Step 3:

[1141] The server analyzes the audio data and transcribes it using speech recognition technology. The recorded audio of someone talking about their trip to Hokkaido is converted into text data. This text data is also analyzed using the aforementioned natural language processing technology.

[1142] Step 4:

[1143] The server analyzes image data and uses image recognition technology to extract key elements and scenes. For example, information such as "snowy landscape" or "family photo" can be obtained from a photograph. This allows for a clear understanding of the specific content of the image.

[1144] Step 5:

[1145] The server generates scenarios based on the collected and analyzed data. Templates are used in this process, for example, to generate a scenario such as "I traveled to Hokkaido with my family and enjoyed the snowy scenery."

[1146] Step 6:

[1147] The server uses the generated scenario to create video content. Specifically, it creates storyboards and then renders them as short films or videos. The rendered video is saved on the server.

[1148] Step 7:

[1149] The server simultaneously creates interactive game content based on the generated scenario. This game includes activities such as the user collecting items while traveling. The generated game content is also saved on the server.

[1150] Step 8:

[1151] The user wears a VR headset using a device. The device plays video content streamed from a server, providing the user with a realistic visual experience. By watching short films, the user can relive their memories with a strong sense of presence.

[1152] Step 9:

[1153] The device runs game content in a VR environment, providing users with an interactive experience. Users can explore a virtual Hokkaido and enjoy a fun time as if they were traveling with their family.

[1154] Through the steps described above, the system of the present invention enables users to transform their memories into a high-quality, realistic form, and to save and share them.

[1155] (Example 1)

[1156] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1157] Existing systems designed to recreate users' memories and deliver them in text, video, and interactive game formats face several challenges in terms of efficiency and quality. For example, incomplete data collection or inaccurate data analysis can lead to a low degree of memory recreation. Furthermore, a lack of quality in the generated content and the realism of the user experience are also problematic. Therefore, there is a need for a system that allows users to relive their memories realistically and in a high-quality manner.

[1158] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[1159] In this invention, the server includes means for collecting one or more data from a user, such as text, voice, and images; means for performing natural language processing and image processing to analyze the collected data; means for automatically generating scenarios based on the analyzed data; means for converting the generated scenarios into video content and interactive game content; means for providing the video content and game content in a virtual reality environment; means for centrally managing and transmitting the collected data to the server; means for analyzing the data using natural language processing technology, speech recognition technology, and image recognition technology; and means for creating storyboards using a template-based method and generating video content and game content based on them. This enables users to relive their memories in high quality and realism, save them, and pass them on to future generations.

[1160] A "user" is an individual or group that uses this system to recreate their own memories.

[1161] A "terminal" is a device used by a user to input data and access a system, and includes smartphones, personal computers, and other similar devices.

[1162] "Data" refers to information in various forms, including text, audio, images, and other information entered by the user.

[1163] "Means of collection" refers to a device or software that has the function of centrally collecting data entered by users and transmitting it to a server.

[1164] "Means of analysis" refers to the process of analyzing collected data using natural language processing, speech recognition, and image recognition technologies.

[1165] "Natural language processing" is a technology that analyzes text data and extracts important keywords and phrases.

[1166] "Image analysis" is a technique that analyzes image data to extract key elements and scenes.

[1167] "Methods for automatically generating scenarios" refer to the process of automatically creating the structure of a story or narrative using a template-based method based on analyzed data.

[1168] "Video content" refers to visual media such as movies and videos that are produced based on a generated scenario.

[1169] "Interactive game content" refers to a type of game where users perform actions and make choices, recreating memories through the experience.

[1170] A "virtual reality environment" is a technology that allows users to experience a virtually recreated environment by wearing devices such as VR headsets.

[1171] "A means of centralized management" refers to the function of a system that unifies and manages collected data and efficiently transmits it to a server.

[1172] "Natural language processing technology" is a technique for analyzing text data and extracting information based on that analysis.

[1173] "Speech recognition technology" is a technology that converts speech data into text data and then analyzes it.

[1174] "Image recognition technology" is a technique that analyzes image data and extracts key elements or scenes from it.

[1175] A "template-based approach" is a method that efficiently generates scenarios by applying data based on a predefined template.

[1176] A "storyboard" is a visual representation of the scene structure of a video or story.

[1177] The system of this invention concretely recreates the user's memories and provides them in the form of text, images, or interactive games. Its detailed configuration and processing are described below.

[1178] Hardware and software configuration

[1179] User input

[1180] The user launches a dedicated application using their device (smartphone or personal computer). The application includes a text input field, voice recording function, and image upload function.

[1181] Data collection

[1182] The device centrally collects text data entered by the user, voice data, and uploaded image data, and sends this data to the server.

[1183] Data analysis and transformation

[1184] Natural Language Processing and Speech Recognition

[1185] The server analyzes the collected text data using natural language processing technology (e.g., Google Cloud Natural Language API) to extract important keywords and phrases. Additionally, the collected audio data is transcribed using speech recognition technology (e.g., Google Speech-to-Text) and further analyzed.

[1186] Image Recognition

[1187] The server analyzes the image data using image recognition technology (e.g., Google Cloud Vision API) to extract key elements and scenes.

[1188] Scenario generation

[1189] The server automatically generates scenarios using a template-based methodology based on the analyzed data. These scenarios enable the efficient creation of stories based on user data.

[1190] Specific example

[1191] For example, if a user enters "memories of a family trip to Hokkaido when they were 10 years old" and uploads photos from the trip and audio recordings of them talking about their memories, the server will extract keywords such as "family," "trip," and "Hokkaido," as well as information such as "snowy scenery" and "family group photo." Then, a scenario themed around "a family trip to Hokkaido" will be generated.

[1192] Generation of video and interactive game content

[1193] Video content

[1194] The server creates storyboards based on the generated scenarios, and then uses these to produce short films or videos (for example, using Adobe Premiere Pro).

[1195] Interactive game content

[1196] The server develops an interactive game that users can experience through actions and choices (for example, using Unity).

[1197] Providing the final experience

[1198] Provided in a VR environment

[1199] The generated video and interactive game content is delivered through the device in a virtual reality (VR) environment. Users wear a VR headset (e.g., Oculus Rift) and experience the content streamed from the device.

[1200] Specific examples and prompt statements

[1201] The user inputs "memories of a family trip to Hokkaido when they were 10 years old" into the system, providing audio and images. The server analyzes the collected data, generates scenarios, and creates video and game content based on them. For example, the following prompt sentences are input into the generating AI model:

[1202] Example of a prompt:

[1203] "Generate memorable scenarios based on the places and events the user has experienced during their travels."

[1204] "Create a Hokkaido travel scenario using the following keywords: family, travel, snowscape."

[1205] In this way, users can relive their precious memories in high quality and with realism, and save them to pass on to future generations.

[1206] The flow of the specific processing in Example 1 will be explained using Figure 11.

[1207] Step 1:

[1208] The user launches a dedicated application using their device.

[1209] Specifically, the user downloads the application and logs in.

[1210] Input: User account information

[1211] Output: Application launched, user authentication successful

[1212] Step 2:

[1213] The user enters data related to their memories.

[1214] Specifically, the user writes their memories in a text input field, records their memories verbally using the voice recording function, and uploads related photos and images.

[1215] Input: Text data, audio data, image data

[1216] Output: Collected memory data

[1217] Step 3:

[1218] The device collects and centrally manages user input data (text, voice, images). Next, the device sends the collected data to a server.

[1219] Specifically, the data is organized within the device and then uploaded to the server via the network.

[1220] Input: User-submitted memory data

[1221] Output: Dataset stored on the server

[1222] Step 4:

[1223] The server analyzes the collected text data using natural language processing techniques.

[1224] In terms of specific operations, the server calls a natural language processing API to extract important keywords and phrases from the text data.

[1225] Input: Text data

[1226] Output: Extracted keywords and phrases

[1227] Step 5:

[1228] The server transcribes the audio data using speech recognition technology and converts it into text data. Next, this text data is analyzed using natural language processing technology.

[1229] Specifically, the server utilizes a speech recognition API to convert speech into text, which is then automatically analyzed.

[1230] Input: Audio data

[1231] Output: Analyzed transcript text data

[1232] Step 6:

[1233] The server analyzes the image data using image recognition technology.

[1234] In terms of specific operations, the server calls an image recognition API to extract key elements and scenes from the image data.

[1235] Input: Image data

[1236] Output: Extracted key elements and scenes

[1237] Step 7:

[1238] The server automatically generates scenarios based on the analyzed data.

[1239] Specifically, a scenario generation algorithm is applied to generate stories using a template-based approach.

[1240] Input: Information from extracted text, audio, and images

[1241] Output: Generated scenario

[1242] Step 8:

[1243] The server creates video content based on the generated scenario.

[1244] Specifically, this involves creating storyboards according to a scenario and then producing short films or videos based on those storyboards. Video editing tools are used for this purpose.

[1245] Input: Generated scenario

[1246] Output: Short film or video content

[1247] Step 9:

[1248] The server develops interactive games based on generated scenarios.

[1249] Specifically, this involves using a game development platform to create interactive games that users can experience through actions and choices.

[1250] Input: Generated scenario

[1251] Output: Interactive game content

[1252] Step 10:

[1253] The device provides generated video and game content in a VR environment.

[1254] Specifically, the device works in conjunction with a VR headset to stream content and deliver it to the user.

[1255] Input: Video content, game content

[1256] Output: User experience in a VR environment

[1257] Through these steps, this system enables users to relive, save, and pass on their precious memories to future generations in high quality and with realism.

[1258] (Application Example 1)

[1259] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1260] Traditionally, there have been limited means to concretely recreate users' memories, making it difficult to preserve individual memories as narratives or images. Furthermore, while methods existed to collect and analyze user memories, there was no way to recreate them in a physical customer experience environment. As a result, users lacked opportunities to easily re-experience their memories in a high-quality format.

[1261] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[1262] In this invention, the server includes means for collecting one or more data from a user, such as text, voice, and images; means for performing natural language processing and image processing to analyze the collected data; means for automatically generating scenarios based on the analyzed data; means for converting the generated scenarios into video content and interactive game content; and means for providing the video content and game content in a physical customer experience environment. This enables users to re-experience their cherished memories in high quality in a physical environment such as a retail store.

[1263] A "user" is an individual or group that uses the system to provide data about memories and receives the recreated content of those memories.

[1264] "Text" refers to data entered by the user, representing the user's memories and information as written text.

[1265] "Voice" refers to data provided by users through speech, which can be converted into text data using speech recognition technology, representing the user's memories and information.

[1266] "Images" refer to photographs and drawings uploaded by users, and are data that includes visual records of memories.

[1267] "Natural language processing" is a technology that analyzes text and audio data collected from users to extract important keywords and phrases.

[1268] "Image analysis" is a technology that analyzes image data collected from users to extract key elements and scenes within the image.

[1269] A "scenario" is a story or narrative generated based on collected and analyzed data, and serves as the basis for video content and game content.

[1270] A "generative AI model" is a type of artificial intelligence that automatically generates stories based on user data, and in particular, it is a model that uses natural language generation technology.

[1271] "Video content" refers to digital content in the form of videos or movies created based on a generated scenario.

[1272] "Interactive game content" refers to digital games based on generated scenarios that users can experience through actions and choices.

[1273] A "physical customer experience environment" refers to a physical store or other physical space where users experience reproduced content.

[1274] This invention is a system that concretely recreates users' memories and provides that recreated content in a physical customer experience environment. The details are described below.

[1275] Collection of user input

[1276] First, users use a device to input data about their memories. The device includes a dedicated application, allowing users to write their memories in a text input field. Users can also provide audio data by speaking using the voice recording function. Furthermore, users can upload related photos and images. This data is collected centrally and sent to a server.

[1277] Data analysis

[1278] The server analyzes text data entered by users using natural language processing techniques to extract important keywords and phrases. Specifically, the text data is analyzed using natural language processing. The server also performs speech recognition techniques to transcribe audio data and convert it into text data. Furthermore, the server analyzes image data and uses image recognition techniques to extract key elements and scenes within the images.

[1279] Scenario generation

[1280] Based on the collected and analyzed data, the server automatically generates scenarios. This uses a generative AI model. The server employs a template-based scenario generation method, for example, by applying keywords such as "travel," "family," and "specific place names" to templates and building stories based on them.

[1281] Generation of video and game content

[1282] After the scenario is generated, the server converts it into video content and interactive game content. For video content, short films and videos are generated based on the scenario. For interactive game content, a game is created that includes elements that users can experience through actions and choices.

[1283] Providing an experience

[1284] The generated video and game content will be delivered in a physical customer experience environment accessible to users. For example, users can experience the content on the spot using tablet devices or robots in a physical store.

[1285] Specific example

[1286] As a concrete example, consider a scenario where a user inputs "memories of a family trip to Hokkaido" into the system. The user uploads text such as "Hokkaido trip," an audio recording of what they said at the time, and photos taken during that trip. The server analyzes this data, extracting keywords such as "family," "trip," and "Hokkaido," and obtaining information from the photos such as "snowy landscape" and "family group photo." Based on this, the server generates a scenario themed around "a family trip to Hokkaido," and then creates a short film and an interactive game based on that scenario.

[1287] Users can use terminals in physical stores to watch movies featuring family members enjoying the snowy landscapes of Hokkaido, and participate in interactive games where they can collect items during their trip.

[1288] Example of a prompt

[1289] For example, when a user is recounting a memory, the following might be used as a prompt:

[1290] "I remember a trip I took to Hokkaido with my family. We enjoyed a warm hot spring bath while admiring the beautiful snowy scenery."

[1291] Based on this prompt, the generative AI model can generate a scenario and provide the user with high-quality content that recreates their memories.

[1292] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[1293] Step 1:

[1294] Users use their devices to input one or more data related to their memories: text, audio, images, and so on. This includes typing into text fields, recording audio using a microphone, and uploading photos and images. The entered data is sent to the server.

[1295] Input: Text, audio, images

[1296] Output: Data to send to the server

[1297] Step 2:

[1298] The server analyzes the received text data using natural language processing techniques to extract important keywords and phrases. Specifically, it tokenizes the text data, performs morphological analysis, and extracts meaningful words and phrases.

[1299] Input: Text data

[1300] Output: Extracted keywords and phrases

[1301] Step 3:

[1302] The server performs speech recognition technology to transcribe the audio data and convert it into text data. The converted text data is similarly analyzed using natural language processing technology to extract keywords and phrases.

[1303] Input: Audio data

[1304] Output: Text data, extracted keywords and phrases

[1305] Step 4:

[1306] The server analyzes the received image data and uses image recognition techniques to extract key elements and scenes from the image. Image analysis employs methods such as object detection and scene classification.

[1307] Input: Image data

[1308] Output: Extracted key elements and scenes

[1309] Step 5:

[1310] The server automatically generates scenarios using a generative AI model based on extracted keywords, phrases, key elements, and scenes. The generative AI model generates a story based on the input prompt sentences.

[1311] Input: Keywords, phrases, key elements or scenes

[1312] Output: Generated scenario

[1313] Step 6:

[1314] The server converts the generated scenarios into video content and interactive game content. For video content, short films and videos are generated based on the scenarios, while interactive games include elements that users can experience through actions and choices.

[1315] Input: Generated scenario

[1316] Output: Video content, interactive game content

[1317] Step 7:

[1318] The generated video and game content will be provided to users to experience on the spot using terminals and robots within physical stores.

[1319] Input: Video content, interactive game content

[1320] Output: Providing experiences in a physical customer experience environment

[1321] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[1322] The system of the present invention concretely recreates the user's memories and converts them into text, images, and interactive game formats. Furthermore, by combining this with an emotion engine that recognizes the user's emotions, it provides a more immersive experience. The following are specific embodiments of this system.

[1323] Collection of user input

[1324] First, the user launches the application using their device and enters data about their memories. Users can write episodes in a text field. They can also provide audio data by speaking about their memories using the voice recording function. Furthermore, they can upload photos and images from that time. This data is collected centrally and sent to the next processing step.

[1325] Data analysis and transformation

[1326] The collected data is analyzed by a server. The server analyzes the text entered by the user using natural language processing technology to extract important keywords and phrases. Next, the server analyzes the audio data and transcribes it using speech recognition technology. The server also analyzes the image data and extracts the main elements and scenes within the images.

[1327] Recognition of emotions

[1328] Next, the server uses an emotion engine to recognize the user's emotions based on the collected audio and text data. From the audio data, it analyzes voice tone, pitch, and volume, and from the text data, it analyzes context and word choice to detect what emotions the user is feeling.

[1329] Scenario generation

[1330] Based on the collected and analyzed data and the output of the emotion engine, the server automatically generates scenarios. Specifically, it employs a template-based scenario generation method to create stories that reflect emotional elements. For example, if a user expresses the emotion of "happiness," a scenario reflecting that emotion will be created.

[1331] Generation of video and game content

[1332] Based on the generated scenario, the server produces video content and interactive game content. For video content, storyboards are created based on the scenario and then rendered as short films or videos. Character expressions and attitudes are also adjusted based on the output of the emotion engine. In game content, interactive elements are created in which the scenario changes according to the user's emotions.

[1333] Providing an experience

[1334] The generated video and game content is delivered through the device in a virtual reality (VR) environment. The device works in conjunction with a VR headset to stream the content, providing users with a realistic experience. This allows users to relive their memories, including emotional elements.

[1335] Specific example

[1336] As a concrete example, consider a scenario where a user inputs "memories of a family trip to Hokkaido when they were 10 years old" into the system. The user uploads text containing "trip to Hokkaido," an audio recording of their conversation at the time, and photos taken during that trip. The server analyzes this data, extracting keywords such as "family," "trip," and "Hokkaido," and obtaining information from the photos such as "snowy landscape" and "family group photo." Furthermore, using an emotion engine, the system detects whether the user was feeling "happy" while recounting this memory.

[1337] Based on this, the server generates a scenario themed around a "family trip to Hokkaido," reflecting the emotion of "happiness." For example, the scenario will include many scenes where the whole family is smiling and enjoying themselves. A short film and an interactive game are then created based on this scenario.

[1338] The user wears a VR headset and watches a short film of a family enjoying a snowball fight in a park. In the game, the user can also experience a task of collecting items with their family, all while smiling. This system allows users to experience their cherished memories in a high-quality, realistic form, including the emotions they evoke.

[1339] The following describes the processing flow.

[1340] Step 1:

[1341] Users launch the application using their devices and input data related to their memories. Specifically, they write episodes in text fields, record audio data using the voice recording function, and upload related photos and images. This data is centrally collected on the device and sent to the next processing step.

[1342] Step 2:

[1343] The device sends the collected data to the server. The server receives the text data, tokenizes it using natural language processing techniques, and extracts important keywords and phrases. For example, keywords such as "family," "travel," and "Hokkaido" are extracted.

[1344] Step 3:

[1345] The server analyzes the received audio data. It uses speech recognition technology to convert the audio into text, and then analyzes the resulting text using natural language processing technology to extract important information.

[1346] Step 4:

[1347] The server analyzes the image data. Using image recognition technology, it extracts key elements and scenes from the image. For example, information such as "snowy landscape" or "family photo" can be obtained from a photograph.

[1348] Step 5:

[1349] The server uses an emotion engine to recognize the user's emotions based on collected audio and text data. From the audio data, it analyzes voice tone, pitch, and volume, and from the text data, it analyzes context and word choice to detect what emotions the user is feeling.

[1350] Step 6:

[1351] The server generates scenarios based on the analyzed data and the output of the emotion engine. Using a template-based scenario generation method, it constructs a story that reflects the emotions expressed by the user. For example, if the user expresses the emotion of "happiness," a scenario reflecting that emotion will be created.

[1352] Step 7:

[1353] The server produces video content using the generated scenario. It creates storyboards based on the scenario and then renders them as short films or videos. Character expressions and attitudes are also adjusted based on the output of the emotion engine.

[1354] Step 8:

[1355] The server creates interactive game content based on the generated scenario. The game is designed to include interactive elements where the scenario changes in response to the user's emotions.

[1356] Step 9:

[1357] The user wears a VR headset using a device. The device plays video content streamed from a server, providing the user with a realistic visual experience. By watching short films, the user can relive their memories with a strong sense of presence.

[1358] Step 10:

[1359] The device runs game content in a VR environment, providing users with an interactive experience. Users can explore a virtual Hokkaido and enjoy a fun time as if they were traveling with their family.

[1360] Through the steps described above, the system of the present invention enables users to transform their memories into a high-quality, realistic form, and to save and share them. Furthermore, the introduction of an emotion engine allows for the provision of an immersive experience that reflects the user's emotions.

[1361] (Example 2)

[1362] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1363] Modern users have a need to relive their memories in a richer way, but this is difficult to achieve with current technology. Existing systems are particularly insufficient when a concrete recreation, including the emotions associated with those memories, is required. Furthermore, there is the challenge of effectively analyzing user input data and converting it into interactive content that reflects those emotions.

[1364] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for collecting one or more data from the user, such as text, voice, and images; means for performing natural language processing and image processing to analyze the collected data; means for recognizing the user's emotions based on the collected voice and text data; means for automatically generating a scenario based on the analyzed data; means for converting the generated scenario into video content and interactive game content; and means for providing the video content and game content in a virtual reality environment. This makes it possible for the user to re-experience their memories in a high-quality, realistic way, including their emotions.

[1365] A "user" is the individual who uses this system to input and re-experience memories.

[1366] "Text" refers to the written data entered by the user.

[1367] "Voice" refers to the sound data that is input when a user speaks.

[1368] "Images" refer to visual data such as photographs and drawings taken by users.

[1369] "Data collection means" refers to means of collecting one or more of the following data from users: text, audio, images, etc.

[1370] "Natural language processing" refers to the technology of analyzing text data to extract meaningful information and keywords.

[1371] "Image analysis" refers to the technique of analyzing image data to extract key elements and scenes.

[1372] "Speech recognition technology" refers to technology used to convert speech data into text data.

[1373] "Emotion recognition means" refers to means for recognizing a user's emotions based on collected audio and text data.

[1374] "Automatic scenario generation method" refers to a method for automatically generating scenarios based on analyzed data.

[1375] A "template" refers to a fixed format or pattern used in automatic scenario generation.

[1376] "Video content conversion means" refers to the means for converting a generated scenario into video content.

[1377] "Interactive game content conversion means" refers to a means for converting a generated scenario into interactive game content.

[1378] "Virtual reality environment provision means" refers to means for providing the aforementioned video content and game content in a virtual reality environment.

[1379] This invention relates to a system that concretely recreates a user's memories and converts them into text, video, or interactive game formats. Furthermore, this system incorporates an emotion engine that recognizes the user's emotions to provide a more immersive experience.

[1380] First, the user launches the application using their device. The user utilizes the device's built-in text input, voice recording, and image upload functions to input data related to their memories. Specifically, the user can write episodes in the text field, provide audio data using the voice recording function, and upload photos or images from that time. This data is then sent from the device to the server.

[1381] Next, the server analyzes the collected data. For text data, it uses Python's natural language processing libraries (such as NLTK and spaCy) to analyze it and extract important keywords and phrases. For audio data, it uses speech recognition technology (such as the Google Speech-to-Text API) to transcribe it. For image data, it uses image analysis libraries (such as OpenCV) to extract key elements and scenes.

[1382] Based on these analysis results, the server uses an emotion engine to recognize the user's emotions. The server analyzes voice tone, pitch, and volume from audio data, and context and word choice from text data to detect what emotions the user is experiencing. IBM Watson Tone Analyzer and Microsoft Azure Text Analytics API are used for this emotion recognition.

[1383] Based on the analyzed data and the results of emotion recognition, the server automatically generates scenarios using a template-based scenario generation method. For example, if the emotion engine detects that the user is feeling "happy," a scenario reflecting that emotion will be generated.

[1384] Based on the generated scenario, the server produces video content and interactive game content. The video content is rendered as short films or videos using video production tools (such as Adobe Premiere Pro and Final Cut Pro). Character expressions and attitudes are also adjusted based on the output of the emotion engine. The game content is created using a game engine (such as Unity or Unreal Engine) to produce a game that includes interactive elements where the scenario changes in response to emotions.

[1385] The final generated video and game content is delivered through the device in a virtual reality (VR) environment. The device works in conjunction with VR headsets (such as Oculus Rift and HTC Vive) to provide users with a realistic experience. This allows users to re-experience their memories, including emotional elements.

[1386] To give a concrete example, if a user inputs "memories of a family trip to Hokkaido when they were 10 years old" into the system, the user would type "trip to Hokkaido" into their terminal and upload audio recordings and photos from that time. The server analyzes this data, extracting keywords such as "family," "trip," and "Hokkaido," and obtaining information from the photos such as "snowy landscape" and "family group photo." Using an emotion engine, it detects the "happy" emotion the user felt while recounting this memory. Based on this, the server generates a scenario themed around "a family trip to Hokkaido" that reflects the "happy" emotion. A short film and an interactive game are created based on the scenario, and the user can experience them through a VR headset.

[1387] Examples of prompts to input into a generative AI model are as follows:

[1388] "Please write about a happy memory from a family trip to Hokkaido when you were 10 years old. Please include photos and audio recordings. Based on this data, please generate a scenario using an emotion-reflecting system and create video content or a game."

[1389] The above describes the form for carrying out the invention. This system allows users to transform their cherished memories into a high-quality, realistic form, including emotions, and re-experience them.

[1390] The flow of the specific processing in Example 2 will be explained using Figure 13.

[1391] Step 1: Collecting User Input

[1392] The user launches the application using their device and inputs data related to their memories. The specific actions are as follows: The user writes episodes in a text field (input: text data), provides audio data using the voice recording function (input: audio data), and uploads photos and images (input: image data). This data is sent centrally from the device to the server (output: transmission of integrated data).

[1393] Step 2: Analysis of the month

[1394] The server receives the transmitted data and begins analysis. First, it uses Python's natural language processing libraries (e.g., NLTK, spaCy) on the text data to extract important keywords and phrases (input: text data, output: important keywords). Next, it uses speech recognition technology such as the Google Speech-to-Text API on the audio data to transcribe it (input: audio data, output: text data). Finally, it uses image analysis libraries such as OpenCV on the image data to extract key elements and scenes (input: image data, output: important elements).

[1395] Step 3: Recognizing Emotions

[1396] The server recognizes emotions based on the analyzed data. For audio data, it analyzes voice tone, pitch, and volume (input: audio data) to identify the user's emotions (output: emotion data). For text data, it uses IBM Watson Tone Analyzer or Microsoft Azure Text Analytics API to analyze context and word choice to identify emotions (input: text data, output: emotion data).

[1397] Step 4: Scenario Generation

[1398] The server automatically generates template-based scenarios based on the analysis data and emotion recognition results. Specifically, it uses a scenario generation algorithm to create stories that reflect emotional elements (input: analysis data and emotion data, output: scenario). For example, if the user indicates the emotion of "happiness," a scenario reflecting that emotion will be automatically generated.

[1399] Step 5: Generating video and game content

[1400] The server produces video content and interactive game content based on the generated scenario. For video content, it uses video production tools (e.g., Adobe Premiere Pro, Final Cut Pro) to render short films and videos based on the scenario (input: scenario, output: video content). For game content, it uses a game engine (e.g., Unity, Unreal Engine) to create interactive games that change in response to emotions (input: scenario, output: game content).

[1401] Step 6: Providing the experience

[1402] The device delivers generated video and game content in a virtual reality (VR) environment. Users wear a VR headset (e.g., Oculus Rift, HTC Vive) and experience this content (input: video and game content, output: VR experience). This allows users to re-experience their memories in a high-quality, realistic form, including the emotions they evoke.

[1403] (Application Example 2)

[1404] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1405] Traditionally, systems designed to recreate users' memories have simply displayed collected data, making it difficult to reproduce users' emotions and subtle nuances. Furthermore, it has been challenging to deliver individual user memories in real time and provide an immersive experience. Therefore, there is a need for a system that can achieve a higher level of realism and individual personalization.

[1406] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for collecting one or more data from the user, such as text, voice, and images; means for performing natural language processing and image processing to analyze the collected data; means for automatically generating a scenario based on the analyzed data; means for converting the generated scenario into video content and interactive game content; and means for streaming the scenario-based content in real time via a terminal such as a smartphone. This makes it possible to reproduce the user's emotions and memories with a high sense of realism and to provide a real-time, individually personalized experience.

[1407] A "user" refers to an individual who uses the system to input and analyze their own memories.

[1408] "Text" refers to the written data entered by the user.

[1409] "Audio" refers to audio data that records what the user has said.

[1410] "Images" refer to visual data such as photos and drawings uploaded by users.

[1411] "Natural language processing" refers to the process of analyzing collected text and audio data to extract keywords and context.

[1412] "Image analysis" refers to the process of analyzing collected image data and extracting key elements and scenes from the image.

[1413] "Automatic scenario generation" refers to a method of automatically creating scenarios that reflect user emotions using templates, based on analyzed data.

[1414] "Video content" refers to visual content rendered based on a generated scenario.

[1415] "Interactive game content" refers to game-style content that progresses through interaction with the user based on a generated scenario.

[1416] A "virtual reality environment" refers to technology that allows users to experience things in a virtual space as if they were reality.

[1417] "Smartphones and other devices" refers to portable information devices used by users to access memory-recreation streaming services.

[1418] "Streaming distribution" refers to the technology that provides generated video and game content to users in real time via a network.

[1419] "Sentiment analysis" refers to the process of analyzing the tone, pitch, and other characteristics of text and audio data collected from users to estimate their emotions.

[1420] The system of this invention reproduces user-inputted memories in real time and delivers them as personalized content that takes emotions into account. This system analyzes data collected from users, generates scenarios, and provides video content and interactive game content based on those scenarios.

[1421] Hardware and software to be used

[1422] 1. Smartphone: The primary device used by users to input text, voice, and image data and receive the results.

[1423] 2. Server / Cloud: A centralized processing system for data analysis, sentiment recognition, scenario generation, and content rendering.

[1424] 3. Natural Language Processing (NLP) libraries (e.g., SpaCy, NLTK): These libraries analyze the collected text data and extract important keywords and phrases.

[1425] 4. Speech recognition library (e.g., Google Cloud Speech-to-Text): Converts audio data into text data and transcribes the audio into text.

[1426] 5. Emotion recognition engine (e.g., IBM Watson Tone Analyzer): Analyzes the emotions in text and audio data.

[1427] 6. Video rendering engine (e.g., Unity): Generates video content and interactive game content based on the scenario.

[1428] Data processing and data calculation workflow

[1429] 1. Collecting user input:

[1430] Users input text, voice, and images using their smartphones. This data is sent to a server and collected centrally.

[1431] 2. Data Analysis:

[1432] The server uses natural language processing technology to analyze text data and extract important keywords and phrases. It also uses speech recognition technology to transcribe audio data and image analysis technology to analyze key elements of image data.

[1433] 3. Recognition of emotions:

[1434] The server uses an emotion recognition engine to analyze the user's emotions from text and voice data. This identifies the tone, pitch, volume, and context of the emotion.

[1435] 4. Scenario generation:

[1436] The server generates emotionally reflective scenarios based on the analyzed data and emotion recognition results. A template-based scenario generation method is employed to create personalized stories.

[1437] 5. Content generation:

[1438] The server creates video content and interactive game content based on the generated scenario. The generated content reflects emotional elements, and the characters' expressions and attitudes are adjusted accordingly.

[1439] 6. Providing experiences:

[1440] Ultimately, the generated video content and interactive game content are streamed in real time via devices such as smartphones. By integrating VR headsets and other devices, a more immersive experience can be provided.

[1441] Specific example

[1442] For example, if a user inputs "memories of a family trip to Hokkaido when they were 10 years old" in text, audio, or image format, the server analyzes this and extracts keywords such as "family," "trip," and "Hokkaido." In addition, it uses an emotion recognition engine to detect that the user is feeling "happy." Based on this information, the server generates a scenario themed around "a family trip to Hokkaido" that reflects the user's emotions.

[1443] Example of a prompt

[1444] Theme: Family trip to Hokkaido

[1445] Emotion: Happy

[1446] Keywords: family, travel, snowscape

[1447] This system allows users to recreate their memories, including the emotions they evoke, and experience them with a high degree of realism.

[1448] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[1449] Step 1:

[1450] The user launches the application using their smartphone and inputs one or more text, audio, or image data related to their memories. For audio input, they use the voice recording function and upload photos or images taken at that time. This data is sent from the device to the server. The input data includes detailed memories (e.g., "Memories of a family trip to Hokkaido when I was 10 years old").

[1451] Step 2:

[1452] The server analyzes the collected text data using natural language processing techniques (e.g., SpaCy, NLTK) to extract important keywords and phrases. Specifically, it extracts keywords such as "family," "travel," and "Hokkaido," and understands the context of the text data. The input is text data, and the output is keywords and contextual information.

[1453] Step 3:

[1454] The server uses speech recognition technology (e.g., Google Cloud Speech-to-Text) to convert audio data into text data and perform transcription. Specifically, it converts what the user says into text format and performs similar natural language processing. The input is audio data, and the output is transcribed text data.

[1455] Step 4:

[1456] The server uses image analysis technology to analyze uploaded image data and extract key elements and scenes. Specifically, it obtains information such as "snowy landscape" or "family group photo" from the image. The input is image data, and the output is information about key elements and scenes.

[1457] Step 5:

[1458] The server analyzes the user's emotions using an emotion recognition engine (e.g., IBM Watson Tone Analyzer) based on the parsed text and audio data. Specifically, it analyzes the context of the text, the tone of the voice, and the pitch to detect whether the user is experiencing emotions such as "happy." The input is the parsed text and audio data, and the output is emotion data.

[1459] Step 6:

[1460] The server automatically generates a story using a template-based scenario generation method, based on extracted keywords, key elements, scene information, and emotion data. Specifically, it creates a scenario themed around a "family trip to Hokkaido" that reflects the emotion of "happiness." The input consists of keywords, key elements, scene information, and emotion data, and the output is the generated scenario.

[1461] Step 7:

[1462] The server produces video content and interactive game content based on the generated scenario. Specifically, it uses a video rendering engine such as Unity to create short films and interactive games while adjusting character expressions and attitudes based on emotional data. The input is the generated scenario, and the output is the video content and game content.

[1463] Step 8:

[1464] The system allows users to stream generated video content and interactive game content in real time using devices such as smartphones. Specifically, it works in conjunction with VR headsets to provide users with a high level of immersion. The input is video content and game content, and the output is the user's experience.

[1465] Through these steps, users can transform their memories into a high-quality, realistic form that includes emotions, and experience them.

[1466] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[1467] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1468] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[1469] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1470] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[1471] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[1472] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[1473] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[1474] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[1475] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[1476] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[1477] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[1478] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[1479] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1480] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[1481] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[1482] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[1483] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[1484] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[1485] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[1486] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.

[1487] The following is further disclosed regarding the embodiments described above.

[1488] (Claim 1)

[1489] Means for collecting one or more data from users, including text, audio, and images,

[1490] Means for performing natural language processing and image processing to analyze collected data,

[1491] A method for automatically generating scenarios based on analyzed data,

[1492] A means for converting the generated scenario into video content and interactive game content,

[1493] A means for providing the aforementioned video content and game content in a virtual reality environment,

[1494] A system that includes this.

[1495] (Claim 2)

[1496] The system according to claim 1, wherein the natural language processing means includes speech recognition technology that converts speech data into text data.

[1497] (Claim 3)

[1498] The system according to claim 1, wherein the scenario generation means includes means for automatically generating a story based on user data using a template.

[1499] "Example 1"

[1500] (Claim 1)

[1501] Means for collecting one or more data from users, including text, audio, and images,

[1502] Means for performing natural language processing and image processing to analyze collected data,

[1503] A method for automatically generating scenarios based on analyzed data,

[1504] A means for converting the generated scenario into video content and interactive game content,

[1505] A means for providing the aforementioned video content and game content in a virtual reality environment,

[1506] A means of centrally managing the collected data and sending it to a server,

[1507] A means for analyzing data using natural language processing technology, speech recognition technology, and image recognition technology,

[1508] A method for creating storyboards using a template-based approach and generating video and game content based on them,

[1509] A system that includes this.

[1510] (Claim 2)

[1511] The system according to claim 1, wherein the natural language processing means includes speech recognition technology that converts speech data into text data.

[1512] (Claim 3)

[1513] The system according to claim 1, wherein the scenario generation means includes means for automatically generating a story based on user data using a template.

[1514] "Application Example 1"

[1515] (Claim 1)

[1516] Means for collecting one or more data from users, including text, audio, and images,

[1517] Means for performing natural language processing and image processing to analyze collected data,

[1518] A method for automatically generating scenarios based on analyzed data,

[1519] A means for converting the generated scenario into video content and interactive game content,

[1520] A means for providing the aforementioned video content and game content in a physical customer experience environment,

[1521] A system that includes this.

[1522] (Claim 2)

[1523] The system according to claim 1, wherein the natural language processing means includes speech recognition technology that converts speech data into text data.

[1524] (Claim 3)

[1525] The system according to claim 1, wherein the scenario generation means includes means for automatically generating a story based on user data using a generation AI model.

[1526] "Example 2 of combining an emotion engine"

[1527] (Claim 1)

[1528] Means for collecting one or more data from users, including text, audio, and images,

[1529] Means for performing natural language processing and image processing to analyze collected data,

[1530] A means of recognizing the user's emotions based on collected audio and text data,

[1531] A method for automatically generating scenarios based on analyzed data,

[1532] A means for converting the generated scenario into video content and interactive game content,

[1533] A means for providing the aforementioned video content and game content in a virtual reality environment,

[1534] A system that includes this.

[1535] (Claim 2)

[1536] The system according to claim 1, wherein the natural language processing means includes speech recognition technology that converts speech data into text data.

[1537] (Claim 3)

[1538] The system according to claim 1, wherein the scenario generation means includes means for automatically generating a story based on user data using a template.

[1539] "Application example 2 when combining with an emotional engine"

[1540] (Claim 1)

[1541] Means for collecting one or more data from users, including text, audio, and images,

[1542] Means for performing natural language processing and image processing to analyze collected data,

[1543] A method for automatically generating scenarios based on analyzed data,

[1544] A means for converting the generated scenario into video content and interactive game content,

[1545] A means for providing the aforementioned video content and game content in a virtual reality environment,

[1546] A means of streaming scenario-based content in real time via devices such as smartphones,

[1547] A system that includes this.

[1548] (Claim 2)

[1549] The system according to claim 1, wherein the natural language processing means includes speech recognition technology that converts speech data into text data.

[1550] (Claim 3)

[1551] The system according to claim 1, wherein the scenario generation means includes means for analyzing the user's emotions and automatically generating a story based on the user's data using a template based on the analysis results. [Explanation of symbols]

[1552] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. Means for collecting one or more data from users, including text, audio, and images, Means for performing natural language processing and image processing to analyze collected data, A method for automatically generating scenarios based on analyzed data, A means for converting the generated scenario into video content and interactive game content, A means for providing the aforementioned video content and game content in a virtual reality environment, A system that includes this.

2. The system according to claim 1, wherein the natural language processing means includes speech recognition technology that converts speech data into text data.

3. The system according to claim 1, wherein the scenario generation means includes means for automatically generating a story based on user data using a template.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A