system
The system addresses the challenge of providing realistic and customized travel experiences for individuals unable to travel by using AI to generate personalized scenarios and videos, integrating virtual shopping, and offering detailed information on demand, thus enhancing user experience.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-08-26
- Publication Date
- 2026-03-10
AI Technical Summary
Individuals unable to travel due to physical limitations, illness, or external factors such as pandemics and natural disasters face psychological stress and loneliness, and existing systems fail to provide realistic and customized travel experiences.
A system that allows users to select a travel destination, generates personalized scenarios and videos, provides detailed information on demand, and integrates virtual shopping, using AI for scenario and video generation, and includes user authentication and information storage.
Enables a realistic and customized travel experience from home or hospital, addressing physical and external travel restrictions, and allows for virtual shopping, providing a fulfilling experience tailored to individual needs.
Smart Images

Figure 2026041213000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] People who cannot travel for physical reasons, such as the elderly, people with disabilities, and those hospitalized due to illness, are unable to experience the local experience and enjoy the joys of travel. Travel may also be restricted by external factors such as pandemics and natural disasters. These issues can cause many people to feel psychological stress and loneliness due to being unable to go out. In response to these circumstances, there is a need to provide a way to provide a realistic travel experience from home or a hospital room. [Means for solving the problem]
[0005] The system of the present invention solves these problems by providing a system including a selection means for a user to select a travel destination they would like to visit, a text generation means for generating a scenario related to the selected travel destination, a video generation means for generating a video of the travel destination based on the generated scenario, a provision means for providing the generated scenario and video to the user, a request means for the user to request detailed information about a specific location based on the scenario or video, and a text generation means for generating the detailed information based on the request.The system also includes a registration means and a storage means for registering and saving basic information such as the user's health condition and travel preferences, an authentication means for the user to log in to the system, and a page provision means for providing a page dedicated to the authenticated user, thereby providing a travel experience customized to the individual needs of the user.
[0006] "User" refers to an individual who uses the system to simulate a travel destination.
[0007] "Selection means" refers to an interface or function that allows a user to select a desired travel destination on the system.
[0008] "Text generation means" refers to a function or system that automatically generates a travel scenario and detailed explanation based on information about the selected travel destination.
[0009] "Video generation means" refers to a function or system for generating videos of travel destinations based on the travel scenario generated by the text generation means.
[0010] "Providing means" refers to an interface or function for displaying and providing the generated scenario and video to the user.
[0011] "Request means" refers to an interface or function for requesting specific locations or detailed information from the scenario or video the user is viewing.
[0012] "Registration means" refers to an interface or function that allows a user to register basic information (health status, travel preferences, etc.) in the system.
[0013] "Storage means" refers to the function for storing and managing basic information of registered users in a database, etc.
[0014] "Authentication means" refers to a function for authenticating a user when accessing a system and confirming that the user is a legitimate user.
[0015] "Page provision means" refers to an interface or function for displaying and providing a page dedicated to an authenticated user.
[0016] "Detailed information" refers to additional information about a specific location or item requested by a user using a request means. [Brief explanation of the drawings]
[0017] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8]FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram illustrating a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0018] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0019] First, the terms used in the following description will be explained.
[0020] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0021] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0022] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0023] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0024] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0025] [First embodiment]
[0026] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0027] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0028] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0029] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0030] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0031] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0032] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0033] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0034] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0035] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0036] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0037] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0038] The present invention provides a system that enables people who cannot travel to experience a simulated trip from their homes, hospital rooms, etc. This system generates realistic travel scenarios and images based on the travel destinations selected by the user and provides them to the user.
[0039] Overall system configuration
[0040] The system mainly consists of the following components:
[0041] 1. User terminal: A device (PC, tablet, smartphone, etc.) on which the user selects a travel destination and views scenarios and footage.
[0042] 2. Server: A back-end system that generates, manages, and provides user information, travel scenarios, and videos.
[0043] 3. Text generation AI: AI that generates scenarios based on information about the selected travel destination.
[0044] 4. Video generation AI: AI that generates video based on a generated scenario.
[0045] 5. Database: A storage system that stores user information, travel scenarios, and videos.
[0046] System processing: Example
[0047] Initial settings and user information registration
[0048] 1. A user accesses the system using a user terminal and creates a new account.
[0049] Enter your name, age, health status, travel preferences, etc., and press the register button.
[0050] 2. The server receives the user information, validates it, and saves it in the database.
[0051] Selecting a travel destination
[0052] 3. After logging in, the user accesses the travel destination selection screen.
[0053] Choose where you want to go from a world map or a list of travel destinations.
[0054] 4. The server receives the selected destination information and records the user's selection.
[0055] Generate a scenario
[0056] 5. The server sends the selected travel destination information to the text generation AI.
[0057] The text generation AI uses that information to generate travel scenarios, such as the history of the Eiffel Tower or where the Mona Lisa is located.
[0058] 6. The server receives the generated scenario and stores it in the database.
[0059] Video generation
[0060] 7. The server requests the video generation AI to generate a video based on the generated scenario.
[0061] The video generation AI generates 3D models and on-site footage based on the scenario.
[0062] For example, the illuminated Eiffel Tower in Paris and the exhibits at the Louvre Museum.
[0063] Providing a user experience
[0064] 8. The terminal provides the user with a screen that displays the generated scenario and images.
[0065] The user can play the video and read the scenario text.
[0066] For example, you can watch a video of the Eiffel Tower while reading its detailed history.
[0067] Get more information
[0068] 9. When the user selects a specific location or event within the video or scenario, more information is requested.
[0069] 10. The server again asks the text generation AI to generate detailed information based on the request.
[0070] For example, the structure of the Eiffel Tower and stories from its construction.
[0071] 11. The terminal displays the generated details to the user.
[0072] Supplementary information is displayed in real time in sync with the video.
[0073] In this way, users can enjoy a realistic travel experience from their own homes or hospital rooms. This system allows users to enjoy the joy of travel regardless of physical constraints or external factors. Furthermore, the system allows for advanced customization according to the individual needs of each user, providing an optimal experience for each individual user.
[0074] The processing flow will be explained below.
[0075] Step 1:
[0076] A user accesses the system and creates a new account by entering information such as their name, age, health status, and travel preferences into the terminal and pressing the register button.
[0077] Step 2:
[0078] The server receives the input information, validates it, checks for errors, and saves the correct information to the database.
[0079] Step 3:
[0080] The device will display a "Registration Complete" message to the user and redirect them to the login screen.
[0081] Step 4:
[0082] The user enters their user ID and password on the login screen and clicks the login button.
[0083] Step 5:
[0084] The server authenticates the user, checking against information in a database and authorizing if there is a match.
[0085] Step 6:
[0086] The device displays a destination selection screen, providing the user with a world map and a list of destinations.
[0087] Step 7:
[0088] The user selects the desired travel destination and presses the select button.
[0089] Step 8:
[0090] The server receives the user's selection information and generates a request to send details of the selected travel destination to a text generation AI.
[0091] Step 9:
[0092] A text generation AI generates a travel scenario based on the selected travel destination, including information on the history, culture, and key attractions of the tourist destination.
[0093] Step 10:
[0094] The server receives the generated scenario and stores it in a database.
[0095] Step 11:
[0096] The server uses the saved scenario information to send a video generation request to the video generation AI.
[0097] Step 12:
[0098] Based on the scenario, the video generation AI generates footage of the target travel destination, including real-time scenery and 3D models.
[0099] Step 13:
[0100] The server receives the generated video and stores it in a database.
[0101] Step 14:
[0102] The terminal displays a screen on which the generated video and scenario can be viewed by the user.
[0103] Step 15:
[0104] The user plays the video, begins reading the scenario text, and clicks on specific locations or events that interest them.
[0105] Step 16:
[0106] The server receives the user's click information and sends a detailed information generation request back to the text generation AI.
[0107] Step 17:
[0108] Text generation AI generates detailed information, including specific history and unique anecdotes about the location.
[0109] Step 18:
[0110] The server receives the generated details and sends them to the terminal.
[0111] Step 19:
[0112] The device displays detailed information to the user and works in conjunction with the video to enhance the travel experience.
[0113] This series of steps allows users to have a realistic travel experience, as if they were actually visiting the location.
[0114] Example 1
[0115] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0116] In modern society, there are many people who find it difficult to actually travel for various reasons. For example, many people are unable to travel due to health conditions, time constraints, or financial reasons. There is also a need for a way to simulate a trip for elderly people and those who find it difficult to go out due to illness. Furthermore, conventional travel experience systems have difficulty customizing realistic travel scenarios and videos, making it impossible to provide an experience tailored to individual users' needs.
[0117] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0118] In this invention, the server includes a selection means for allowing a user to select a travel destination they wish to go to, a text generation means for generating a scenario related to the selected travel destination, a video generation means for generating a video of the travel destination based on the generated scenario, a provision means for providing the generated scenario and video to the user, a request means for allowing the user to request detailed information about a specific location based on the scenario or video, a text generation means for generating the detailed information based on the request, a display means for displaying the detailed information generated in response to the user's request on a user terminal, and a storage means for registering and saving user information by initial setting. This enables users who have difficulty traveling in person due to health conditions, time constraints, financial reasons, etc. to enjoy a realistic and customized travel experience from the comfort of their own home or hospital room.
[0119] "User" refers to an individual who intends to use the system to have a simulated travel experience.
[0120] "Selection means" refers to an interface or tool that allows a user to select a travel destination they wish to go to.
[0121] "Text generation means" refers to a mechanism that utilizes natural language processing technology to generate scenarios and detailed information about the selected travel destination.
[0122] "Video generation means" refers to the technology and tools used to generate realistic travel destination footage based on the generated scenario.
[0123] "Providing means" refers to the interface or media for providing the generated scenario and video to the user.
[0124] "Request means" refers to an interface or tool that allows a user to request detailed information about a specific location or event based on a scenario or video.
[0125] "Display means" refers to a mechanism for displaying detailed information and images generated in response to a user's request on a user terminal.
[0126] "Storage means" refers to the technology or tools used to register user information in the initial settings and store it in a database, etc.
[0127] MODE FOR CARRYING OUT THE INVENTION
[0128] The present invention relates to a system that allows people who cannot travel to have a simulated real-life travel experience from their homes, hospital rooms, etc. This system generates realistic travel scenarios and images based on travel destinations selected by the user and provides them to the user. The specific configuration and processing flow of the system are described below.
[0129] Overall system configuration
[0130] The system mainly consists of the following components:
[0131] 1. User terminal: A device (PC, tablet, smartphone, etc.) on which the user selects a travel destination and views scenarios and footage.
[0132] 2. Server: A back-end system that generates, manages, and provides user information, travel scenarios, and videos.
[0133] 3. Text generation AI: AI that generates scenarios based on information about the selected travel destination.
[0134] 4. Video generation AI: AI that generates video based on a generated scenario.
[0135] 5. Database: A storage system that stores user information, travel scenarios, and videos.
[0136] System Operation
[0137] Initial settings and user information registration
[0138] A user accesses the system using a user terminal and creates a new account. Specifically, they enter basic information such as name, age, health status, and travel preferences, and then press the register button. At this stage, the server receives the user information, validates the input data, and stores the information in a database. This stored information is used to provide a customized travel experience in the future.
[0139] Selecting a travel destination
[0140] After logging in, the user accesses the travel destination selection screen and selects a destination. The user selects the place they want to go from a world map or a list of travel destinations, and once the selection is complete, the information is sent to the server. The server records this selection information in a database.
[0141] Generate a scenario
[0142] Based on the information about the selected travel destination, the server sends the information to the text generation AI. The text generation AI generates a travel scenario based on this information. For example, it generates information such as "The history of the Eiffel Tower" and "The location where the Mona Lisa is exhibited." The generated scenario is saved in a database by the server.
[0143] Video generation
[0144] The server requests the video generation AI to generate a video based on the generated scenario. The video generation AI generates a 3D model and on-site video in accordance with the scenario. For example, it generates video of the "illuminated Eiffel Tower" or "exhibits at the Louvre Museum." This video is also stored in a database by the server.
[0145] Providing a user experience
[0146] The user device provides a screen that displays the generated scenario and video to the user, allowing the user to experience a simulated trip. For example, a realistic travel experience can be provided, in which a video of the Eiffel Tower is played while a text about its detailed history is read.
[0147] Get more information
[0148] When a user selects a specific location or event within a video or scenario, they can request additional details. Based on this request, the server again asks the text generation AI to generate more information. For example, specific information such as "the structure of the Eiffel Tower" or "episodes from its construction" is generated. The user's device displays this information to the user, providing supplementary information in real time.
[0149] Examples and prompts
[0150] As a concrete example, consider the case where a user selects "Paris" as a travel destination and wants to search for more information about the Eiffel Tower. An example prompt sentence is:
[0151] Generate Scenario prompt:
[0152] Prompt: What is the detailed history of the Eiffel Tower in Paris and what are its popular tourist attractions?
[0153] Video generation prompt:
[0154] Prompt: Generate a video that includes the Eiffel Tower lit up and a museum exhibit.
[0155] summary
[0156] This system allows users to enjoy a realistic travel experience from the comfort of their own home or hospital room. Furthermore, by using generative AI models and prompts, it is possible to customize the system to meet the needs of each individual user, providing a high-quality simulated travel experience.
[0157] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0158] Step 1: Register user information
[0159] A user accesses the system from their terminal and creates a new account. They enter their name, age, health condition, travel preferences, etc., and press the "Register" button. The server receives this information and performs validation (checks the validity of the data). If the input data is correct, the server saves the information in the database. This registers the user information.
[0160] Input: User's name, age, health status, travel preferences
[0161] Data processing / data calculation: User information validation
[0162] Output: Registered user information
[0163] Step 2: Choose your destination
[0164] The user logs in to the system and accesses the travel destination selection screen. The user selects the place they want to go to from a world map or a list of travel destinations and presses the "Select" button. The server receives the selected travel destination information and records it in the database.
[0165] Input: Travel destination selection information
[0166] Data processing / data calculation: Recording selected travel destination information
[0167] Output: Recorded travel destination information
[0168] Step 3: Generate a scenario
[0169] The server retrieves the selected travel destination information from the database and sends the information to the text generation AI based on this. The text generation AI processes the generation prompt and generates a detailed travel scenario. For example, it sends a prompt sentence such as "Prompt: Please tell me the detailed history of the Eiffel Tower in Paris and its popular tourist attractions." The server receives the generated scenario and stores it in the database.
[0170] Input: Selected travel destination information
[0171] Data processing / data calculation: Travel scenario generation
[0172] Output: Generated travel scenarios
[0173] Step 4: Generate footage
[0174] The server requests the video generation AI to generate a video based on the generated scenario. A prompt based on the scenario is sent to the video generation AI, for example, "Prompt: Generate a video that includes the illuminated Eiffel Tower and museum exhibits." The video generation AI generates a 3D model and on-site video based on the scenario, and the server stores the generated video in a database.
[0175] Input: Generated travel scenarios
[0176] Data processing / data calculation: Video generation based on travel scenarios
[0177] Output: Generated video
[0178] Step 5: Delivering the user experience
[0179] The device provides the user with a screen that displays the generated scenarios and videos. The user can enjoy the travel scenarios and videos through the device. For example, the user can play a video of the Eiffel Tower while reading a text about its history.
[0180] Input: Travel scenario and footage
[0181] Data processing / data calculation: Scenario and video display
[0182] Output: Travel scenario and video displayed
[0183] Step 6: Get more information
[0184] The user requests more information about a specific location or event in a video or scenario. For example, if the user wants more information about the structure of the Eiffel Tower, they click on that part. The server receives the request and again asks the text generation AI to generate the detailed information. The text generation AI generates the detailed information based on the specified prompt, and the server sends it to the device. The device then displays the generated detailed information to the user.
[0185] Input: Request for more information
[0186] Data processing / data calculation: generating detailed information based on demand
[0187] Output: Detailed information generated
[0188] In this way, at each processing step, appropriate processing is performed based on the input data, and the results are linked to the next step, providing the user with a realistic and customized travel experience.
[0189] (Application example 1)
[0190] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0191] In recent years, there has been a demand for systems that allow people who cannot travel to simulate travel experiences from their homes, hospital rooms, etc. There is also a need for systems that allow users to enjoy shopping at travel destinations using virtual environments. However, these systems have not yet become widespread, making it difficult for users to simultaneously enjoy realistic travel and shopping experiences. Therefore, the present invention aims to provide a system that integrates simulated travel experiences with virtual shopping, thereby enabling users to have a more realistic and fulfilling experience.
[0192] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0193] In this invention, the server includes a selection means for allowing a user to select a travel destination they wish to go to, a text generation means for generating a scenario related to the selected travel destination, a video generation means for generating a video of the travel destination based on the generated scenario, a provision means for providing the generated scenario and video to the user, a request means for allowing the user to request detailed information about a specific place based on the scenario or video, a text generation means for generating the detailed information based on the request, a virtual shopping means for allowing the user to purchase products in the virtual environment, and a delivery means for delivering the products selected in the virtual environment to the real world. This allows the user to enjoy a virtual trip while shopping from home or a hospital room.
[0194] "User" refers to any individual or entity that uses the System.
[0195] "Destination" refers to a virtual or real geographic location that a user wishes to visit.
[0196] "Selection means" refers to an interface or tool that allows a user to select a travel destination they wish to visit.
[0197] "Text generation means" refers to an artificial intelligence or algorithm that automatically generates a scenario based on specific information.
[0198] "Video generation means" refers to artificial intelligence or a method that automatically creates video based on a generated scenario.
[0199] "Providing means" refers to a technical method or device for displaying or playing the generated scenario and video to the user.
[0200] "Request means" refers to an interface or tool that allows a user to request detailed information about a specific location based on a scenario or video.
[0201] "More Information" refers to additional and specific information about a particular location or event.
[0202] "Virtual shopping vehicle" refers to an interface or tool that allows a user to select and purchase items within a virtual environment.
[0203] "Delivery Method" refers to a method or service for delivering a product selected in a virtual environment to a user in the real world.
[0204] A specific embodiment for carrying out the present invention is described below. This system is designed to provide an integrated experience that allows users to enjoy virtual travel and virtual shopping. The details are described below.
[0205] 1. Overall system configuration
[0206] This system mainly consists of a user terminal, a server, a generative AI model, a database, and a delivery service.
[0207] 1.1 User terminal
[0208] The user terminal is a device that allows users to select travel destinations, view scenarios and videos, and even conduct virtual shopping. Specifically, this applies to smartphones and head-mounted displays.
[0209] 1.2 Server
[0210] The server is a back-end system that generates, manages, and provides user information, travel scenarios, videos, and shopping information. It has the following means:
[0211] 1. Selection means: Provides an interface for the user to select the travel destination they wish to go to.
[0212] 2. Text generation means: Use a generative AI model (e.g., GPT-3 (registered trademark) from OpenAI (registered trademark)).5) to generate travel destination scenarios.
[0213] 3. Video generation method: A generative AI model for generating video based on a scenario.
[0214] 4. Provision means: The generated scenario and video are provided to the user.
[0215] 5. Request means: Provides an interface for users to request more information.
[0216] 6. Virtual Shopping Vehicle: Provides an interface for users to purchase products within a virtual environment.
[0217] 7. Delivery Method: Providing a service for delivering the items selected in the virtual environment in the real world.
[0218] 1.3 Database
[0219] The database is a storage system for storing user information, travel scenarios, videos, and product information.
[0220] 1.4 Delivery Services
[0221] The delivery service arranges for the items purchased by the user through virtual shopping to be delivered in the real world.
[0222] 2. Processing Overview
[0223] 2.1 User Registration and Settings
[0224] A user accesses the system using a user terminal and creates a new account. The server receives the user information, validates it, and stores it in the database.
[0225] 2.2 Choosing a travel destination
[0226] After logging in, users select the places they want to virtually visit from a world map or a list of destinations. The server receives the selected travel destination information and records the user's selection.
[0227] 2.3 Scenario and image generation
[0228] The server transmits information about the selected travel destination to the text generation means, receives the generated scenario, and stores it in a database.The server then transmits the scenario to the video generation means and receives the generated video.
[0229] 2.4 Virtual Shopping
[0230] Users can access the virtual store in the generated video and purchase products, which are then delivered to the real world using a delivery method.
[0231] 3. Examples of concrete examples and prompts
[0232] For example, if the user selects "Paris," the system performs the following process:
[0233] 1. The scenario generation AI generates a scenario that includes information about the Eiffel Tower and the Louvre Museum.
[0234] 2. Image generation AI generates images of actual Paris cityscapes and tourist attractions.
[0235] 3. Users can purchase limited edition keychains and delicious local chocolates at the Eiffel Tower's virtual store.
[0236] Prompt Sentence Examples
[0237] "You're currently on a virtual tour of Paris. See the beautifully lit Eiffel Tower and shop for exclusive keychains and delicious local chocolates at a nearby souvenir shop."
[0238] This system allows users to enjoy virtual travel and shopping from the comfort of their own home or hospital room.
[0239] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0240] Step 1:
[0241] A user accesses the system using a terminal and creates a new account. The user enters basic information such as name, age, and travel preferences, and presses the register button. The basic information is sent from the terminal to the server.
[0242] Input: Basic information you enter into your device, such as your name, age, and travel preferences.
[0243] Data processing: The server validates basic information.
[0244] Output: The server saves the basic information that passes validation to the database.
[0245] Step 2:
[0246] The user accesses the login screen and logs in using the authentication method. The server checks the user's authentication information against the database and provides a dedicated page for the authenticated user.
[0247] Input: User credentials (username, password).
[0248] Data calculation: The server performs authentication by checking against information in the database.
[0249] Output: Authenticated users are served a dedicated page.
[0250] Step 3:
[0251] The user accesses the travel destination selection screen from a dedicated page and selects the desired travel destination from a map or list. The selected information is sent from the device to the server.
[0252] Input: Information about the travel destination selected by the user.
[0253] Data processing: The server records the selected travel destination information.
[0254] Output: Save the travel destination information to the database.
[0255] Step 4:
[0256] The server transmits information about the selected travel destination to the text generation means to generate a scenario, and the text generation means generates a scenario about the travel destination using the generative AI model.
[0257] Input: Travel destination information.
[0258] Data calculation: A generative AI model generates scenarios based on travel destination information.
[0259] Output: The server receives the generated scenario and stores it in a database.
[0260] Step 5:
[0261] The server transmits the generated scenario to the video generation means, which generates a video. The video generation means generates a video based on the scenario using a generative AI model.
[0262] Input: The generated scenario.
[0263] Data calculation: The generative AI model generates images based on the scenario.
[0264] Output: The server receives the generated video and stores it in a database.
[0265] Step 6:
[0266] The user uses the terminal to play back the generated scenario and video, and the video and scenario are displayed on the user terminal via the providing means.
[0267] Input: Scenarios and footage stored in a database.
[0268] Data processing: The server sends the scenario and video to the user terminal via the provision means.
[0269] Output: The user reads the scenario and watches the video.
[0270] Step 7:
[0271] The user accesses the virtual store in the video and performs virtual shopping. Product information selected by the user is sent from the terminal to the server.
[0272] Input: Product information selected by the user.
[0273] Data processing: The server records the product information and processes the purchase through the virtual shopping means.
[0274] Output: Save product information after purchase completion to the database.
[0275] Step 8:
[0276] The server requests a delivery means to deliver the product selected in the virtual environment and arranges for delivery in the real world.
[0277] Input: Product information selected and purchased by the user.
[0278] Data calculation: The server works with the delivery means to carry out the delivery procedure.
[0279] Output: The product is arranged to be delivered to the address specified by the user.
[0280] Through the above steps, a system is realized in which a user can enjoy virtual travel, do virtual shopping, and receive purchased products in the real world.
[0281] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0282] This invention is a system that allows users to virtually experience travel destinations that they cannot visit, and by combining it with an emotion engine that recognizes the user's emotions, it provides a more personalized experience. The main components of the system of this invention are as follows: a user terminal, a server, a text generation AI, a video generation AI, an emotion engine, and a database.
[0283] Overall system configuration
[0284] Main components of the system
[0285] 1. User terminal: A device (PC, tablet, smartphone, etc.) on which the user selects a travel destination and watches scenarios and videos.
[0286] 2. Server: A back-end system that generates, manages, and provides user information, travel scenarios, and videos.
[0287] 3. Text generation AI: An AI that generates travel scenarios based on selected travel destinations.
[0288] 4. Video generation AI: AI that generates videos of travel destinations based on generated scenarios.
[0289] 5. Emotion engine: A function that recognizes the user's emotions and optimizes scenarios and images according to those emotions.
[0290] 6. Database: A storage system that stores user information, travel scenarios, videos, and emotion data.
[0291] System processing: Example
[0292] Initial settings and user information registration
[0293] 1. A user accesses the system using a user terminal and creates a new account. The user enters information such as name, age, health condition, and travel preferences, and presses the registration button.
[0294] 2. The server receives the input information, validates it, and then saves the information to the database.
[0295] 3. The device displays a "Registration complete" message to the user and redirects them to the login screen.
[0296] Selecting a travel destination
[0297] 4. After logging in, the user accesses the travel destination selection screen and chooses the place they want to go.
[0298] 5. The server receives the selection information and records the details of the selected travel destination in a database.
[0299] Scenario and video generation
[0300] 6. The server sends the selected travel destination information to the text generation AI and generates a travel scenario.
[0301] 7. The text generation AI generates a scenario for the selected travel destination and sends it to the server.
[0302] 8. The server sends a video generation request to the video generation AI based on the generated scenario.
[0303] 9. The video generation AI generates video based on the scenario and sends it to the server.
[0304] 10. The server stores the generated scenario and video in a database.
[0305] Emotion Engine Operation
[0306] 11. The terminal uses the camera and microphone built into the user's device to transmit the user's emotions to the emotion engine.
[0307] 12. The emotion engine analyzes the user's facial expressions, voice, and movements in real time to generate emotion data.
[0308] 13. Based on the emotion data received from the emotion engine, the server sends a request to the text generation AI and video generation AI again to adjust the optimal scenario and video.
[0309] 14. The server provides the adjusted scenario and video to the user's terminal.
[0310] Providing a user experience
[0311] 15. The terminal provides a screen that displays the adjusted video and scenario to the user.
[0312] 16. As the user plays the video and reads the scenario text, the experience is further customized based on feedback from the emotion engine.
[0313] The system recognizes users' emotions in real time and can optimize their travel experience accordingly. For example, if a user is moved by a video of the Eiffel Tower, the system automatically generates and provides related videos and scenarios to further enhance the emotion. In this way, users can transcend physical constraints and enjoy a richer, more personalized travel experience.
[0314] The processing flow will be explained below.
[0315] Step 1:
[0316] A user accesses the system using a user terminal and creates a new account. The user enters information such as name, age, health status, and travel preferences, and presses the registration button.
[0317] Step 2:
[0318] The server receives the input information, validates it, checks for errors, and saves the correct information to the database.
[0319] Step 3:
[0320] The device will display a "Registration Complete" message to the user and redirect them to the login screen.
[0321] Step 4:
[0322] The user enters their user ID and password on the login screen and clicks the login button.
[0323] Step 5:
[0324] The server authenticates the user, checking against information in a database and authorizing if there is a match.
[0325] Step 6:
[0326] The device displays a destination selection screen, providing the user with a world map and a list of destinations.
[0327] Step 7:
[0328] The user selects the desired travel destination and presses the select button.
[0329] Step 8:
[0330] The server receives the user's selection information and generates a request to send details of the selected travel destination to a text generation AI.
[0331] Step 9:
[0332] A text generation AI generates a travel scenario based on the selected travel destination, including information on the history, culture, and key attractions of the tourist destination.
[0333] Step 10:
[0334] The server receives the generated scenario and stores it in a database.
[0335] Step 11:
[0336] The server uses the saved scenario information to send a video generation request to the video generation AI.
[0337] Step 12:
[0338] Based on the scenario, video generation AI generates video footage of the travel destination, including real-time scenery and 3D models.
[0339] Step 13:
[0340] The server receives the generated video and stores it in a database.
[0341] Step 14:
[0342] The terminal displays a screen on which the generated video and scenario can be viewed by the user.
[0343] Step 15:
[0344] The user plays the video, begins reading the scenario text, and clicks on specific locations or events that interest them.
[0345] Step 16:
[0346] The terminal records the user's facial expressions and voice through the camera and microphone built into the user device and transmits them to the emotion engine.
[0347] Step 17:
[0348] The emotion engine analyzes the user's facial expressions, voice, and movements in real time to generate emotion data, which includes the user's emotional state, such as joy, surprise, or boredom.
[0349] Step 18:
[0350] Based on the emotional data received from the emotion engine, the server sends requests to the text generation AI and video generation AI to adjust the optimal scenario and video.
[0351] Step 19:
[0352] The text generation AI and video generation AI generate adjusted scenarios and images based on the emotional data and send them to the server.
[0353] Step 20:
[0354] The server stores the adjusted scenario and video in a database to provide to the user.
[0355] Step 21:
[0356] The device displays the adjusted new scenario and video to the user. For example, if the user is moved, additional video and scenarios will be added to further deepen the user's emotion.
[0357] Based on the above processing steps, a system that provides users with a personalized travel experience is realized, allowing users to enjoy a realistic and rich travel experience optimized according to their emotions, even from the comfort of their own home or hospital room.
[0358] Example 2
[0359] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0360] Conventional virtual travel systems have had difficulty personalizing the user experience. They also lacked the ability to adjust the travel experience in real time to reflect the user's emotions. This resulted in problems such as users becoming bored with the experience midway through and their satisfaction decreasing.
[0361] The specification process by the specification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a selection means for selecting a place the user wants to go to, a text generation means for generating a scenario related to the selected place, a video generation means for generating a video of the place based on the generated scenario, a provision means for providing the generated scenario and video to the user, a request means for the user to request detailed information about the destination based on the scenario or video, a text generation means for generating the detailed information based on the request, an emotion recognition means for recognizing the user's emotion, and an adjustment means for adjusting the scenario and video based on the recognized emotion data. This allows the user's emotions to be reflected in real time, enabling a more personalized and rich travel experience.
[0362] The "selection means" is a function that allows the user to select the place they want to go.
[0363] The "text generation means" is a function for generating scenarios and detailed information about a selected location.
[0364] The "video generation means" is a function for generating video of a location based on the generated scenario.
[0365] The "provision means" is a function for providing the generated scenario and video to the user.
[0366] The "request means" is a function that allows a user to request detailed information about a customer based on a scenario or image.
[0367] The "emotion recognition means" is a function for recognizing the user's emotions.
[0368] The "adjustment means" is a function for adjusting the scenario and video based on the recognized emotion data.
[0369] The "registration means" is a function for registering basic information about a user.
[0370] The "storage means" is a function for storing the registered basic information.
[0371] "Authentication means" is a function that allows a user to access a system.
[0372] The "page providing means" is a function for providing a page exclusively for an authenticated user.
[0373] MODE FOR CARRYING OUT THE INVENTION
[0374] This invention is a system that allows users to virtually experience travel destinations that are difficult to visit in person. This system generates personalized travel experiences based on the user's individual requests, and by incorporating emotion recognition technology, it provides an optimal experience that responds to the user's real-time emotions.
[0375] System Components
[0376] 1. User terminal: A device (e.g., PC, tablet, smartphone) that provides an interface for users to select the destination and experience the virtual journey.
[0377] 2. Server: A back-end system that generates, manages, and provides user information, travel scenarios, and videos.
[0378] 3. Text Generation AI: An artificial intelligence technology that generates travel scenarios based on the locations selected by the user.
[0379] 4. Video Generation AI: Artificial intelligence technology that generates footage of a selected location based on a generated scenario.
[0380] 5. Emotion Recognition Engine: An engine that recognizes user emotions in real time and optimizes the travel experience based on the results.
[0381] 6. Database: A storage system that stores user information, travel scenarios, video, and emotion data.
[0382] System Operation Overview
[0383] When a user accesses the system and selects a place they want to go, the server first retrieves information about the selected place from the database and sends it to a text generation AI. The text generation AI generates a detailed travel scenario based on the selected place and sends the scenario back to the server. The server then sends the generated scenario to a video generation AI, which generates a video of the place based on it.
[0384] The generated scenarios and videos are stored on a server and provided to the user's device. The user can view the scenarios and videos through their device. The emotion recognition engine also uses the user's camera and microphone to analyze emotions from facial expressions and voice in real time and sends the data to the server.
[0385] Optimizing experiences through emotions
[0386] When the server receives the emotion data sent from the emotion recognition engine, it optimizes the travel experience based on that data. For example, if the user is moved, related videos and scenarios are automatically generated to further deepen the emotion, enriching the user's experience. This series of processes is repeated in real time, ensuring that the user always receives an optimized travel experience.
[0387] Specific examples
[0388] For example, if the user selects "The Eiffel Tower in France," the system will send the following prompt to the generative AI model:
[0389] "Generate a scenario in which the user visits the Eiffel Tower in France. The user is a woman in her 20s who likes sightseeing."
[0390] Based on this prompt, the text generation AI generates a detailed scenario about the Eiffel Tower (history, tourist attractions, etc.), and the video generation AI then generates a 360-degree panoramic video based on that scenario.
[0391] In this way, the present invention can reflect the user's emotions in real time and provide a personalized and enriched travel experience.
[0392] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0393] Step 1:
[0394] A user accesses the system using their own device (PC, tablet, smartphone). Using a web browser or a dedicated app, the user enters basic information such as name, age, health status, and travel preferences on the sign-up screen and clicks the "Register" button (input: user's basic information). This information is then sent to the server (output: sending user's basic information).
[0395] Step 2:
[0396] The server receives the basic information sent by the user and performs a validation check to ensure that the format and content are correct (input: basic user information). If validation is successful, this information is saved in the database (output: saved user information). If validation fails, a message is generated to prompt the user to re-enter the invalid information and sent to the user's terminal (output: message prompting re-entry).
[0397] Step 3:
[0398] The terminal receives the "Registration Complete" message from the server and displays it to the user (input: registration complete message). At the same time, it redirects the user to the login screen (output: display of login screen).
[0399] Step 4:
[0400] The user logs in using their account information. After logging in, the user accesses the destination selection screen and chooses the place they want to go (input: login information and destination selection).
[0401] Step 5:
[0402] The server authenticates the user's login information and receives the selected travel destination information (input: login information and travel destination selection). The server records the selected travel destination information in the database and prepares for the next scenario generation (output: saving travel destination information).
[0403] Step 6:
[0404] The server sends travel destination information to the text generation AI and requests it to generate a detailed travel scenario (input: travel destination information) (example prompt: "Please generate a scenario for visiting the Eiffel Tower in France. The user is a woman in her twenties who enjoys sightseeing."). The text generation AI generates a detailed scenario based on the prompt and sends the scenario back to the server (output: generated travel scenario).
[0405] Step 7:
[0406] The server receives the generated scenario and sends it to the video generation AI as a video generation request (input: travel scenario). The video generation AI generates a video of the travel destination based on the scenario and sends it back to the server (output: generated video).
[0407] Step 8:
[0408] The server stores the generated scenario and video in a database (input: generated scenario and video), and then provides the scenario and video to the user's device (output: providing scenario and video).
[0409] Step 9:
[0410] The device sends facial expressions and voice data collected through the user's camera and microphone to the emotion recognition engine in real time (input: facial and voice data). The emotion recognition engine analyzes the data, generates the user's emotion data, and sends it to the server (output: emotion data).
[0411] Step 10:
[0412] The server receives the emotion data and sends a request to the text generation AI and video generation AI again to adjust the optimal scenario and video based on that data (input: emotion data). The generated new scenario and video are sent back to the server (output: adjusted scenario and video).
[0413] Step 11:
[0414] The server provides the adjusted scenario and video back to the user's device (input: adjusted scenario and video), allowing the user to receive a more personalized and enriched travel experience (output: re-provision of scenario and video).
[0415] The above is the specific operation and data processing flow in each processing step of this system.
[0416] (Application example 2)
[0417] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0418] Conventional virtual shopping experience systems allow users to simply browse products in a virtual space, making it difficult to provide a personalized experience based on the user's emotions and preferences. Furthermore, they lack the ability to identify user emotions in real time and optimize the content provided based on those emotions. Furthermore, it is difficult for users to experience the same realistic shopping experience as in a physical store from the comfort of their own home.
[0419] The specification processing by the specification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes a selection means for the user to select a destination they wish to go to, a text generation means for generating a scenario related to the selected destination, a video generation means for generating a video of the destination based on the generated scenario, a provision means for providing the generated scenario and video to the user, a request means for the user to request detailed information about a specific place based on the scenario or video, a text generation means for generating detailed information based on the request, an emotion recognition means for identifying the user's emotion and optimizing the scenario and video to be provided based on the emotion, and an optimization means for regenerating the scenario and video to be provided based on the generated emotion data. This allows the user to have a realistic shopping experience as if they were visiting a physical store, even from the comfort of their own home, and further enables them to enjoy personalized services based on emotion recognition.
[0420] "User" refers to an individual who uses the system.
[0421] "Destination" refers to a virtual location selected by the user that the user wants to visit.
[0422] The "selection means" refers to a means for the user to select the place they want to go.
[0423] "Text generation means" refers to a means for automatically generating a scenario related to a selected destination.
[0424] "Video generation means" refers to means for generating video based on a generated scenario.
[0425] "Providing means" refers to a means for providing the generated scenario and video to the user.
[0426] The "request means" refers to a means for a user to request detailed information about a specific location based on a scenario or video.
[0427] The "emotion recognition means" refers to a means for identifying a user's emotions in real time and optimizing the scenarios and videos provided based on those emotions.
[0428] The "optimization means" refers to a means for regenerating the scenario and video to be provided based on the generated emotion data, and optimizing the user experience.
[0429] "Basic information" refers to personal information registered in the system, such as the user's health status and travel preferences.
[0430] "Registration means" refers to a means for registering basic information of a user in the system.
[0431] "Storage means" refers to the means for storing registered basic information.
[0432] "Authentication means" refers to a means for authenticating a user when logging into a system.
[0433] The "page providing means" refers to a means for providing a page exclusively for an authenticated user.
[0434] "Virtual visit means" refers to a means for allowing a user to progress through the experience as if they had visited the destination.
[0435] The "personalized provision means" refers to a means for providing more personalized products and information to a user based on the emotion identified by the emotion recognition means.
[0436] This invention is a system that allows users to virtually visit a virtual store and have a shopping experience. Furthermore, by combining it with an emotion recognition engine, it provides a personalized experience based on the user's emotions.
[0437] System Components
[0438] 1. User Device:
[0439] This is a device that allows users to select a virtual store and view scenarios and videos. When using a smartphone, a dedicated application is installed. This application is developed using React Native.
[0440] 2. Server:
[0441] This is the backend system that generates, manages, and provides user information, scenarios, and videos. The server uses AWS (registered trademark) EC2 or GCP. The following main components are placed on the server:
[0442] Text generation AI (OpenAI GPT-4 (registered trademark))
[0443] Video generation AI (DALL-E or other 3D video generation models)
[0444] Emotion Engine (Affectiva SDK)
[0445] Database (MySQL (registered trademark) or PostgreSQL)
[0446] Program processing
[0447] 1. Initial settings and user information registration
[0448] A user installs the application and creates a new account, entering personal information such as name, age, preferences, etc., which is then saved to a database. This process is done on the front-end (React Native) and back-end (Node.js / Express).
[0449] 2. Select a store
[0450] After logging in, the user selects the virtual store they want to visit. The store list is retrieved from the server via API communication. The API uses REST or GraphQL.
[0451] 3. Scenario and video generation
[0452] Based on the information of the selected store, the text generation AI (GPT-4) generates a guidance scenario. The generated scenario is sent to the video generation AI (DALL-E), which generates a video of the virtual store. This processing is performed using the OpenAI API and serverless processing (AWS Lambda).
[0453] 4. Operation of the Emotion Engine
[0454] The camera and microphone on the user's device (smartphone) are used to transmit the user's facial expressions and voice to the emotion engine (Affectiva SDK). Emotional data is sent to the server in real time, and the scenario and video are optimized by text generation AI and video generation AI.
[0455] 5. Providing a user experience
[0456] The optimized scenarios and videos are displayed on the user's device, and based on emotion recognition, product listings and custom messages that pique the user's interest are dynamically updated to provide a personalized shopping experience.
[0457] Examples of concrete examples and prompts
[0458] Prompt for text generation AI:
[0459] Enter the user's age, preferences, and selected store information:
[0460] Age: 30
[0461] Interests: Fashion, accessories
[0462] Selected store: High-end fashion mall in the city
[0463] Use this information to generate product lists and scenarios that may be of interest to your users.
[0464] Prompts for video generation AI:
[0465] Generate a 3D image of the interior of a high-end fashion mall in the city, highlighting the product shelves that users are most interested in.
[0466] This system allows users to transcend physical constraints and enjoy a rich and personalized virtual shopping experience. The emotion recognition engine provides optimal content according to the user's emotions, ensuring a highly satisfying experience.
[0467] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0468] Step 1: Initial settings and user information registration
[0469] Input: After users install the application, they enter personal information such as their name, age, preferences, etc.
[0470] What happens: When a user fills out a form and clicks the "Submit" button, the information is sent from the frontend (React Native) to the backend (Node.js / Express), where it is validated and saved to a MySQL or PostgreSQL database.
[0471] Output: The user's personal information is saved in the database and a "Successful registration" message is displayed.
[0472] Step 2: Select a store
[0473] Input: After the user logs in, they select the virtual store they want to visit.
[0474] Specific operation: The store list screen is displayed on the front end, and the user selects the desired store. The selection information is sent to the server, and the information of the selected store is saved in the database.
[0475] Output: The store information selected by the user is saved on the server and used in the next step.
[0476] Step 3: Generate a scenario
[0477] Input: Selected store information.
[0478] Specific operation: The server generates a prompt sentence based on the selected store information and sends it to OpenAI's GPT-4 API. The text generation AI (GPT-4) generates a store guide scenario and returns the result to the server.
[0479] Output: The scenario text generated by GPT-4 is saved on the server.
[0480] Step 4: Generate footage
[0481] Input: Scenario text generated by GPT-4.
[0482] Specific operation: The server generates prompt sentences based on the scenario text and sends them to the video generation AI (DALL-E or other 3D video generation model). The video generation AI generates images of the virtual store based on the scenario and sends the results back to the server.
[0483] Output: The video generated by DALL-E is stored on the server.
[0484] Step 5: Emotion Recognition in Action
[0485] Input: User's facial and voice data.
[0486] Specific operation: The camera and microphone on the user device (smartphone) capture facial expressions and voice, and send them to the emotion recognition engine (Affectiva SDK). The emotion engine analyzes the data, generates emotion data, and sends it to the server.
[0487] Output: The emotion data generated by the emotion recognition engine is stored on the server.
[0488] Step 6: Scenario and footage optimization
[0489] Input: Emotion data, initial scenario and video.
[0490] Specific operation: The server again sends a request to the text generation AI and video generation AI to optimize the scenario and video generated based on the emotion data. As a result, the scenario and video are regenerated based on the user's emotions.
[0491] Output: The new optimized scenario and footage are saved on the server.
[0492] Step 7: Delivering the user experience
[0493] Input: Optimized scenario and footage.
[0494] Specific operation: The user device retrieves the optimized scenario and video from the server and provides it to the user. Product lists and custom messages based on the user's emotion are dynamically displayed.
[0495] Output: The user watches the optimized scenario and video, and enjoys a personalized shopping experience.
[0496] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0497] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (registered trademark) (Internet search engine).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0498] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0499] [Second embodiment]
[0500] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0501] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0502] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0503] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0504] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0505] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0506] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0507] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0508] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0509] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0510] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0511] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0512] The present invention provides a system that enables people who cannot travel to experience a simulated trip from their homes, hospital rooms, etc. This system generates realistic travel scenarios and images based on the travel destinations selected by the user and provides them to the user.
[0513] Overall system configuration
[0514] The system mainly consists of the following components:
[0515] 1. User terminal: A device (PC, tablet, smartphone, etc.) on which the user selects a travel destination and views scenarios and footage.
[0516] 2. Server: A back-end system that generates, manages, and provides user information, travel scenarios, and videos.
[0517] 3. Text generation AI: AI that generates scenarios based on information about the selected travel destination.
[0518] 4. Video generation AI: AI that generates video based on a generated scenario.
[0519] 5. Database: A storage system that stores user information, travel scenarios, and videos.
[0520] System processing: Example
[0521] Initial settings and user information registration
[0522] 1. A user accesses the system using a user terminal and creates a new account.
[0523] Enter your name, age, health status, travel preferences, etc., and press the register button.
[0524] 2. The server receives the user information, validates it, and saves it in the database.
[0525] Selecting a travel destination
[0526] 3. After logging in, the user accesses the travel destination selection screen.
[0527] Choose where you want to go from a world map or a list of travel destinations.
[0528] 4. The server receives the selected destination information and records the user's selection.
[0529] Generate a scenario
[0530] 5. The server sends the selected travel destination information to the text generation AI.
[0531] The text generation AI uses that information to generate travel scenarios, such as the history of the Eiffel Tower or where the Mona Lisa is located.
[0532] 6. The server receives the generated scenario and stores it in the database.
[0533] Video generation
[0534] 7. The server requests the video generation AI to generate a video based on the generated scenario.
[0535] The video generation AI generates 3D models and on-site footage based on the scenario.
[0536] For example, the illuminated Eiffel Tower in Paris and the exhibits at the Louvre Museum.
[0537] Providing a user experience
[0538] 8. The terminal provides the user with a screen that displays the generated scenario and images.
[0539] The user can play the video and read the scenario text.
[0540] For example, you can watch a video of the Eiffel Tower while reading its detailed history.
[0541] Get more information
[0542] 9. When the user selects a specific location or event within the video or scenario, more information is requested.
[0543] 10. The server again asks the text generation AI to generate detailed information based on the request.
[0544] For example, the structure of the Eiffel Tower and stories from its construction.
[0545] 11. The terminal displays the generated details to the user.
[0546] Supplementary information is displayed in real time in sync with the video.
[0547] In this way, users can enjoy a realistic travel experience from their own homes or hospital rooms. This system allows users to enjoy the joy of travel regardless of physical constraints or external factors. Furthermore, the system allows for advanced customization according to the individual needs of each user, providing an optimal experience for each individual user.
[0548] The processing flow will be explained below.
[0549] Step 1:
[0550] A user accesses the system and creates a new account by entering information such as their name, age, health status, and travel preferences into the terminal and pressing the register button.
[0551] Step 2:
[0552] The server receives the input information, validates it, checks for errors, and saves the correct information to the database.
[0553] Step 3:
[0554] The device will display a "Registration Complete" message to the user and redirect them to the login screen.
[0555] Step 4:
[0556] The user enters their user ID and password on the login screen and clicks the login button.
[0557] Step 5:
[0558] The server authenticates the user, checking against information in a database and authorizing if there is a match.
[0559] Step 6:
[0560] The device displays a destination selection screen, providing the user with a world map and a list of destinations.
[0561] Step 7:
[0562] The user selects the desired travel destination and presses the select button.
[0563] Step 8:
[0564] The server receives the user's selection information and generates a request to send details of the selected travel destination to a text generation AI.
[0565] Step 9:
[0566] A text generation AI generates a travel scenario based on the selected travel destination, including information on the history, culture, and key attractions of the tourist destination.
[0567] Step 10:
[0568] The server receives the generated scenario and stores it in a database.
[0569] Step 11:
[0570] The server uses the saved scenario information to send a video generation request to the video generation AI.
[0571] Step 12:
[0572] Based on the scenario, the video generation AI generates footage of the target travel destination, including real-time scenery and 3D models.
[0573] Step 13:
[0574] The server receives the generated video and stores it in a database.
[0575] Step 14:
[0576] The terminal displays a screen on which the generated video and scenario can be viewed by the user.
[0577] Step 15:
[0578] The user plays the video, begins reading the scenario text, and clicks on specific locations or events that interest them.
[0579] Step 16:
[0580] The server receives the user's click information and sends a detailed information generation request back to the text generation AI.
[0581] Step 17:
[0582] Text generation AI generates detailed information, including specific history and unique anecdotes about the location.
[0583] Step 18:
[0584] The server receives the generated details and sends them to the terminal.
[0585] Step 19:
[0586] The device displays detailed information to the user and works in conjunction with the video to enhance the travel experience.
[0587] This series of steps allows users to have a realistic travel experience, as if they were actually visiting the location.
[0588] Example 1
[0589] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0590] In modern society, there are many people who find it difficult to actually travel for various reasons. For example, many people are unable to travel due to health conditions, time constraints, or financial reasons. There is also a need for a way to simulate a trip for elderly people and those who find it difficult to go out due to illness. Furthermore, conventional travel experience systems have difficulty customizing realistic travel scenarios and videos, making it impossible to provide an experience tailored to individual users' needs.
[0591] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0592] In this invention, the server includes a selection means for allowing a user to select a travel destination they wish to go to, a text generation means for generating a scenario related to the selected travel destination, a video generation means for generating a video of the travel destination based on the generated scenario, a provision means for providing the generated scenario and video to the user, a request means for allowing the user to request detailed information about a specific location based on the scenario or video, a text generation means for generating the detailed information based on the request, a display means for displaying the detailed information generated in response to the user's request on a user terminal, and a storage means for registering and saving user information by initial setting. This enables users who have difficulty traveling in person due to health conditions, time constraints, financial reasons, etc. to enjoy a realistic and customized travel experience from the comfort of their own home or hospital room.
[0593] "User" refers to an individual who intends to use the system to have a simulated travel experience.
[0594] "Selection means" refers to an interface or tool that allows a user to select a travel destination they wish to go to.
[0595] "Text generation means" refers to a mechanism that utilizes natural language processing technology to generate scenarios and detailed information about the selected travel destination.
[0596] "Video generation means" refers to the technology and tools used to generate realistic travel destination footage based on the generated scenario.
[0597] "Providing means" refers to the interface or media for providing the generated scenario and video to the user.
[0598] "Request means" refers to an interface or tool that allows a user to request detailed information about a specific location or event based on a scenario or video.
[0599] "Display means" refers to a mechanism for displaying detailed information and images generated in response to a user's request on a user terminal.
[0600] "Storage means" refers to the technology or tools used to register user information in the initial settings and store it in a database, etc.
[0601] MODE FOR CARRYING OUT THE INVENTION
[0602] The present invention relates to a system that allows people who cannot travel to have a simulated real-life travel experience from their homes, hospital rooms, etc. This system generates realistic travel scenarios and images based on travel destinations selected by the user and provides them to the user. The specific configuration and processing flow of the system are described below.
[0603] Overall system configuration
[0604] The system mainly consists of the following components:
[0605] 1. User terminal: A device (PC, tablet, smartphone, etc.) on which the user selects a travel destination and views scenarios and footage.
[0606] 2. Server: A back-end system that generates, manages, and provides user information, travel scenarios, and videos.
[0607] 3. Text generation AI: AI that generates scenarios based on information about the selected travel destination.
[0608] 4. Video generation AI: AI that generates video based on a generated scenario.
[0609] 5. Database: A storage system that stores user information, travel scenarios, and videos.
[0610] System Operation
[0611] Initial settings and user information registration
[0612] A user accesses the system using a user terminal and creates a new account. Specifically, they enter basic information such as name, age, health status, and travel preferences, and then press the register button. At this stage, the server receives the user information, validates the input data, and stores the information in a database. This stored information is used to provide a customized travel experience in the future.
[0613] Selecting a travel destination
[0614] After logging in, the user accesses the travel destination selection screen and selects a destination. The user selects the place they want to go from a world map or a list of travel destinations, and once the selection is complete, the information is sent to the server. The server records this selection information in a database.
[0615] Generate a scenario
[0616] Based on the information about the selected travel destination, the server sends the information to the text generation AI. The text generation AI generates a travel scenario based on this information. For example, it generates information such as "The history of the Eiffel Tower" and "The location where the Mona Lisa is exhibited." The generated scenario is saved in a database by the server.
[0617] Video generation
[0618] The server requests the video generation AI to generate a video based on the generated scenario. The video generation AI generates a 3D model and on-site video in accordance with the scenario. For example, it generates video of the "illuminated Eiffel Tower" or "exhibits at the Louvre Museum." This video is also stored in a database by the server.
[0619] Providing a user experience
[0620] The user device provides a screen that displays the generated scenario and video to the user, allowing the user to experience a simulated trip. For example, a realistic travel experience can be provided, in which a video of the Eiffel Tower is played while a text about its detailed history is read.
[0621] Get more information
[0622] When a user selects a specific location or event within a video or scenario, they can request additional details. Based on this request, the server again asks the text generation AI to generate more information. For example, specific information such as "the structure of the Eiffel Tower" or "episodes from its construction" is generated. The user's device displays this information to the user, providing supplementary information in real time.
[0623] Examples and prompts
[0624] As a concrete example, consider the case where a user selects "Paris" as a travel destination and wants to search for more information about the Eiffel Tower. An example prompt sentence is:
[0625] Generate Scenario prompt:
[0626] Prompt: What is the detailed history of the Eiffel Tower in Paris and what are its popular tourist attractions?
[0627] Video generation prompt:
[0628] Prompt: Generate a video that includes the Eiffel Tower lit up and a museum exhibit.
[0629] summary
[0630] This system allows users to enjoy a realistic travel experience from the comfort of their own home or hospital room. Furthermore, by using generative AI models and prompts, it is possible to customize the system to meet the needs of each individual user, providing a high-quality simulated travel experience.
[0631] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0632] Step 1: Register user information
[0633] A user accesses the system from their terminal and creates a new account. They enter their name, age, health condition, travel preferences, etc., and press the "Register" button. The server receives this information and performs validation (checks the validity of the data). If the input data is correct, the server saves the information in the database. This registers the user information.
[0634] Input: User's name, age, health status, travel preferences
[0635] Data processing / data calculation: User information validation
[0636] Output: Registered user information
[0637] Step 2: Choose your destination
[0638] The user logs in to the system and accesses the travel destination selection screen. The user selects the place they want to go to from a world map or a list of travel destinations and presses the "Select" button. The server receives the selected travel destination information and records it in the database.
[0639] Input: Travel destination selection information
[0640] Data processing / data calculation: Recording selected travel destination information
[0641] Output: Recorded travel destination information
[0642] Step 3: Generate a scenario
[0643] The server retrieves the selected travel destination information from the database and sends the information to the text generation AI based on this. The text generation AI processes the generation prompt and generates a detailed travel scenario. For example, it sends a prompt sentence such as "Prompt: Please tell me the detailed history of the Eiffel Tower in Paris and its popular tourist attractions." The server receives the generated scenario and stores it in the database.
[0644] Input: Selected travel destination information
[0645] Data processing / data calculation: Travel scenario generation
[0646] Output: Generated travel scenarios
[0647] Step 4: Generate footage
[0648] The server requests the video generation AI to generate a video based on the generated scenario. A prompt based on the scenario is sent to the video generation AI, for example, "Prompt: Generate a video that includes the illuminated Eiffel Tower and museum exhibits." The video generation AI generates a 3D model and on-site video based on the scenario, and the server stores the generated video in a database.
[0649] Input: Generated travel scenarios
[0650] Data processing / data calculation: Video generation based on travel scenarios
[0651] Output: Generated video
[0652] Step 5: Delivering the user experience
[0653] The device provides the user with a screen that displays the generated scenarios and videos. The user can enjoy the travel scenarios and videos through the device. For example, the user can play a video of the Eiffel Tower while reading a text about its history.
[0654] Input: Travel scenario and footage
[0655] Data processing / data calculation: Scenario and video display
[0656] Output: Travel scenario and video displayed
[0657] Step 6: Get more information
[0658] The user requests more information about a specific location or event in a video or scenario. For example, if the user wants more information about the structure of the Eiffel Tower, they click on that part. The server receives the request and again asks the text generation AI to generate the detailed information. The text generation AI generates the detailed information based on the specified prompt, and the server sends it to the device. The device then displays the generated detailed information to the user.
[0659] Input: Request for more information
[0660] Data processing / data calculation: generating detailed information based on demand
[0661] Output: Detailed information generated
[0662] In this way, at each processing step, appropriate processing is performed based on the input data, and the results are linked to the next step, providing the user with a realistic and customized travel experience.
[0663] (Application example 1)
[0664] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0665] In recent years, there has been a demand for systems that allow people who cannot travel to simulate travel experiences from their homes, hospital rooms, etc. There is also a need for systems that allow users to enjoy shopping at travel destinations using virtual environments. However, these systems have not yet become widespread, making it difficult for users to simultaneously enjoy realistic travel and shopping experiences. Therefore, the present invention aims to provide a system that integrates simulated travel experiences with virtual shopping, thereby enabling users to have a more realistic and fulfilling experience.
[0666] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0667] In this invention, the server includes a selection means for allowing a user to select a travel destination they wish to go to, a text generation means for generating a scenario related to the selected travel destination, a video generation means for generating a video of the travel destination based on the generated scenario, a provision means for providing the generated scenario and video to the user, a request means for allowing the user to request detailed information about a specific place based on the scenario or video, a text generation means for generating the detailed information based on the request, a virtual shopping means for allowing the user to purchase products in the virtual environment, and a delivery means for delivering the products selected in the virtual environment to the real world. This allows the user to enjoy a virtual trip while shopping from home or a hospital room.
[0668] "User" refers to any individual or entity that uses the System.
[0669] "Destination" refers to a virtual or real geographic location that a user wishes to visit.
[0670] "Selection means" refers to an interface or tool that allows a user to select a travel destination they wish to visit.
[0671] "Text generation means" refers to an artificial intelligence or algorithm that automatically generates a scenario based on specific information.
[0672] "Video generation means" refers to artificial intelligence or a method that automatically creates video based on a generated scenario.
[0673] "Providing means" refers to a technical method or device for displaying or playing the generated scenario and video to the user.
[0674] "Request means" refers to an interface or tool that allows a user to request detailed information about a specific location based on a scenario or video.
[0675] "More Information" refers to additional and specific information about a particular location or event.
[0676] "Virtual shopping vehicle" refers to an interface or tool that allows a user to select and purchase items within a virtual environment.
[0677] "Delivery Method" refers to a method or service for delivering a product selected in a virtual environment to a user in the real world.
[0678] A specific embodiment for carrying out the present invention is described below. This system is designed to provide an integrated experience that allows users to enjoy virtual travel and virtual shopping. The details are described below.
[0679] 1. Overall system configuration
[0680] This system mainly consists of a user terminal, a server, a generative AI model, a database, and a delivery service.
[0681] 1.1 User terminal
[0682] The user terminal is a device that allows users to select travel destinations, view scenarios and videos, and even conduct virtual shopping. Specifically, this applies to smartphones and head-mounted displays.
[0683] 1.2 Server
[0684] The server is a back-end system that generates, manages, and provides user information, travel scenarios, videos, and shopping information. It has the following means:
[0685] 1. Selection means: Provides an interface for the user to select the travel destination they wish to go to.
[0686] 2. Text generation method: Use a generative AI model (e.g., OpenAI's GPT-3.5) to generate travel destination scenarios.
[0687] 3. Video generation method: A generative AI model for generating video based on a scenario.
[0688] 4. Provision means: The generated scenario and video are provided to the user.
[0689] 5. Request means: Provides an interface for users to request more information.
[0690] 6. Virtual Shopping Vehicle: Provides an interface for users to purchase products within a virtual environment.
[0691] 7. Delivery Method: Providing a service for delivering the items selected in the virtual environment in the real world.
[0692] 1.3 Database
[0693] The database is a storage system for storing user information, travel scenarios, videos, and product information.
[0694] 1.4 Delivery Services
[0695] The delivery service arranges for the items purchased by the user through virtual shopping to be delivered in the real world.
[0696] 2. Processing Overview
[0697] 2.1 User Registration and Settings
[0698] A user accesses the system using a user terminal and creates a new account. The server receives the user information, validates it, and stores it in the database.
[0699] 2.2 Choosing a travel destination
[0700] After logging in, users select the places they want to virtually visit from a world map or a list of destinations. The server receives the selected travel destination information and records the user's selection.
[0701] 2.3 Scenario and image generation
[0702] The server transmits information about the selected travel destination to the text generation means, receives the generated scenario, and stores it in a database.The server then transmits the scenario to the video generation means and receives the generated video.
[0703] 2.4 Virtual Shopping
[0704] Users can access the virtual store in the generated video and purchase products, which are then delivered to the real world using a delivery method.
[0705] 3. Examples of concrete examples and prompts
[0706] For example, if the user selects "Paris," the system performs the following process:
[0707] 1. The scenario generation AI generates a scenario that includes information about the Eiffel Tower and the Louvre Museum.
[0708] 2. Image generation AI generates images of actual Paris cityscapes and tourist attractions.
[0709] 3. Users can purchase limited edition keychains and delicious local chocolates at the Eiffel Tower's virtual store.
[0710] Prompt Sentence Examples
[0711] "You're currently on a virtual tour of Paris. See the beautifully lit Eiffel Tower and shop for exclusive keychains and delicious local chocolates at a nearby souvenir shop."
[0712] This system allows users to enjoy virtual travel and shopping from the comfort of their own home or hospital room.
[0713] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0714] Step 1:
[0715] A user accesses the system using a terminal and creates a new account. The user enters basic information such as name, age, and travel preferences, and presses the register button. The basic information is sent from the terminal to the server.
[0716] Input: Basic information you enter into your device, such as your name, age, and travel preferences.
[0717] Data processing: The server validates basic information.
[0718] Output: The server saves the basic information that passes validation to the database.
[0719] Step 2:
[0720] The user accesses the login screen and logs in using the authentication method. The server checks the user's authentication information against the database and provides a dedicated page for the authenticated user.
[0721] Input: User credentials (username, password).
[0722] Data calculation: The server performs authentication by checking against information in the database.
[0723] Output: Authenticated users are served a dedicated page.
[0724] Step 3:
[0725] The user accesses the travel destination selection screen from a dedicated page and selects the desired travel destination from a map or list. The selected information is sent from the device to the server.
[0726] Input: Information about the travel destination selected by the user.
[0727] Data processing: The server records the selected travel destination information.
[0728] Output: Save the travel destination information to the database.
[0729] Step 4:
[0730] The server transmits information about the selected travel destination to the text generation means to generate a scenario, and the text generation means generates a scenario about the travel destination using the generative AI model.
[0731] Input: Travel destination information.
[0732] Data calculation: A generative AI model generates scenarios based on travel destination information.
[0733] Output: The server receives the generated scenario and stores it in a database.
[0734] Step 5:
[0735] The server transmits the generated scenario to the video generation means, which generates a video. The video generation means generates a video based on the scenario using a generative AI model.
[0736] Input: The generated scenario.
[0737] Data calculation: The generative AI model generates images based on the scenario.
[0738] Output: The server receives the generated video and stores it in a database.
[0739] Step 6:
[0740] The user uses the terminal to play back the generated scenario and video, and the video and scenario are displayed on the user terminal via the providing means.
[0741] Input: Scenarios and footage stored in a database.
[0742] Data processing: The server sends the scenario and video to the user terminal via the provision means.
[0743] Output: The user reads the scenario and watches the video.
[0744] Step 7:
[0745] The user accesses the virtual store in the video and performs virtual shopping. Product information selected by the user is sent from the terminal to the server.
[0746] Input: Product information selected by the user.
[0747] Data processing: The server records the product information and processes the purchase through the virtual shopping means.
[0748] Output: Save product information after purchase completion to the database.
[0749] Step 8:
[0750] The server requests a delivery means to deliver the product selected in the virtual environment and arranges for delivery in the real world.
[0751] Input: Product information selected and purchased by the user.
[0752] Data calculation: The server works with the delivery means to carry out the delivery procedure.
[0753] Output: The product is arranged to be delivered to the address specified by the user.
[0754] Through the above steps, a system is realized in which a user can enjoy virtual travel, do virtual shopping, and receive purchased products in the real world.
[0755] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0756] This invention is a system that allows users to virtually experience travel destinations that they cannot visit, and by combining it with an emotion engine that recognizes the user's emotions, it provides a more personalized experience. The main components of the system of this invention are as follows: a user terminal, a server, a text generation AI, a video generation AI, an emotion engine, and a database.
[0757] Overall system configuration
[0758] Main components of the system
[0759] 1. User terminal: A device (PC, tablet, smartphone, etc.) on which the user selects a travel destination and watches scenarios and videos.
[0760] 2. Server: A back-end system that generates, manages, and provides user information, travel scenarios, and videos.
[0761] 3. Text generation AI: An AI that generates travel scenarios based on selected travel destinations.
[0762] 4. Video generation AI: AI that generates videos of travel destinations based on generated scenarios.
[0763] 5. Emotion engine: A function that recognizes the user's emotions and optimizes scenarios and images according to those emotions.
[0764] 6. Database: A storage system that stores user information, travel scenarios, videos, and emotion data.
[0765] System processing: Example
[0766] Initial settings and user information registration
[0767] 1. A user accesses the system using a user terminal and creates a new account. The user enters information such as name, age, health condition, and travel preferences, and presses the registration button.
[0768] 2. The server receives the input information, validates it, and then saves the information to the database.
[0769] 3. The device displays a "Registration complete" message to the user and redirects them to the login screen.
[0770] Selecting a travel destination
[0771] 4. After logging in, the user accesses the travel destination selection screen and chooses the place they want to go.
[0772] 5. The server receives the selection information and records the details of the selected travel destination in a database.
[0773] Scenario and video generation
[0774] 6. The server sends the selected travel destination information to the text generation AI and generates a travel scenario.
[0775] 7. The text generation AI generates a scenario for the selected travel destination and sends it to the server.
[0776] 8. The server sends a video generation request to the video generation AI based on the generated scenario.
[0777] 9. The video generation AI generates video based on the scenario and sends it to the server.
[0778] 10. The server stores the generated scenario and video in a database.
[0779] Emotion Engine Operation
[0780] 11. The terminal uses the camera and microphone built into the user's device to transmit the user's emotions to the emotion engine.
[0781] 12. The emotion engine analyzes the user's facial expressions, voice, and movements in real time to generate emotion data.
[0782] 13. Based on the emotion data received from the emotion engine, the server sends a request to the text generation AI and video generation AI again to adjust the optimal scenario and video.
[0783] 14. The server provides the adjusted scenario and video to the user's terminal.
[0784] Providing a user experience
[0785] 15. The terminal provides a screen that displays the adjusted video and scenario to the user.
[0786] 16. As the user plays the video and reads the scenario text, the experience is further customized based on feedback from the emotion engine.
[0787] The system recognizes users' emotions in real time and can optimize their travel experience accordingly. For example, if a user is moved by a video of the Eiffel Tower, the system automatically generates and provides related videos and scenarios to further enhance the emotion. In this way, users can transcend physical constraints and enjoy a richer, more personalized travel experience.
[0788] The processing flow will be explained below.
[0789] Step 1:
[0790] A user accesses the system using a user terminal and creates a new account. The user enters information such as name, age, health status, and travel preferences, and presses the registration button.
[0791] Step 2:
[0792] The server receives the input information, validates it, checks for errors, and saves the correct information to the database.
[0793] Step 3:
[0794] The device will display a "Registration Complete" message to the user and redirect them to the login screen.
[0795] Step 4:
[0796] The user enters their user ID and password on the login screen and clicks the login button.
[0797] Step 5:
[0798] The server authenticates the user, checking against information in a database and authorizing if there is a match.
[0799] Step 6:
[0800] The device displays a destination selection screen, providing the user with a world map and a list of destinations.
[0801] Step 7:
[0802] The user selects the desired travel destination and presses the select button.
[0803] Step 8:
[0804] The server receives the user's selection information and generates a request to send details of the selected travel destination to a text generation AI.
[0805] Step 9:
[0806] A text generation AI generates a travel scenario based on the selected travel destination, including information on the history, culture, and key attractions of the tourist destination.
[0807] Step 10:
[0808] The server receives the generated scenario and stores it in a database.
[0809] Step 11:
[0810] The server uses the saved scenario information to send a video generation request to the video generation AI.
[0811] Step 12:
[0812] Based on the scenario, video generation AI generates video footage of the travel destination, including real-time scenery and 3D models.
[0813] Step 13:
[0814] The server receives the generated video and stores it in a database.
[0815] Step 14:
[0816] The terminal displays a screen on which the generated video and scenario can be viewed by the user.
[0817] Step 15:
[0818] The user plays the video, begins reading the scenario text, and clicks on specific locations or events that interest them.
[0819] Step 16:
[0820] The terminal records the user's facial expressions and voice through the camera and microphone built into the user device and transmits them to the emotion engine.
[0821] Step 17:
[0822] The emotion engine analyzes the user's facial expressions, voice, and movements in real time to generate emotion data, which includes the user's emotional state, such as joy, surprise, or boredom.
[0823] Step 18:
[0824] Based on the emotional data received from the emotion engine, the server sends requests to the text generation AI and video generation AI to adjust the optimal scenario and video.
[0825] Step 19:
[0826] The text generation AI and video generation AI generate adjusted scenarios and images based on the emotional data and send them to the server.
[0827] Step 20:
[0828] The server stores the adjusted scenario and video in a database to provide to the user.
[0829] Step 21:
[0830] The device displays the adjusted new scenario and video to the user. For example, if the user is moved, additional video and scenarios will be added to further deepen the user's emotion.
[0831] Based on the above processing steps, a system that provides users with a personalized travel experience is realized, allowing users to enjoy a realistic and rich travel experience optimized according to their emotions, even from the comfort of their own home or hospital room.
[0832] Example 2
[0833] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0834] Conventional virtual travel systems have had difficulty personalizing the user experience. They also lacked the ability to adjust the travel experience in real time to reflect the user's emotions. This resulted in problems such as users becoming bored with the experience midway through and their satisfaction decreasing.
[0835] The specification process by the specification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a selection means for selecting a place the user wants to go to, a text generation means for generating a scenario related to the selected place, a video generation means for generating a video of the place based on the generated scenario, a provision means for providing the generated scenario and video to the user, a request means for the user to request detailed information about the destination based on the scenario or video, a text generation means for generating the detailed information based on the request, an emotion recognition means for recognizing the user's emotion, and an adjustment means for adjusting the scenario and video based on the recognized emotion data. This allows the user's emotions to be reflected in real time, enabling a more personalized and rich travel experience.
[0836] The "selection means" is a function that allows the user to select the place they want to go.
[0837] The "text generation means" is a function for generating scenarios and detailed information about a selected location.
[0838] The "video generation means" is a function for generating video of a location based on the generated scenario.
[0839] The "provision means" is a function for providing the generated scenario and video to the user.
[0840] The "request means" is a function that allows a user to request detailed information about a customer based on a scenario or image.
[0841] The "emotion recognition means" is a function for recognizing the user's emotions.
[0842] The "adjustment means" is a function for adjusting the scenario and video based on the recognized emotion data.
[0843] The "registration means" is a function for registering basic information about a user.
[0844] The "storage means" is a function for storing the registered basic information.
[0845] "Authentication means" is a function that allows a user to access a system.
[0846] The "page providing means" is a function for providing a page exclusively for an authenticated user.
[0847] MODE FOR CARRYING OUT THE INVENTION
[0848] This invention is a system that allows users to virtually experience travel destinations that are difficult to visit in person. This system generates personalized travel experiences based on the user's individual requests, and by incorporating emotion recognition technology, it provides an optimal experience that responds to the user's real-time emotions.
[0849] System Components
[0850] 1. User terminal: A device (e.g., PC, tablet, smartphone) that provides an interface for users to select the destination and experience the virtual journey.
[0851] 2. Server: A back-end system that generates, manages, and provides user information, travel scenarios, and videos.
[0852] 3. Text Generation AI: An artificial intelligence technology that generates travel scenarios based on the locations selected by the user.
[0853] 4. Video Generation AI: Artificial intelligence technology that generates footage of a selected location based on a generated scenario.
[0854] 5. Emotion Recognition Engine: An engine that recognizes user emotions in real time and optimizes the travel experience based on the results.
[0855] 6. Database: A storage system that stores user information, travel scenarios, video, and emotion data.
[0856] System Operation Overview
[0857] When a user accesses the system and selects a place they want to go, the server first retrieves information about the selected place from the database and sends it to a text generation AI. The text generation AI generates a detailed travel scenario based on the selected place and sends the scenario back to the server. The server then sends the generated scenario to a video generation AI, which generates a video of the place based on it.
[0858] The generated scenarios and videos are stored on a server and provided to the user's device. The user can view the scenarios and videos through their device. The emotion recognition engine also uses the user's camera and microphone to analyze emotions from facial expressions and voice in real time and sends the data to the server.
[0859] Optimizing experiences through emotions
[0860] When the server receives the emotion data sent from the emotion recognition engine, it optimizes the travel experience based on that data. For example, if the user is moved, related videos and scenarios are automatically generated to further deepen the emotion, enriching the user's experience. This series of processes is repeated in real time, ensuring that the user always receives an optimized travel experience.
[0861] Specific examples
[0862] For example, if the user selects "The Eiffel Tower in France," the system will send the following prompt to the generative AI model:
[0863] "Generate a scenario in which the user visits the Eiffel Tower in France. The user is a woman in her 20s who likes sightseeing."
[0864] Based on this prompt, the text generation AI generates a detailed scenario about the Eiffel Tower (history, tourist attractions, etc.), and the video generation AI then generates a 360-degree panoramic video based on that scenario.
[0865] In this way, the present invention can reflect the user's emotions in real time and provide a personalized and enriched travel experience.
[0866] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0867] Step 1:
[0868] A user accesses the system using their own device (PC, tablet, smartphone). Using a web browser or a dedicated app, the user enters basic information such as name, age, health status, and travel preferences on the sign-up screen and clicks the "Register" button (input: user's basic information). This information is then sent to the server (output: sending user's basic information).
[0869] Step 2:
[0870] The server receives the basic information sent by the user and performs a validation check to ensure that the format and content are correct (input: basic user information). If validation is successful, this information is saved in the database (output: saved user information). If validation fails, a message is generated to prompt the user to re-enter the invalid information and sent to the user's terminal (output: message prompting re-entry).
[0871] Step 3:
[0872] The terminal receives the "Registration Complete" message from the server and displays it to the user (input: registration complete message). At the same time, it redirects the user to the login screen (output: display of login screen).
[0873] Step 4:
[0874] The user logs in using their account information. After logging in, the user accesses the destination selection screen and chooses the place they want to go (input: login information and destination selection).
[0875] Step 5:
[0876] The server authenticates the user's login information and receives the selected travel destination information (input: login information and travel destination selection). The server records the selected travel destination information in the database and prepares for the next scenario generation (output: saving travel destination information).
[0877] Step 6:
[0878] The server sends travel destination information to the text generation AI and requests it to generate a detailed travel scenario (input: travel destination information) (example prompt: "Please generate a scenario for visiting the Eiffel Tower in France. The user is a woman in her twenties who enjoys sightseeing."). The text generation AI generates a detailed scenario based on the prompt and sends the scenario back to the server (output: generated travel scenario).
[0879] Step 7:
[0880] The server receives the generated scenario and sends it to the video generation AI as a video generation request (input: travel scenario). The video generation AI generates a video of the travel destination based on the scenario and sends it back to the server (output: generated video).
[0881] Step 8:
[0882] The server stores the generated scenario and video in a database (input: generated scenario and video), and then provides the scenario and video to the user's device (output: providing scenario and video).
[0883] Step 9:
[0884] The device sends facial expressions and voice data collected through the user's camera and microphone to the emotion recognition engine in real time (input: facial and voice data). The emotion recognition engine analyzes the data, generates the user's emotion data, and sends it to the server (output: emotion data).
[0885] Step 10:
[0886] The server receives the emotion data and sends a request to the text generation AI and video generation AI again to adjust the optimal scenario and video based on that data (input: emotion data). The generated new scenario and video are sent back to the server (output: adjusted scenario and video).
[0887] Step 11:
[0888] The server provides the adjusted scenario and video back to the user's device (input: adjusted scenario and video), allowing the user to receive a more personalized and enriched travel experience (output: re-provision of scenario and video).
[0889] The above is the specific operation and data processing flow in each processing step of this system.
[0890] (Application example 2)
[0891] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0892] Conventional virtual shopping experience systems allow users to simply browse products in a virtual space, making it difficult to provide a personalized experience based on the user's emotions and preferences. Furthermore, they lack the ability to identify user emotions in real time and optimize the content provided based on those emotions. Furthermore, it is difficult for users to experience the same realistic shopping experience as in a physical store from the comfort of their own home.
[0893] The specification processing by the specification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes a selection means for the user to select a destination they wish to go to, a text generation means for generating a scenario related to the selected destination, a video generation means for generating a video of the destination based on the generated scenario, a provision means for providing the generated scenario and video to the user, a request means for the user to request detailed information about a specific place based on the scenario or video, a text generation means for generating detailed information based on the request, an emotion recognition means for identifying the user's emotion and optimizing the scenario and video to be provided based on the emotion, and an optimization means for regenerating the scenario and video to be provided based on the generated emotion data. This allows the user to have a realistic shopping experience as if they were visiting a physical store, even from the comfort of their own home, and further enables them to enjoy personalized services based on emotion recognition.
[0894] "User" refers to an individual who uses the system.
[0895] "Destination" refers to a virtual location selected by the user that the user wants to visit.
[0896] The "selection means" refers to a means for the user to select the place they want to go.
[0897] "Text generation means" refers to a means for automatically generating a scenario related to a selected destination.
[0898] "Video generation means" refers to means for generating video based on a generated scenario.
[0899] "Providing means" refers to a means for providing the generated scenario and video to the user.
[0900] The "request means" refers to a means for a user to request detailed information about a specific location based on a scenario or video.
[0901] The "emotion recognition means" refers to a means for identifying a user's emotions in real time and optimizing the scenarios and videos provided based on those emotions.
[0902] The "optimization means" refers to a means for regenerating the scenario and video to be provided based on the generated emotion data, and optimizing the user experience.
[0903] "Basic information" refers to personal information registered in the system, such as the user's health status and travel preferences.
[0904] "Registration means" refers to a means for registering basic information of a user in the system.
[0905] "Storage means" refers to the means for storing registered basic information.
[0906] "Authentication means" refers to a means for authenticating a user when logging into a system.
[0907] The "page providing means" refers to a means for providing a page exclusively for an authenticated user.
[0908] "Virtual visit means" refers to a means for allowing a user to progress through the experience as if they had visited the destination.
[0909] The "personalized provision means" refers to a means for providing more personalized products and information to a user based on the emotion identified by the emotion recognition means.
[0910] This invention is a system that allows users to virtually visit a virtual store and have a shopping experience. Furthermore, by combining it with an emotion recognition engine, it provides a personalized experience based on the user's emotions.
[0911] System Components
[0912] 1. User Device:
[0913] This is a device that allows users to select a virtual store and view scenarios and videos. When using a smartphone, a dedicated application is installed. This application is developed using React Native.
[0914] 2. Server:
[0915] This is the backend system that generates, manages, and provides user information, scenarios, and videos. The server uses AWS EC2 or GCP. The following main components are placed on the server:
[0916] Text generation AI (OpenAI GPT-4)
[0917] Video generation AI (DALL-E or other 3D video generation models)
[0918] Emotion Engine (Affectiva SDK)
[0919] Database (MySQL or PostgreSQL)
[0920] Program processing
[0921] 1. Initial settings and user information registration
[0922] A user installs the application and creates a new account, entering personal information such as name, age, preferences, etc., which is then saved to a database. This process is done on the front-end (React Native) and back-end (Node.js / Express).
[0923] 2. Select a store
[0924] After logging in, the user selects the virtual store they want to visit. The store list is retrieved from the server via API communication. The API uses REST or GraphQL.
[0925] 3. Scenario and video generation
[0926] Based on the information of the selected store, the text generation AI (GPT-4) generates a guidance scenario. The generated scenario is sent to the video generation AI (DALL-E), which generates a video of the virtual store. This processing is performed using the OpenAI API and serverless processing (AWS Lambda).
[0927] 4. Operation of the Emotion Engine
[0928] The camera and microphone on the user's device (smartphone) are used to transmit the user's facial expressions and voice to the emotion engine (Affectiva SDK). Emotional data is sent to the server in real time, and the scenario and video are optimized by text generation AI and video generation AI.
[0929] 5. Providing a user experience
[0930] The optimized scenarios and videos are displayed on the user's device, and based on emotion recognition, product listings and custom messages that pique the user's interest are dynamically updated to provide a personalized shopping experience.
[0931] Examples of concrete examples and prompts
[0932] Prompt for text generation AI:
[0933] Enter the user's age, preferences, and selected store information:
[0934] Age: 30
[0935] Interests: Fashion, accessories
[0936] Selected store: High-end fashion mall in the city
[0937] Use this information to generate product lists and scenarios that may be of interest to your users.
[0938] Prompts for video generation AI:
[0939] Generate a 3D image of the interior of a high-end fashion mall in the city, highlighting the product shelves that users are most interested in.
[0940] This system allows users to transcend physical constraints and enjoy a rich and personalized virtual shopping experience. The emotion recognition engine provides optimal content according to the user's emotions, ensuring a highly satisfying experience.
[0941] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0942] Step 1: Initial settings and user information registration
[0943] Input: After users install the application, they enter personal information such as their name, age, preferences, etc.
[0944] What happens: When a user fills out a form and clicks the "Submit" button, the information is sent from the frontend (React Native) to the backend (Node.js / Express), where it is validated and saved to a MySQL or PostgreSQL database.
[0945] Output: The user's personal information is saved in the database and a "Successful registration" message is displayed.
[0946] Step 2: Select a store
[0947] Input: After the user logs in, they select the virtual store they want to visit.
[0948] Specific operation: The store list screen is displayed on the front end, and the user selects the desired store. The selection information is sent to the server, and the information of the selected store is saved in the database.
[0949] Output: The store information selected by the user is saved on the server and used in the next step.
[0950] Step 3: Generate a scenario
[0951] Input: Selected store information.
[0952] Specific operation: The server generates a prompt sentence based on the selected store information and sends it to OpenAI's GPT-4 API. The text generation AI (GPT-4) generates a store guide scenario and returns the result to the server.
[0953] Output: The scenario text generated by GPT-4 is saved on the server.
[0954] Step 4: Generate footage
[0955] Input: Scenario text generated by GPT-4.
[0956] Specific operation: The server generates prompt sentences based on the scenario text and sends them to the video generation AI (DALL-E or other 3D video generation model). The video generation AI generates images of the virtual store based on the scenario and sends the results back to the server.
[0957] Output: The video generated by DALL-E is stored on the server.
[0958] Step 5: Emotion Recognition in Action
[0959] Input: User's facial and voice data.
[0960] Specific operation: The camera and microphone on the user device (smartphone) capture facial expressions and voice, and send them to the emotion recognition engine (Affectiva SDK). The emotion engine analyzes the data, generates emotion data, and sends it to the server.
[0961] Output: The emotion data generated by the emotion recognition engine is stored on the server.
[0962] Step 6: Scenario and footage optimization
[0963] Input: Emotion data, initial scenario and video.
[0964] Specific operation: The server again sends a request to the text generation AI and video generation AI to optimize the scenario and video generated based on the emotion data. As a result, the scenario and video are regenerated based on the user's emotions.
[0965] Output: The new optimized scenario and footage are saved on the server.
[0966] Step 7: Delivering the user experience
[0967] Input: Optimized scenario and footage.
[0968] Specific operation: The user device retrieves the optimized scenario and video from the server and provides it to the user. Product lists and custom messages based on the user's emotion are dynamically displayed.
[0969] Output: The user watches the optimized scenario and video, and enjoys a personalized shopping experience.
[0970] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0971] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0972] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0973] [Third embodiment]
[0974] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0975] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[0976] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0977] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0978] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0979] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0980] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0981] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0982] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0983] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0984] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0985] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0986] The present invention provides a system that enables people who cannot travel to experience a simulated trip from their homes, hospital rooms, etc. This system generates realistic travel scenarios and images based on the travel destinations selected by the user and provides them to the user.
[0987] Overall system configuration
[0988] The system mainly consists of the following components:
[0989] 1. User terminal: A device (PC, tablet, smartphone, etc.) on which the user selects a travel destination and views scenarios and footage.
[0990] 2. Server: A back-end system that generates, manages, and provides user information, travel scenarios, and videos.
[0991] 3. Text generation AI: AI that generates scenarios based on information about the selected travel destination.
[0992] 4. Video generation AI: AI that generates video based on a generated scenario.
[0993] 5. Database: A storage system that stores user information, travel scenarios, and videos.
[0994] System processing: Example
[0995] Initial settings and user information registration
[0996] 1. A user accesses the system using a user terminal and creates a new account.
[0997] Enter your name, age, health status, travel preferences, etc., and press the register button.
[0998] 2. The server receives the user information, validates it, and saves it in the database.
[0999] Selecting a travel destination
[1000] 3. After logging in, the user accesses the travel destination selection screen.
[1001] Choose where you want to go from a world map or a list of travel destinations.
[1002] 4. The server receives the selected destination information and records the user's selection.
[1003] Generate a scenario
[1004] 5. The server sends the selected travel destination information to the text generation AI.
[1005] The text generation AI uses that information to generate travel scenarios, such as the history of the Eiffel Tower or where the Mona Lisa is located.
[1006] 6. The server receives the generated scenario and stores it in the database.
[1007] Video generation
[1008] 7. The server requests the video generation AI to generate a video based on the generated scenario.
[1009] The video generation AI generates 3D models and on-site footage based on the scenario.
[1010] For example, the illuminated Eiffel Tower in Paris and the exhibits at the Louvre Museum.
[1011] Providing a user experience
[1012] 8. The terminal provides the user with a screen that displays the generated scenario and images.
[1013] The user can play the video and read the scenario text.
[1014] For example, you can watch a video of the Eiffel Tower while reading its detailed history.
[1015] Get more information
[1016] 9. When the user selects a specific location or event within the video or scenario, more information is requested.
[1017] 10. The server again asks the text generation AI to generate detailed information based on the request.
[1018] For example, the structure of the Eiffel Tower and stories from its construction.
[1019] 11. The terminal displays the generated details to the user.
[1020] Supplementary information is displayed in real time in sync with the video.
[1021] In this way, users can enjoy a realistic travel experience from their own homes or hospital rooms. This system allows users to enjoy the joy of travel regardless of physical constraints or external factors. Furthermore, the system allows for advanced customization according to the individual needs of each user, providing an optimal experience for each individual user.
[1022] The processing flow will be explained below.
[1023] Step 1:
[1024] A user accesses the system and creates a new account by entering information such as their name, age, health status, and travel preferences into the terminal and pressing the register button.
[1025] Step 2:
[1026] The server receives the input information, validates it, checks for errors, and saves the correct information to the database.
[1027] Step 3:
[1028] The device will display a "Registration Complete" message to the user and redirect them to the login screen.
[1029] Step 4:
[1030] The user enters their user ID and password on the login screen and clicks the login button.
[1031] Step 5:
[1032] The server authenticates the user, checking against information in a database and authorizing if there is a match.
[1033] Step 6:
[1034] The device displays a destination selection screen, providing the user with a world map and a list of destinations.
[1035] Step 7:
[1036] The user selects the desired travel destination and presses the select button.
[1037] Step 8:
[1038] The server receives the user's selection information and generates a request to send details of the selected travel destination to a text generation AI.
[1039] Step 9:
[1040] A text generation AI generates a travel scenario based on the selected travel destination, including information on the history, culture, and key attractions of the tourist destination.
[1041] Step 10:
[1042] The server receives the generated scenario and stores it in a database.
[1043] Step 11:
[1044] The server uses the saved scenario information to send a video generation request to the video generation AI.
[1045] Step 12:
[1046] Based on the scenario, the video generation AI generates footage of the target travel destination, including real-time scenery and 3D models.
[1047] Step 13:
[1048] The server receives the generated video and stores it in a database.
[1049] Step 14:
[1050] The terminal displays a screen on which the generated video and scenario can be viewed by the user.
[1051] Step 15:
[1052] The user plays the video, begins reading the scenario text, and clicks on specific locations or events that interest them.
[1053] Step 16:
[1054] The server receives the user's click information and sends a detailed information generation request back to the text generation AI.
[1055] Step 17:
[1056] Text generation AI generates detailed information, including specific history and unique anecdotes about the location.
[1057] Step 18:
[1058] The server receives the generated details and sends them to the terminal.
[1059] Step 19:
[1060] The device displays detailed information to the user and works in conjunction with the video to enhance the travel experience.
[1061] This series of steps allows users to have a realistic travel experience, as if they were actually visiting the location.
[1062] Example 1
[1063] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1064] In modern society, there are many people who find it difficult to actually travel for various reasons. For example, many people are unable to travel due to health conditions, time constraints, or financial reasons. There is also a need for a way to simulate a trip for elderly people and those who find it difficult to go out due to illness. Furthermore, conventional travel experience systems have difficulty customizing realistic travel scenarios and videos, making it impossible to provide an experience tailored to individual users' needs.
[1065] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1066] In this invention, the server includes a selection means for allowing a user to select a travel destination they wish to go to, a text generation means for generating a scenario related to the selected travel destination, a video generation means for generating a video of the travel destination based on the generated scenario, a provision means for providing the generated scenario and video to the user, a request means for allowing the user to request detailed information about a specific location based on the scenario or video, a text generation means for generating the detailed information based on the request, a display means for displaying the detailed information generated in response to the user's request on a user terminal, and a storage means for registering and saving user information by initial setting. This enables users who have difficulty traveling in person due to health conditions, time constraints, financial reasons, etc. to enjoy a realistic and customized travel experience from the comfort of their own home or hospital room.
[1067] "User" refers to an individual who intends to use the system to have a simulated travel experience.
[1068] "Selection means" refers to an interface or tool that allows a user to select a travel destination they wish to go to.
[1069] "Text generation means" refers to a mechanism that utilizes natural language processing technology to generate scenarios and detailed information about the selected travel destination.
[1070] "Video generation means" refers to the technology and tools used to generate realistic travel destination footage based on the generated scenario.
[1071] "Providing means" refers to the interface or media for providing the generated scenario and video to the user.
[1072] "Request means" refers to an interface or tool that allows a user to request detailed information about a specific location or event based on a scenario or video.
[1073] "Display means" refers to a mechanism for displaying detailed information and images generated in response to a user's request on a user terminal.
[1074] "Storage means" refers to the technology or tools used to register user information in the initial settings and store it in a database, etc.
[1075] MODE FOR CARRYING OUT THE INVENTION
[1076] The present invention relates to a system that allows people who cannot travel to have a simulated real-life travel experience from their homes, hospital rooms, etc. This system generates realistic travel scenarios and images based on travel destinations selected by the user and provides them to the user. The specific configuration and processing flow of the system are described below.
[1077] Overall system configuration
[1078] The system mainly consists of the following components:
[1079] 1. User terminal: A device (PC, tablet, smartphone, etc.) on which the user selects a travel destination and views scenarios and footage.
[1080] 2. Server: A back-end system that generates, manages, and provides user information, travel scenarios, and videos.
[1081] 3. Text generation AI: AI that generates scenarios based on information about the selected travel destination.
[1082] 4. Video generation AI: AI that generates video based on a generated scenario.
[1083] 5. Database: A storage system that stores user information, travel scenarios, and videos.
[1084] System Operation
[1085] Initial settings and user information registration
[1086] A user accesses the system using a user terminal and creates a new account. Specifically, they enter basic information such as name, age, health status, and travel preferences, and then press the register button. At this stage, the server receives the user information, validates the input data, and stores the information in a database. This stored information is used to provide a customized travel experience in the future.
[1087] Selecting a travel destination
[1088] After logging in, the user accesses the travel destination selection screen and selects a destination. The user selects the place they want to go from a world map or a list of travel destinations, and once the selection is complete, the information is sent to the server. The server records this selection information in a database.
[1089] Generate a scenario
[1090] Based on the information about the selected travel destination, the server sends the information to the text generation AI. The text generation AI generates a travel scenario based on this information. For example, it generates information such as "The history of the Eiffel Tower" and "The location where the Mona Lisa is exhibited." The generated scenario is saved in a database by the server.
[1091] Video generation
[1092] The server requests the video generation AI to generate a video based on the generated scenario. The video generation AI generates a 3D model and on-site video in accordance with the scenario. For example, it generates video of the "illuminated Eiffel Tower" or "exhibits at the Louvre Museum." This video is also stored in a database by the server.
[1093] Providing a user experience
[1094] The user device provides a screen that displays the generated scenario and video to the user, allowing the user to experience a simulated trip. For example, a realistic travel experience can be provided, in which a video of the Eiffel Tower is played while a text about its detailed history is read.
[1095] Get more information
[1096] When a user selects a specific location or event within a video or scenario, they can request additional details. Based on this request, the server again asks the text generation AI to generate more information. For example, specific information such as "the structure of the Eiffel Tower" or "episodes from its construction" is generated. The user's device displays this information to the user, providing supplementary information in real time.
[1097] Examples and prompts
[1098] As a concrete example, consider the case where a user selects "Paris" as a travel destination and wants to search for more information about the Eiffel Tower. An example prompt sentence is:
[1099] Generate Scenario prompt:
[1100] Prompt: What is the detailed history of the Eiffel Tower in Paris and what are its popular tourist attractions?
[1101] Video generation prompt:
[1102] Prompt: Generate a video that includes the Eiffel Tower lit up and a museum exhibit.
[1103] summary
[1104] This system allows users to enjoy a realistic travel experience from the comfort of their own home or hospital room. Furthermore, by using generative AI models and prompts, it is possible to customize the system to meet the needs of each individual user, providing a high-quality simulated travel experience.
[1105] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1106] Step 1: Register user information
[1107] A user accesses the system from their terminal and creates a new account. They enter their name, age, health condition, travel preferences, etc., and press the "Register" button. The server receives this information and performs validation (checks the validity of the data). If the input data is correct, the server saves the information in the database. This registers the user information.
[1108] Input: User's name, age, health status, travel preferences
[1109] Data processing / data calculation: User information validation
[1110] Output: Registered user information
[1111] Step 2: Choose your destination
[1112] The user logs in to the system and accesses the travel destination selection screen. The user selects the place they want to go to from a world map or a list of travel destinations and presses the "Select" button. The server receives the selected travel destination information and records it in the database.
[1113] Input: Travel destination selection information
[1114] Data processing / data calculation: Recording selected travel destination information
[1115] Output: Recorded travel destination information
[1116] Step 3: Generate a scenario
[1117] The server retrieves the selected travel destination information from the database and sends the information to the text generation AI based on this. The text generation AI processes the generation prompt and generates a detailed travel scenario. For example, it sends a prompt sentence such as "Prompt: Please tell me the detailed history of the Eiffel Tower in Paris and its popular tourist attractions." The server receives the generated scenario and stores it in the database.
[1118] Input: Selected travel destination information
[1119] Data processing / data calculation: Travel scenario generation
[1120] Output: Generated travel scenarios
[1121] Step 4: Generate footage
[1122] The server requests the video generation AI to generate a video based on the generated scenario. A prompt based on the scenario is sent to the video generation AI, for example, "Prompt: Generate a video that includes the illuminated Eiffel Tower and museum exhibits." The video generation AI generates a 3D model and on-site video based on the scenario, and the server stores the generated video in a database.
[1123] Input: Generated travel scenarios
[1124] Data processing / data calculation: Video generation based on travel scenarios
[1125] Output: Generated video
[1126] Step 5: Delivering the user experience
[1127] The device provides the user with a screen that displays the generated scenarios and videos. The user can enjoy the travel scenarios and videos through the device. For example, the user can play a video of the Eiffel Tower while reading a text about its history.
[1128] Input: Travel scenario and footage
[1129] Data processing / data calculation: Scenario and video display
[1130] Output: Travel scenario and video displayed
[1131] Step 6: Get more information
[1132] The user requests more information about a specific location or event in a video or scenario. For example, if the user wants more information about the structure of the Eiffel Tower, they click on that part. The server receives the request and again asks the text generation AI to generate the detailed information. The text generation AI generates the detailed information based on the specified prompt, and the server sends it to the device. The device then displays the generated detailed information to the user.
[1133] Input: Request for more information
[1134] Data processing / data calculation: generating detailed information based on demand
[1135] Output: Detailed information generated
[1136] In this way, at each processing step, appropriate processing is performed based on the input data, and the results are linked to the next step, providing the user with a realistic and customized travel experience.
[1137] (Application example 1)
[1138] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1139] In recent years, there has been a demand for systems that allow people who cannot travel to simulate travel experiences from their homes, hospital rooms, etc. There is also a need for systems that allow users to enjoy shopping at travel destinations using virtual environments. However, these systems have not yet become widespread, making it difficult for users to simultaneously enjoy realistic travel and shopping experiences. Therefore, the present invention aims to provide a system that integrates simulated travel experiences with virtual shopping, thereby enabling users to have a more realistic and fulfilling experience.
[1140] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1141] In this invention, the server includes a selection means for allowing a user to select a travel destination they wish to go to, a text generation means for generating a scenario related to the selected travel destination, a video generation means for generating a video of the travel destination based on the generated scenario, a provision means for providing the generated scenario and video to the user, a request means for allowing the user to request detailed information about a specific place based on the scenario or video, a text generation means for generating the detailed information based on the request, a virtual shopping means for allowing the user to purchase products in the virtual environment, and a delivery means for delivering the products selected in the virtual environment to the real world. This allows the user to enjoy a virtual trip while shopping from home or a hospital room.
[1142] "User" refers to any individual or entity that uses the System.
[1143] "Destination" refers to a virtual or real geographic location that a user wishes to visit.
[1144] "Selection means" refers to an interface or tool that allows a user to select a travel destination they wish to visit.
[1145] "Text generation means" refers to an artificial intelligence or algorithm that automatically generates a scenario based on specific information.
[1146] "Video generation means" refers to artificial intelligence or a method that automatically creates video based on a generated scenario.
[1147] "Providing means" refers to a technical method or device for displaying or playing the generated scenario and video to the user.
[1148] "Request means" refers to an interface or tool that allows a user to request detailed information about a specific location based on a scenario or video.
[1149] "More Information" refers to additional and specific information about a particular location or event.
[1150] "Virtual shopping vehicle" refers to an interface or tool that allows a user to select and purchase items within a virtual environment.
[1151] "Delivery Method" refers to a method or service for delivering a product selected in a virtual environment to a user in the real world.
[1152] A specific embodiment for carrying out the present invention is described below. This system is designed to provide an integrated experience that allows users to enjoy virtual travel and virtual shopping. The details are described below.
[1153] 1. Overall system configuration
[1154] This system mainly consists of a user terminal, a server, a generative AI model, a database, and a delivery service.
[1155] 1.1 User terminal
[1156] The user terminal is a device that allows users to select travel destinations, view scenarios and videos, and even conduct virtual shopping. Specifically, this applies to smartphones and head-mounted displays.
[1157] 1.2 Server
[1158] The server is a back-end system that generates, manages, and provides user information, travel scenarios, videos, and shopping information. It has the following means:
[1159] 1. Selection means: Provides an interface for the user to select the travel destination they wish to go to.
[1160] 2. Text generation method: Use a generative AI model (e.g., OpenAI's GPT-3.5) to generate travel destination scenarios.
[1161] 3. Video generation method: A generative AI model for generating video based on a scenario.
[1162] 4. Provision means: The generated scenario and video are provided to the user.
[1163] 5. Request means: Provides an interface for users to request more information.
[1164] 6. Virtual Shopping Vehicle: Provides an interface for users to purchase products within a virtual environment.
[1165] 7. Delivery Method: Providing a service for delivering the items selected in the virtual environment in the real world.
[1166] 1.3 Database
[1167] The database is a storage system for storing user information, travel scenarios, videos, and product information.
[1168] 1.4 Delivery Services
[1169] The delivery service arranges for the items purchased by the user through virtual shopping to be delivered in the real world.
[1170] 2. Processing Overview
[1171] 2.1 User Registration and Settings
[1172] A user accesses the system using a user terminal and creates a new account. The server receives the user information, validates it, and stores it in the database.
[1173] 2.2 Choosing a travel destination
[1174] After logging in, users select the places they want to virtually visit from a world map or a list of destinations. The server receives the selected travel destination information and records the user's selection.
[1175] 2.3 Scenario and image generation
[1176] The server transmits information about the selected travel destination to the text generation means, receives the generated scenario, and stores it in a database.The server then transmits the scenario to the video generation means and receives the generated video.
[1177] 2.4 Virtual Shopping
[1178] Users can access the virtual store in the generated video and purchase products, which are then delivered to the real world using a delivery method.
[1179] 3. Examples of concrete examples and prompts
[1180] For example, if the user selects "Paris," the system performs the following process:
[1181] 1. The scenario generation AI generates a scenario that includes information about the Eiffel Tower and the Louvre Museum.
[1182] 2. Image generation AI generates images of actual Paris cityscapes and tourist attractions.
[1183] 3. Users can purchase limited edition keychains and delicious local chocolates at the Eiffel Tower's virtual store.
[1184] Prompt Sentence Examples
[1185] "You're currently on a virtual tour of Paris. See the beautifully lit Eiffel Tower and shop for exclusive keychains and delicious local chocolates at a nearby souvenir shop."
[1186] This system allows users to enjoy virtual travel and shopping from the comfort of their own home or hospital room.
[1187] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1188] Step 1:
[1189] A user accesses the system using a terminal and creates a new account. The user enters basic information such as name, age, and travel preferences, and presses the register button. The basic information is sent from the terminal to the server.
[1190] Input: Basic information you enter into your device, such as your name, age, and travel preferences.
[1191] Data processing: The server validates basic information.
[1192] Output: The server saves the basic information that passes validation to the database.
[1193] Step 2:
[1194] The user accesses the login screen and logs in using the authentication method. The server checks the user's authentication information against the database and provides a dedicated page for the authenticated user.
[1195] Input: User credentials (username, password).
[1196] Data calculation: The server performs authentication by checking against information in the database.
[1197] Output: Authenticated users are served a dedicated page.
[1198] Step 3:
[1199] The user accesses the travel destination selection screen from a dedicated page and selects the desired travel destination from a map or list. The selected information is sent from the device to the server.
[1200] Input: Information about the travel destination selected by the user.
[1201] Data processing: The server records the selected travel destination information.
[1202] Output: Save the travel destination information to the database.
[1203] Step 4:
[1204] The server transmits information about the selected travel destination to the text generation means to generate a scenario, and the text generation means generates a scenario about the travel destination using the generative AI model.
[1205] Input: Travel destination information.
[1206] Data calculation: A generative AI model generates scenarios based on travel destination information.
[1207] Output: The server receives the generated scenario and stores it in a database.
[1208] Step 5:
[1209] The server transmits the generated scenario to the video generation means, which generates a video. The video generation means generates a video based on the scenario using a generative AI model.
[1210] Input: The generated scenario.
[1211] Data calculation: The generative AI model generates images based on the scenario.
[1212] Output: The server receives the generated video and stores it in a database.
[1213] Step 6:
[1214] The user uses the terminal to play back the generated scenario and video, and the video and scenario are displayed on the user terminal via the providing means.
[1215] Input: Scenarios and footage stored in a database.
[1216] Data processing: The server sends the scenario and video to the user terminal via the provision means.
[1217] Output: The user reads the scenario and watches the video.
[1218] Step 7:
[1219] The user accesses the virtual store in the video and performs virtual shopping. Product information selected by the user is sent from the terminal to the server.
[1220] Input: Product information selected by the user.
[1221] Data processing: The server records the product information and processes the purchase through the virtual shopping means.
[1222] Output: Save product information after purchase completion to the database.
[1223] Step 8:
[1224] The server requests a delivery means to deliver the product selected in the virtual environment and arranges for delivery in the real world.
[1225] Input: Product information selected and purchased by the user.
[1226] Data calculation: The server works with the delivery means to carry out the delivery procedure.
[1227] Output: The product is arranged to be delivered to the address specified by the user.
[1228] Through the above steps, a system is realized in which a user can enjoy virtual travel, do virtual shopping, and receive purchased products in the real world.
[1229] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1230] This invention is a system that allows users to virtually experience travel destinations that they cannot visit, and by combining it with an emotion engine that recognizes the user's emotions, it provides a more personalized experience. The main components of the system of this invention are as follows: a user terminal, a server, a text generation AI, a video generation AI, an emotion engine, and a database.
[1231] Overall system configuration
[1232] Main components of the system
[1233] 1. User terminal: A device (PC, tablet, smartphone, etc.) on which the user selects a travel destination and watches scenarios and videos.
[1234] 2. Server: A back-end system that generates, manages, and provides user information, travel scenarios, and videos.
[1235] 3. Text generation AI: An AI that generates travel scenarios based on selected travel destinations.
[1236] 4. Video generation AI: AI that generates videos of travel destinations based on generated scenarios.
[1237] 5. Emotion engine: A function that recognizes the user's emotions and optimizes scenarios and images according to those emotions.
[1238] 6. Database: A storage system that stores user information, travel scenarios, videos, and emotion data.
[1239] System processing: Example
[1240] Initial settings and user information registration
[1241] 1. A user accesses the system using a user terminal and creates a new account. The user enters information such as name, age, health condition, and travel preferences, and presses the registration button.
[1242] 2. The server receives the input information, validates it, and then saves the information to the database.
[1243] 3. The device displays a "Registration complete" message to the user and redirects them to the login screen.
[1244] Selecting a travel destination
[1245] 4. After logging in, the user accesses the travel destination selection screen and chooses the place they want to go.
[1246] 5. The server receives the selection information and records the details of the selected travel destination in a database.
[1247] Scenario and video generation
[1248] 6. The server sends the selected travel destination information to the text generation AI and generates a travel scenario.
[1249] 7. The text generation AI generates a scenario for the selected travel destination and sends it to the server.
[1250] 8. The server sends a video generation request to the video generation AI based on the generated scenario.
[1251] 9. The video generation AI generates video based on the scenario and sends it to the server.
[1252] 10. The server stores the generated scenario and video in a database.
[1253] Emotion Engine Operation
[1254] 11. The terminal uses the camera and microphone built into the user's device to transmit the user's emotions to the emotion engine.
[1255] 12. The emotion engine analyzes the user's facial expressions, voice, and movements in real time to generate emotion data.
[1256] 13. Based on the emotion data received from the emotion engine, the server sends a request to the text generation AI and video generation AI again to adjust the optimal scenario and video.
[1257] 14. The server provides the adjusted scenario and video to the user's terminal.
[1258] Providing a user experience
[1259] 15. The terminal provides a screen that displays the adjusted video and scenario to the user.
[1260] 16. As the user plays the video and reads the scenario text, the experience is further customized based on feedback from the emotion engine.
[1261] The system recognizes users' emotions in real time and can optimize their travel experience accordingly. For example, if a user is moved by a video of the Eiffel Tower, the system automatically generates and provides related videos and scenarios to further enhance the emotion. In this way, users can transcend physical constraints and enjoy a richer, more personalized travel experience.
[1262] The processing flow will be explained below.
[1263] Step 1:
[1264] A user accesses the system using a user terminal and creates a new account. The user enters information such as name, age, health status, and travel preferences, and presses the registration button.
[1265] Step 2:
[1266] The server receives the input information, validates it, checks for errors, and saves the correct information to the database.
[1267] Step 3:
[1268] The device will display a "Registration Complete" message to the user and redirect them to the login screen.
[1269] Step 4:
[1270] The user enters their user ID and password on the login screen and clicks the login button.
[1271] Step 5:
[1272] The server authenticates the user, checking against information in a database and authorizing if there is a match.
[1273] Step 6:
[1274] The device displays a destination selection screen, providing the user with a world map and a list of destinations.
[1275] Step 7:
[1276] The user selects the desired travel destination and presses the select button.
[1277] Step 8:
[1278] The server receives the user's selection information and generates a request to send details of the selected travel destination to a text generation AI.
[1279] Step 9:
[1280] A text generation AI generates a travel scenario based on the selected travel destination, including information on the history, culture, and key attractions of the tourist destination.
[1281] Step 10:
[1282] The server receives the generated scenario and stores it in a database.
[1283] Step 11:
[1284] The server uses the saved scenario information to send a video generation request to the video generation AI.
[1285] Step 12:
[1286] Based on the scenario, video generation AI generates video footage of the travel destination, including real-time scenery and 3D models.
[1287] Step 13:
[1288] The server receives the generated video and stores it in a database.
[1289] Step 14:
[1290] The terminal displays a screen on which the generated video and scenario can be viewed by the user.
[1291] Step 15:
[1292] The user plays the video, begins reading the scenario text, and clicks on specific locations or events that interest them.
[1293] Step 16:
[1294] The terminal records the user's facial expressions and voice through the camera and microphone built into the user device and transmits them to the emotion engine.
[1295] Step 17:
[1296] The emotion engine analyzes the user's facial expressions, voice, and movements in real time to generate emotion data, which includes the user's emotional state, such as joy, surprise, or boredom.
[1297] Step 18:
[1298] Based on the emotional data received from the emotion engine, the server sends requests to the text generation AI and video generation AI to adjust the optimal scenario and video.
[1299] Step 19:
[1300] The text generation AI and video generation AI generate adjusted scenarios and images based on the emotional data and send them to the server.
[1301] Step 20:
[1302] The server stores the adjusted scenario and video in a database to provide to the user.
[1303] Step 21:
[1304] The device displays the adjusted new scenario and video to the user. For example, if the user is moved, additional video and scenarios will be added to further deepen the user's emotion.
[1305] Based on the above processing steps, a system that provides users with a personalized travel experience is realized, allowing users to enjoy a realistic and rich travel experience optimized according to their emotions, even from the comfort of their own home or hospital room.
[1306] Example 2
[1307] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1308] Conventional virtual travel systems have had difficulty personalizing the user experience. They also lacked the ability to adjust the travel experience in real time to reflect the user's emotions. This resulted in problems such as users becoming bored with the experience midway through and their satisfaction decreasing.
[1309] The specification process by the specification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a selection means for selecting a place the user wants to go to, a text generation means for generating a scenario related to the selected place, a video generation means for generating a video of the place based on the generated scenario, a provision means for providing the generated scenario and video to the user, a request means for the user to request detailed information about the destination based on the scenario or video, a text generation means for generating the detailed information based on the request, an emotion recognition means for recognizing the user's emotion, and an adjustment means for adjusting the scenario and video based on the recognized emotion data. This allows the user's emotions to be reflected in real time, enabling a more personalized and rich travel experience.
[1310] The "selection means" is a function that allows the user to select the place they want to go.
[1311] The "text generation means" is a function for generating scenarios and detailed information about a selected location.
[1312] The "video generation means" is a function for generating video of a location based on the generated scenario.
[1313] The "provision means" is a function for providing the generated scenario and video to the user.
[1314] The "request means" is a function that allows a user to request detailed information about a customer based on a scenario or image.
[1315] The "emotion recognition means" is a function for recognizing the user's emotions.
[1316] The "adjustment means" is a function for adjusting the scenario and video based on the recognized emotion data.
[1317] The "registration means" is a function for registering basic information about a user.
[1318] The "storage means" is a function for storing the registered basic information.
[1319] "Authentication means" is a function that allows a user to access a system.
[1320] The "page providing means" is a function for providing a page exclusively for an authenticated user.
[1321] MODE FOR CARRYING OUT THE INVENTION
[1322] This invention is a system that allows users to virtually experience travel destinations that are difficult to visit in person. This system generates personalized travel experiences based on the user's individual requests, and by incorporating emotion recognition technology, it provides an optimal experience that responds to the user's real-time emotions.
[1323] System Components
[1324] 1. User terminal: A device (e.g., PC, tablet, smartphone) that provides an interface for users to select the destination and experience the virtual journey.
[1325] 2. Server: A back-end system that generates, manages, and provides user information, travel scenarios, and videos.
[1326] 3. Text Generation AI: An artificial intelligence technology that generates travel scenarios based on the locations selected by the user.
[1327] 4. Video Generation AI: Artificial intelligence technology that generates footage of a selected location based on a generated scenario.
[1328] 5. Emotion Recognition Engine: An engine that recognizes user emotions in real time and optimizes the travel experience based on the results.
[1329] 6. Database: A storage system that stores user information, travel scenarios, video, and emotion data.
[1330] System Operation Overview
[1331] When a user accesses the system and selects a place they want to go, the server first retrieves information about the selected place from the database and sends it to a text generation AI. The text generation AI generates a detailed travel scenario based on the selected place and sends the scenario back to the server. The server then sends the generated scenario to a video generation AI, which generates a video of the place based on it.
[1332] The generated scenarios and videos are stored on a server and provided to the user's device. The user can view the scenarios and videos through their device. The emotion recognition engine also uses the user's camera and microphone to analyze emotions from facial expressions and voice in real time and sends the data to the server.
[1333] Optimizing experiences through emotions
[1334] When the server receives the emotion data sent from the emotion recognition engine, it optimizes the travel experience based on that data. For example, if the user is moved, related videos and scenarios are automatically generated to further deepen the emotion, enriching the user's experience. This series of processes is repeated in real time, ensuring that the user always receives an optimized travel experience.
[1335] Specific examples
[1336] For example, if the user selects "The Eiffel Tower in France," the system will send the following prompt to the generative AI model:
[1337] "Generate a scenario in which the user visits the Eiffel Tower in France. The user is a woman in her 20s who likes sightseeing."
[1338] Based on this prompt, the text generation AI generates a detailed scenario about the Eiffel Tower (history, tourist attractions, etc.), and the video generation AI then generates a 360-degree panoramic video based on that scenario.
[1339] In this way, the present invention can reflect the user's emotions in real time and provide a personalized and enriched travel experience.
[1340] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1341] Step 1:
[1342] A user accesses the system using their own device (PC, tablet, smartphone). Using a web browser or a dedicated app, the user enters basic information such as name, age, health status, and travel preferences on the sign-up screen and clicks the "Register" button (input: user's basic information). This information is then sent to the server (output: sending user's basic information).
[1343] Step 2:
[1344] The server receives the basic information sent by the user and performs a validation check to ensure that the format and content are correct (input: basic user information). If validation is successful, this information is saved in the database (output: saved user information). If validation fails, a message is generated to prompt the user to re-enter the invalid information and sent to the user's terminal (output: message prompting re-entry).
[1345] Step 3:
[1346] The terminal receives the "Registration Complete" message from the server and displays it to the user (input: registration complete message). At the same time, it redirects the user to the login screen (output: display of login screen).
[1347] Step 4:
[1348] The user logs in using their account information. After logging in, the user accesses the destination selection screen and chooses the place they want to go (input: login information and destination selection).
[1349] Step 5:
[1350] The server authenticates the user's login information and receives the selected travel destination information (input: login information and travel destination selection). The server records the selected travel destination information in the database and prepares for the next scenario generation (output: saving travel destination information).
[1351] Step 6:
[1352] The server sends travel destination information to the text generation AI and requests it to generate a detailed travel scenario (input: travel destination information) (example prompt: "Please generate a scenario for visiting the Eiffel Tower in France. The user is a woman in her twenties who enjoys sightseeing."). The text generation AI generates a detailed scenario based on the prompt and sends the scenario back to the server (output: generated travel scenario).
[1353] Step 7:
[1354] The server receives the generated scenario and sends it to the video generation AI as a video generation request (input: travel scenario). The video generation AI generates a video of the travel destination based on the scenario and sends it back to the server (output: generated video).
[1355] Step 8:
[1356] The server stores the generated scenario and video in a database (input: generated scenario and video), and then provides the scenario and video to the user's device (output: providing scenario and video).
[1357] Step 9:
[1358] The device sends facial expressions and voice data collected through the user's camera and microphone to the emotion recognition engine in real time (input: facial and voice data). The emotion recognition engine analyzes the data, generates the user's emotion data, and sends it to the server (output: emotion data).
[1359] Step 10:
[1360] The server receives the emotion data and sends a request to the text generation AI and video generation AI again to adjust the optimal scenario and video based on that data (input: emotion data). The generated new scenario and video are sent back to the server (output: adjusted scenario and video).
[1361] Step 11:
[1362] The server provides the adjusted scenario and video back to the user's device (input: adjusted scenario and video), allowing the user to receive a more personalized and enriched travel experience (output: re-provision of scenario and video).
[1363] The above is the specific operation and data processing flow in each processing step of this system.
[1364] (Application example 2)
[1365] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1366] Conventional virtual shopping experience systems allow users to simply browse products in a virtual space, making it difficult to provide a personalized experience based on the user's emotions and preferences. Furthermore, they lack the ability to identify user emotions in real time and optimize the content provided based on those emotions. Furthermore, it is difficult for users to experience the same realistic shopping experience as in a physical store from the comfort of their own home.
[1367] The specification processing by the specification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes a selection means for the user to select a destination they wish to go to, a text generation means for generating a scenario related to the selected destination, a video generation means for generating a video of the destination based on the generated scenario, a provision means for providing the generated scenario and video to the user, a request means for the user to request detailed information about a specific place based on the scenario or video, a text generation means for generating detailed information based on the request, an emotion recognition means for identifying the user's emotion and optimizing the scenario and video to be provided based on the emotion, and an optimization means for regenerating the scenario and video to be provided based on the generated emotion data. This allows the user to have a realistic shopping experience as if they were visiting a physical store, even from the comfort of their own home, and further enables them to enjoy personalized services based on emotion recognition.
[1368] "User" refers to an individual who uses the system.
[1369] "Destination" refers to a virtual location selected by the user that the user wants to visit.
[1370] The "selection means" refers to a means for the user to select the place they want to go.
[1371] "Text generation means" refers to a means for automatically generating a scenario related to a selected destination.
[1372] "Video generation means" refers to means for generating video based on a generated scenario.
[1373] "Providing means" refers to a means for providing the generated scenario and video to the user.
[1374] The "request means" refers to a means for a user to request detailed information about a specific location based on a scenario or video.
[1375] The "emotion recognition means" refers to a means for identifying a user's emotions in real time and optimizing the scenarios and videos provided based on those emotions.
[1376] The "optimization means" refers to a means for regenerating the scenario and video to be provided based on the generated emotion data, and optimizing the user experience.
[1377] "Basic information" refers to personal information registered in the system, such as the user's health status and travel preferences.
[1378] "Registration means" refers to a means for registering basic information of a user in the system.
[1379] "Storage means" refers to the means for storing registered basic information.
[1380] "Authentication means" refers to a means for authenticating a user when logging into a system.
[1381] The "page providing means" refers to a means for providing a page exclusively for an authenticated user.
[1382] "Virtual visit means" refers to a means for allowing a user to progress through the experience as if they had visited the destination.
[1383] The "personalized provision means" refers to a means for providing more personalized products and information to a user based on the emotion identified by the emotion recognition means.
[1384] This invention is a system that allows users to virtually visit a virtual store and have a shopping experience. Furthermore, by combining it with an emotion recognition engine, it provides a personalized experience based on the user's emotions.
[1385] System Components
[1386] 1. User Device:
[1387] This is a device that allows users to select a virtual store and view scenarios and videos. When using a smartphone, a dedicated application is installed. This application is developed using React Native.
[1388] 2. Server:
[1389] This is the backend system that generates, manages, and provides user information, scenarios, and videos. The server uses AWS EC2 or GCP. The following main components are placed on the server:
[1390] Text generation AI (OpenAI GPT-4)
[1391] Video generation AI (DALL-E or other 3D video generation models)
[1392] Emotion Engine (Affectiva SDK)
[1393] Database (MySQL or PostgreSQL)
[1394] Program processing
[1395] 1. Initial settings and user information registration
[1396] A user installs the application and creates a new account, entering personal information such as name, age, preferences, etc., which is then saved to a database. This process is done on the front-end (React Native) and back-end (Node.js / Express).
[1397] 2. Select a store
[1398] After logging in, the user selects the virtual store they want to visit. The store list is retrieved from the server via API communication. The API uses REST or GraphQL.
[1399] 3. Scenario and video generation
[1400] Based on the information of the selected store, the text generation AI (GPT-4) generates a guidance scenario. The generated scenario is sent to the video generation AI (DALL-E), which generates a video of the virtual store. This processing is performed using the OpenAI API and serverless processing (AWS Lambda).
[1401] 4. Operation of the Emotion Engine
[1402] The camera and microphone on the user's device (smartphone) are used to transmit the user's facial expressions and voice to the emotion engine (Affectiva SDK). Emotional data is sent to the server in real time, and the scenario and video are optimized by text generation AI and video generation AI.
[1403] 5. Providing a user experience
[1404] The optimized scenarios and videos are displayed on the user's device, and based on emotion recognition, product listings and custom messages that pique the user's interest are dynamically updated to provide a personalized shopping experience.
[1405] Examples of concrete examples and prompts
[1406] Prompt for text generation AI:
[1407] Enter the user's age, preferences, and selected store information:
[1408] Age: 30
[1409] Interests: Fashion, accessories
[1410] Selected store: High-end fashion mall in the city
[1411] Use this information to generate product lists and scenarios that may be of interest to your users.
[1412] Prompts for video generation AI:
[1413] Generate a 3D image of the interior of a high-end fashion mall in the city, highlighting the product shelves that users are most interested in.
[1414] This system allows users to transcend physical constraints and enjoy a rich and personalized virtual shopping experience. The emotion recognition engine provides optimal content according to the user's emotions, ensuring a highly satisfying experience.
[1415] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1416] Step 1: Initial settings and user information registration
[1417] Input: After users install the application, they enter personal information such as their name, age, preferences, etc.
[1418] What happens: When a user fills out a form and clicks the "Submit" button, the information is sent from the frontend (React Native) to the backend (Node.js / Express), where it is validated and saved to a MySQL or PostgreSQL database.
[1419] Output: The user's personal information is saved in the database and a "Successful registration" message is displayed.
[1420] Step 2: Select a store
[1421] Input: After the user logs in, they select the virtual store they want to visit.
[1422] Specific operation: The store list screen is displayed on the front end, and the user selects the desired store. The selection information is sent to the server, and the information of the selected store is saved in the database.
[1423] Output: The store information selected by the user is saved on the server and used in the next step.
[1424] Step 3: Generate a scenario
[1425] Input: Selected store information.
[1426] Specific operation: The server generates a prompt sentence based on the selected store information and sends it to OpenAI's GPT-4 API. The text generation AI (GPT-4) generates a store guide scenario and returns the result to the server.
[1427] Output: The scenario text generated by GPT-4 is saved on the server.
[1428] Step 4: Generate footage
[1429] Input: Scenario text generated by GPT-4.
[1430] Specific operation: The server generates prompt sentences based on the scenario text and sends them to the video generation AI (DALL-E or other 3D video generation model). The video generation AI generates images of the virtual store based on the scenario and sends the results back to the server.
[1431] Output: The video generated by DALL-E is stored on the server.
[1432] Step 5: Emotion Recognition in Action
[1433] Input: User's facial and voice data.
[1434] Specific operation: The camera and microphone on the user device (smartphone) capture facial expressions and voice, and send them to the emotion recognition engine (Affectiva SDK). The emotion engine analyzes the data, generates emotion data, and sends it to the server.
[1435] Output: The emotion data generated by the emotion recognition engine is stored on the server.
[1436] Step 6: Scenario and footage optimization
[1437] Input: Emotion data, initial scenario and video.
[1438] Specific operation: The server sends a request to the text generation AI and video generation AI again to optimize the scenario and video generated based on the emotion data. As a result, the scenario and video are regenerated based on the user's emotions.
[1439] Output: The new optimized scenario and footage are saved on the server.
[1440] Step 7: Delivering the user experience
[1441] Input: Optimized scenario and footage.
[1442] Specific operation: The user device retrieves the optimized scenario and video from the server and provides it to the user. Product lists and custom messages based on the user's emotion are dynamically displayed.
[1443] Output: The user watches the optimized scenario and video, and enjoys a personalized shopping experience.
[1444] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1445] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1446] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1447] [Fourth embodiment]
[1448] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1449] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1450] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1451] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1452] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1453] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1454] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1455] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1456] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1457] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1458] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1459] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1460] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1461] The present invention provides a system that enables people who cannot travel to experience a simulated trip from their homes, hospital rooms, etc. This system generates realistic travel scenarios and images based on the travel destinations selected by the user and provides them to the user.
[1462] Overall system configuration
[1463] The system mainly consists of the following components:
[1464] 1. User terminal: A device (PC, tablet, smartphone, etc.) on which the user selects a travel destination and views scenarios and footage.
[1465] 2. Server: A back-end system that generates, manages, and provides user information, travel scenarios, and videos.
[1466] 3. Text generation AI: AI that generates scenarios based on information about the selected travel destination.
[1467] 4. Video generation AI: AI that generates video based on a generated scenario.
[1468] 5. Database: A storage system that stores user information, travel scenarios, and videos.
[1469] System processing: Example
[1470] Initial settings and user information registration
[1471] 1. A user accesses the system using a user terminal and creates a new account.
[1472] Enter your name, age, health status, travel preferences, etc., and press the register button.
[1473] 2. The server receives the user information, validates it, and saves it in the database.
[1474] Selecting a travel destination
[1475] 3. After logging in, the user accesses the travel destination selection screen.
[1476] Choose where you want to go from a world map or a list of travel destinations.
[1477] 4. The server receives the selected destination information and records the user's selection.
[1478] Generate a scenario
[1479] 5. The server sends the selected travel destination information to the text generation AI.
[1480] The text generation AI uses that information to generate travel scenarios, such as the history of the Eiffel Tower or where the Mona Lisa is located.
[1481] 6. The server receives the generated scenario and stores it in the database.
[1482] Video generation
[1483] 7. The server requests the video generation AI to generate a video based on the generated scenario.
[1484] The video generation AI generates 3D models and on-site footage based on the scenario.
[1485] For example, the illuminated Eiffel Tower in Paris and the exhibits at the Louvre Museum.
[1486] Providing a user experience
[1487] 8. The terminal provides the user with a screen that displays the generated scenario and images.
[1488] The user can play the video and read the scenario text.
[1489] For example, you can watch a video of the Eiffel Tower while reading its detailed history.
[1490] Get more information
[1491] 9. When the user selects a specific location or event within the video or scenario, more information is requested.
[1492] 10. The server again asks the text generation AI to generate detailed information based on the request.
[1493] For example, the structure of the Eiffel Tower and stories from its construction.
[1494] 11. The terminal displays the generated details to the user.
[1495] Supplementary information is displayed in real time in sync with the video.
[1496] In this way, users can enjoy a realistic travel experience from their own homes or hospital rooms. This system allows users to enjoy the joy of travel regardless of physical constraints or external factors. Furthermore, the system allows for advanced customization according to the individual needs of each user, providing an optimal experience for each individual user.
[1497] The processing flow will be explained below.
[1498] Step 1:
[1499] A user accesses the system and creates a new account by entering information such as their name, age, health status, and travel preferences into the terminal and pressing the register button.
[1500] Step 2:
[1501] The server receives the input information, validates it, checks for errors, and saves the correct information to the database.
[1502] Step 3:
[1503] The device will display a "Registration Complete" message to the user and redirect them to the login screen.
[1504] Step 4:
[1505] The user enters their user ID and password on the login screen and clicks the login button.
[1506] Step 5:
[1507] The server authenticates the user, checking against information in a database and authorizing if there is a match.
[1508] Step 6:
[1509] The device displays a destination selection screen, providing the user with a world map and a list of destinations.
[1510] Step 7:
[1511] The user selects the desired travel destination and presses the select button.
[1512] Step 8:
[1513] The server receives the user's selection information and generates a request to send details of the selected travel destination to a text generation AI.
[1514] Step 9:
[1515] A text generation AI generates a travel scenario based on the selected travel destination, including information on the history, culture, and key attractions of the tourist destination.
[1516] Step 10:
[1517] The server receives the generated scenario and stores it in a database.
[1518] Step 11:
[1519] The server uses the saved scenario information to send a video generation request to the video generation AI.
[1520] Step 12:
[1521] Based on the scenario, the video generation AI generates footage of the target travel destination, including real-time scenery and 3D models.
[1522] Step 13:
[1523] The server receives the generated video and stores it in a database.
[1524] Step 14:
[1525] The terminal displays a screen on which the generated video and scenario can be viewed by the user.
[1526] Step 15:
[1527] The user plays the video, begins reading the scenario text, and clicks on specific locations or events that interest them.
[1528] Step 16:
[1529] The server receives the user's click information and sends a detailed information generation request back to the text generation AI.
[1530] Step 17:
[1531] Text generation AI generates detailed information, including specific history and unique anecdotes about the location.
[1532] Step 18:
[1533] The server receives the generated details and sends them to the terminal.
[1534] Step 19:
[1535] The device displays detailed information to the user and works in conjunction with the video to enhance the travel experience.
[1536] This series of steps allows users to have a realistic travel experience, as if they were actually visiting the location.
[1537] Example 1
[1538] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1539] In modern society, there are many people who find it difficult to actually travel for various reasons. For example, many people are unable to travel due to health conditions, time constraints, or financial reasons. There is also a need for a way to simulate a trip for elderly people and those who find it difficult to go out due to illness. Furthermore, conventional travel experience systems have difficulty customizing realistic travel scenarios and videos, making it impossible to provide an experience tailored to individual users' needs.
[1540] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1541] In this invention, the server includes a selection means for allowing a user to select a travel destination they wish to go to, a text generation means for generating a scenario related to the selected travel destination, a video generation means for generating a video of the travel destination based on the generated scenario, a provision means for providing the generated scenario and video to the user, a request means for allowing the user to request detailed information about a specific location based on the scenario or video, a text generation means for generating the detailed information based on the request, a display means for displaying the detailed information generated in response to the user's request on a user terminal, and a storage means for registering and saving user information by initial setting. This enables users who have difficulty traveling in person due to health conditions, time constraints, financial reasons, etc. to enjoy a realistic and customized travel experience from the comfort of their own home or hospital room.
[1542] "User" refers to an individual who intends to use the system to have a simulated travel experience.
[1543] "Selection means" refers to an interface or tool that allows a user to select a travel destination they wish to go to.
[1544] "Text generation means" refers to a mechanism that utilizes natural language processing technology to generate scenarios and detailed information about the selected travel destination.
[1545] "Video generation means" refers to the technology and tools used to generate realistic travel destination footage based on the generated scenario.
[1546] "Providing means" refers to the interface or media for providing the generated scenario and video to the user.
[1547] "Request means" refers to an interface or tool that allows a user to request detailed information about a specific location or event based on a scenario or video.
[1548] "Display means" refers to a mechanism for displaying detailed information and images generated in response to a user's request on a user terminal.
[1549] "Storage means" refers to the technology or tools used to register user information in the initial settings and store it in a database, etc.
[1550] MODE FOR CARRYING OUT THE INVENTION
[1551] The present invention relates to a system that allows people who cannot travel to have a simulated real-life travel experience from their homes, hospital rooms, etc. This system generates realistic travel scenarios and images based on travel destinations selected by the user and provides them to the user. The specific configuration and processing flow of the system are described below.
[1552] Overall system configuration
[1553] The system mainly consists of the following components:
[1554] 1. User terminal: A device (PC, tablet, smartphone, etc.) on which the user selects a travel destination and views scenarios and footage.
[1555] 2. Server: A back-end system that generates, manages, and provides user information, travel scenarios, and videos.
[1556] 3. Text generation AI: AI that generates scenarios based on information about the selected travel destination.
[1557] 4. Video generation AI: AI that generates video based on a generated scenario.
[1558] 5. Database: A storage system that stores user information, travel scenarios, and videos.
[1559] System Operation
[1560] Initial settings and user information registration
[1561] A user accesses the system using a user terminal and creates a new account. Specifically, they enter basic information such as name, age, health status, and travel preferences, and then press the register button. At this stage, the server receives the user information, validates the input data, and stores the information in a database. This stored information is used to provide a customized travel experience in the future.
[1562] Selecting a travel destination
[1563] After logging in, the user accesses the travel destination selection screen and selects a destination. The user selects the place they want to go from a world map or a list of travel destinations, and once the selection is complete, the information is sent to the server. The server records this selection information in a database.
[1564] Generate a scenario
[1565] Based on the information about the selected travel destination, the server sends the information to the text generation AI. The text generation AI generates a travel scenario based on this information. For example, it generates information such as "The history of the Eiffel Tower" and "The location where the Mona Lisa is exhibited." The generated scenario is saved in a database by the server.
[1566] Video generation
[1567] The server requests the video generation AI to generate a video based on the generated scenario. The video generation AI generates a 3D model and on-site video in accordance with the scenario. For example, it generates video of the "illuminated Eiffel Tower" or "exhibits at the Louvre Museum." This video is also stored in a database by the server.
[1568] Providing a user experience
[1569] The user device provides a screen that displays the generated scenario and video to the user, allowing the user to experience a simulated trip. For example, a realistic travel experience can be provided, in which a video of the Eiffel Tower is played while a text about its detailed history is read.
[1570] Get more information
[1571] When a user selects a specific location or event within a video or scenario, they can request additional details. Based on this request, the server again asks the text generation AI to generate more information. For example, specific information such as "the structure of the Eiffel Tower" or "episodes from its construction" is generated. The user's device displays this information to the user, providing supplementary information in real time.
[1572] Examples and prompts
[1573] As a concrete example, consider the case where a user selects "Paris" as a travel destination and wants to search for more information about the Eiffel Tower. An example prompt sentence is:
[1574] Generate Scenario prompt:
[1575] Prompt: What is the detailed history of the Eiffel Tower in Paris and what are its popular tourist attractions?
[1576] Video generation prompt:
[1577] Prompt: Generate a video that includes the Eiffel Tower lit up and a museum exhibit.
[1578] summary
[1579] This system allows users to enjoy a realistic travel experience from the comfort of their own home or hospital room. Furthermore, by using generative AI models and prompts, it is possible to customize the system to meet the needs of each individual user, providing a high-quality simulated travel experience.
[1580] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1581] Step 1: Register user information
[1582] A user accesses the system from their terminal and creates a new account. They enter their name, age, health condition, travel preferences, etc., and press the "Register" button. The server receives this information and performs validation (checks the validity of the data). If the input data is correct, the server saves the information in the database. This registers the user information.
[1583] Input: User's name, age, health status, travel preferences
[1584] Data processing / data calculation: User information validation
[1585] Output: Registered user information
[1586] Step 2: Choose your destination
[1587] The user logs in to the system and accesses the travel destination selection screen. The user selects the place they want to go to from a world map or a list of travel destinations and presses the "Select" button. The server receives the selected travel destination information and records it in the database.
[1588] Input: Travel destination selection information
[1589] Data processing / data calculation: Recording selected travel destination information
[1590] Output: Recorded travel destination information
[1591] Step 3: Generate a scenario
[1592] The server retrieves the selected travel destination information from the database and sends the information to the text generation AI based on this. The text generation AI processes the generation prompt and generates a detailed travel scenario. For example, it sends a prompt sentence such as "Prompt: Please tell me the detailed history of the Eiffel Tower in Paris and its popular tourist attractions." The server receives the generated scenario and stores it in the database.
[1593] Input: Selected travel destination information
[1594] Data processing / data calculation: Travel scenario generation
[1595] Output: Generated travel scenarios
[1596] Step 4: Generate footage
[1597] The server requests the video generation AI to generate a video based on the generated scenario. A prompt based on the scenario is sent to the video generation AI, for example, "Prompt: Generate a video that includes the illuminated Eiffel Tower and museum exhibits." The video generation AI generates a 3D model and on-site video based on the scenario, and the server stores the generated video in a database.
[1598] Input: Generated travel scenarios
[1599] Data processing / data calculation: Video generation based on travel scenarios
[1600] Output: Generated video
[1601] Step 5: Delivering the user experience
[1602] The device provides the user with a screen that displays the generated scenarios and videos. The user can enjoy the travel scenarios and videos through the device. For example, the user can play a video of the Eiffel Tower while reading a text about its history.
[1603] Input: Travel scenario and footage
[1604] Data processing / data calculation: Scenario and video display
[1605] Output: Travel scenario and video displayed
[1606] Step 6: Get more information
[1607] The user requests more information about a specific location or event in a video or scenario. For example, if the user wants more information about the structure of the Eiffel Tower, they click on that part. The server receives the request and again asks the text generation AI to generate the detailed information. The text generation AI generates the detailed information based on the specified prompt, and the server sends it to the device. The device then displays the generated detailed information to the user.
[1608] Input: Request for more information
[1609] Data processing / data calculation: generating detailed information based on demand
[1610] Output: Detailed information generated
[1611] In this way, at each processing step, appropriate processing is performed based on the input data, and the results are linked to the next step, providing the user with a realistic and customized travel experience.
[1612] (Application example 1)
[1613] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1614] In recent years, there has been a demand for systems that allow people who cannot travel to simulate travel experiences from their homes, hospital rooms, etc. There is also a need for systems that allow users to enjoy shopping at travel destinations using virtual environments. However, these systems have not yet become widespread, making it difficult for users to simultaneously enjoy realistic travel and shopping experiences. Therefore, the present invention aims to provide a system that integrates simulated travel experiences with virtual shopping, thereby enabling users to have a more realistic and fulfilling experience.
[1615] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1616] In this invention, the server includes a selection means for allowing a user to select a travel destination they wish to go to, a text generation means for generating a scenario related to the selected travel destination, a video generation means for generating a video of the travel destination based on the generated scenario, a provision means for providing the generated scenario and video to the user, a request means for allowing the user to request detailed information about a specific place based on the scenario or video, a text generation means for generating the detailed information based on the request, a virtual shopping means for allowing the user to purchase products in the virtual environment, and a delivery means for delivering the products selected in the virtual environment to the real world. This allows the user to enjoy a virtual trip while shopping from home or a hospital room.
[1617] "User" refers to any individual or entity that uses the System.
[1618] "Destination" refers to a virtual or real geographic location that a user wishes to visit.
[1619] "Selection means" refers to an interface or tool that allows a user to select a travel destination they wish to visit.
[1620] "Text generation means" refers to an artificial intelligence or algorithm that automatically generates a scenario based on specific information.
[1621] "Video generation means" refers to artificial intelligence or a method that automatically creates video based on a generated scenario.
[1622] "Providing means" refers to a technical method or device for displaying or playing the generated scenario and video to the user.
[1623] "Request means" refers to an interface or tool that allows a user to request detailed information about a specific location based on a scenario or video.
[1624] "More Information" refers to additional and specific information about a particular location or event.
[1625] "Virtual shopping vehicle" refers to an interface or tool that allows a user to select and purchase items within a virtual environment.
[1626] "Delivery Method" refers to a method or service for delivering a product selected in a virtual environment to a user in the real world.
[1627] A specific embodiment for carrying out the present invention is described below. This system is designed to provide an integrated experience that allows users to enjoy virtual travel and virtual shopping. The details are described below.
[1628] 1. Overall system configuration
[1629] This system mainly consists of a user terminal, a server, a generative AI model, a database, and a delivery service.
[1630] 1.1 User terminal
[1631] The user terminal is a device that allows users to select travel destinations, view scenarios and videos, and even conduct virtual shopping. Specifically, this applies to smartphones and head-mounted displays.
[1632] 1.2 Server
[1633] The server is a back-end system that generates, manages, and provides user information, travel scenarios, videos, and shopping information. It has the following means:
[1634] 1. Selection means: Provides an interface for the user to select the travel destination they wish to go to.
[1635] 2. Text generation method: Use a generative AI model (e.g., OpenAI's GPT-3.5) to generate travel destination scenarios.
[1636] 3. Video generation method: A generative AI model for generating video based on a scenario.
[1637] 4. Provision means: The generated scenario and video are provided to the user.
[1638] 5. Request means: Provides an interface for users to request more information.
[1639] 6. Virtual Shopping Vehicle: Provides an interface for users to purchase products within a virtual environment.
[1640] 7. Delivery Method: Providing a service for delivering the items selected in the virtual environment in the real world.
[1641] 1.3 Database
[1642] The database is a storage system for storing user information, travel scenarios, videos, and product information.
[1643] 1.4 Delivery Services
[1644] The delivery service arranges for the items purchased by the user through virtual shopping to be delivered in the real world.
[1645] 2. Processing Overview
[1646] 2.1 User Registration and Settings
[1647] A user accesses the system using a user terminal and creates a new account. The server receives the user information, validates it, and stores it in the database.
[1648] 2.2 Choosing a travel destination
[1649] After logging in, users select the places they want to virtually visit from a world map or a list of destinations. The server receives the selected travel destination information and records the user's selection.
[1650] 2.3 Scenario and image generation
[1651] The server transmits information about the selected travel destination to the text generation means, receives the generated scenario, and stores it in a database.The server then transmits the scenario to the video generation means and receives the generated video.
[1652] 2.4 Virtual Shopping
[1653] Users can access the virtual store in the generated video and purchase products, which are then delivered to the real world using a delivery method.
[1654] 3. Examples of concrete examples and prompts
[1655] For example, if the user selects "Paris," the system performs the following process:
[1656] 1. The scenario generation AI generates a scenario that includes information about the Eiffel Tower and the Louvre Museum.
[1657] 2. Image generation AI generates images of actual Paris cityscapes and tourist attractions.
[1658] 3. Users can purchase limited edition keychains and delicious local chocolates at the Eiffel Tower's virtual store.
[1659] Prompt Sentence Examples
[1660] "You're currently on a virtual tour of Paris. See the beautifully lit Eiffel Tower and shop for exclusive keychains and delicious local chocolates at a nearby souvenir shop."
[1661] This system allows users to enjoy virtual travel and shopping from the comfort of their own home or hospital room.
[1662] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1663] Step 1:
[1664] A user accesses the system using a terminal and creates a new account. The user enters basic information such as name, age, and travel preferences, and presses the register button. The basic information is sent from the terminal to the server.
[1665] Input: Basic information you enter into your device, such as your name, age, and travel preferences.
[1666] Data processing: The server validates basic information.
[1667] Output: The server saves the basic information that passes validation to the database.
[1668] Step 2:
[1669] The user accesses the login screen and logs in using the authentication method. The server checks the user's authentication information against the database and provides a dedicated page for the authenticated user.
[1670] Input: User credentials (username, password).
[1671] Data calculation: The server performs authentication by checking against information in the database.
[1672] Output: Authenticated users are served a dedicated page.
[1673] Step 3:
[1674] The user accesses the travel destination selection screen from a dedicated page and selects the desired travel destination from a map or list. The selected information is sent from the device to the server.
[1675] Input: Information about the travel destination selected by the user.
[1676] Data processing: The server records the selected travel destination information.
[1677] Output: Save the travel destination information to the database.
[1678] Step 4:
[1679] The server transmits information about the selected travel destination to the text generation means to generate a scenario, and the text generation means generates a scenario about the travel destination using the generative AI model.
[1680] Input: Travel destination information.
[1681] Data calculation: A generative AI model generates scenarios based on travel destination information.
[1682] Output: The server receives the generated scenario and stores it in a database.
[1683] Step 5:
[1684] The server transmits the generated scenario to the video generation means, which generates a video. The video generation means generates a video based on the scenario using a generative AI model.
[1685] Input: The generated scenario.
[1686] Data calculation: The generative AI model generates images based on the scenario.
[1687] Output: The server receives the generated video and stores it in a database.
[1688] Step 6:
[1689] The user uses the terminal to play back the generated scenario and video, and the video and scenario are displayed on the user terminal via the providing means.
[1690] Input: Scenarios and footage stored in a database.
[1691] Data processing: The server sends the scenario and video to the user terminal via the provision means.
[1692] Output: The user reads the scenario and watches the video.
[1693] Step 7:
[1694] The user accesses the virtual store in the video and performs virtual shopping. Product information selected by the user is sent from the terminal to the server.
[1695] Input: Product information selected by the user.
[1696] Data processing: The server records the product information and processes the purchase through the virtual shopping means.
[1697] Output: Save product information after purchase completion to the database.
[1698] Step 8:
[1699] The server requests a delivery means to deliver the product selected in the virtual environment and arranges for delivery in the real world.
[1700] Input: Product information selected and purchased by the user.
[1701] Data calculation: The server works with the delivery means to carry out the delivery procedure.
[1702] Output: The product is arranged to be delivered to the address specified by the user.
[1703] Through the above steps, a system is realized in which a user can enjoy virtual travel, do virtual shopping, and receive purchased products in the real world.
[1704] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1705] This invention is a system that allows users to virtually experience travel destinations that they cannot visit, and by combining it with an emotion engine that recognizes the user's emotions, it provides a more personalized experience. The main components of the system of this invention are as follows: a user terminal, a server, a text generation AI, a video generation AI, an emotion engine, and a database.
[1706] Overall system configuration
[1707] Main components of the system
[1708] 1. User terminal: A device (PC, tablet, smartphone, etc.) on which the user selects a travel destination and watches scenarios and videos.
[1709] 2. Server: A back-end system that generates, manages, and provides user information, travel scenarios, and videos.
[1710] 3. Text generation AI: An AI that generates travel scenarios based on selected travel destinations.
[1711] 4. Video generation AI: AI that generates videos of travel destinations based on generated scenarios.
[1712] 5. Emotion engine: A function that recognizes the user's emotions and optimizes scenarios and images according to those emotions.
[1713] 6. Database: A storage system that stores user information, travel scenarios, videos, and emotion data.
[1714] System processing: Example
[1715] Initial settings and user information registration
[1716] 1. A user accesses the system using a user terminal and creates a new account. The user enters information such as name, age, health condition, and travel preferences, and presses the registration button.
[1717] 2. The server receives the input information, validates it, and then saves the information to the database.
[1718] 3. The device displays a "Registration complete" message to the user and redirects them to the login screen.
[1719] Selecting a travel destination
[1720] 4. After logging in, the user accesses the travel destination selection screen and chooses the place they want to go.
[1721] 5. The server receives the selection information and records the details of the selected travel destination in a database.
[1722] Scenario and video generation
[1723] 6. The server sends the selected travel destination information to the text generation AI and generates a travel scenario.
[1724] 7. The text generation AI generates a scenario for the selected travel destination and sends it to the server.
[1725] 8. The server sends a video generation request to the video generation AI based on the generated scenario.
[1726] 9. The video generation AI generates video based on the scenario and sends it to the server.
[1727] 10. The server stores the generated scenario and video in a database.
[1728] Emotion Engine Operation
[1729] 11. The terminal uses the camera and microphone built into the user's device to transmit the user's emotions to the emotion engine.
[1730] 12. The emotion engine analyzes the user's facial expressions, voice, and movements in real time to generate emotion data.
[1731] 13. Based on the emotion data received from the emotion engine, the server sends a request to the text generation AI and video generation AI again to adjust the optimal scenario and video.
[1732] 14. The server provides the adjusted scenario and video to the user's terminal.
[1733] Providing a user experience
[1734] 15. The terminal provides a screen that displays the adjusted video and scenario to the user.
[1735] 16. As the user plays the video and reads the scenario text, the experience is further customized based on feedback from the emotion engine.
[1736] The system recognizes users' emotions in real time and can optimize their travel experience accordingly. For example, if a user is moved by a video of the Eiffel Tower, the system automatically generates and provides related videos and scenarios to further enhance the emotion. In this way, users can transcend physical constraints and enjoy a richer, more personalized travel experience.
[1737] The processing flow will be explained below.
[1738] Step 1:
[1739] A user accesses the system using a user terminal and creates a new account. The user enters information such as name, age, health status, and travel preferences, and presses the registration button.
[1740] Step 2:
[1741] The server receives the input information, validates it, checks for errors, and saves the correct information to the database.
[1742] Step 3:
[1743] The device will display a "Registration Complete" message to the user and redirect them to the login screen.
[1744] Step 4:
[1745] The user enters their user ID and password on the login screen and clicks the login button.
[1746] Step 5:
[1747] The server authenticates the user, checking against information in a database and authorizing if there is a match.
[1748] Step 6:
[1749] The device displays a destination selection screen, providing the user with a world map and a list of destinations.
[1750] Step 7:
[1751] The user selects the desired travel destination and presses the select button.
[1752] Step 8:
[1753] The server receives the user's selection information and generates a request to send details of the selected travel destination to a text generation AI.
[1754] Step 9:
[1755] A text generation AI generates a travel scenario based on the selected travel destination, including information on the history, culture, and key attractions of the tourist destination.
[1756] Step 10:
[1757] The server receives the generated scenario and stores it in a database.
[1758] Step 11:
[1759] The server uses the saved scenario information to send a video generation request to the video generation AI.
[1760] Step 12:
[1761] Based on the scenario, video generation AI generates video footage of the travel destination, including real-time scenery and 3D models.
[1762] Step 13:
[1763] The server receives the generated video and stores it in a database.
[1764] Step 14:
[1765] The terminal displays a screen on which the generated video and scenario can be viewed by the user.
[1766] Step 15:
[1767] The user plays the video, begins reading the scenario text, and clicks on specific locations or events that interest them.
[1768] Step 16:
[1769] The terminal records the user's facial expressions and voice through the camera and microphone built into the user device and transmits them to the emotion engine.
[1770] Step 17:
[1771] The emotion engine analyzes the user's facial expressions, voice, and movements in real time to generate emotion data, which includes the user's emotional state, such as joy, surprise, or boredom.
[1772] Step 18:
[1773] Based on the emotional data received from the emotion engine, the server sends requests to the text generation AI and video generation AI to adjust the optimal scenario and video.
[1774] Step 19:
[1775] The text generation AI and video generation AI generate adjusted scenarios and images based on the emotional data and send them to the server.
[1776] Step 20:
[1777] The server stores the adjusted scenario and video in a database to provide to the user.
[1778] Step 21:
[1779] The device displays the adjusted new scenario and video to the user. For example, if the user is moved, additional video and scenarios will be added to further deepen the user's emotion.
[1780] Based on the above processing steps, a system that provides users with a personalized travel experience is realized, allowing users to enjoy a realistic and rich travel experience optimized according to their emotions, even from the comfort of their own home or hospital room.
[1781] Example 2
[1782] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1783] Conventional virtual travel systems have had difficulty personalizing the user experience. They also lacked the ability to adjust the travel experience in real time to reflect the user's emotions. This resulted in problems such as users becoming bored with the experience midway through and their satisfaction decreasing.
[1784] The specification process by the specification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a selection means for selecting a place the user wants to go to, a text generation means for generating a scenario related to the selected place, a video generation means for generating a video of the place based on the generated scenario, a provision means for providing the generated scenario and video to the user, a request means for the user to request detailed information about the destination based on the scenario or video, a text generation means for generating the detailed information based on the request, an emotion recognition means for recognizing the user's emotion, and an adjustment means for adjusting the scenario and video based on the recognized emotion data. This allows the user's emotions to be reflected in real time, enabling a more personalized and rich travel experience.
[1785] The "selection means" is a function that allows the user to select the place they want to go.
[1786] The "text generation means" is a function for generating scenarios and detailed information about a selected location.
[1787] The "video generation means" is a function for generating video of a location based on the generated scenario.
[1788] The "provision means" is a function for providing the generated scenario and video to the user.
[1789] The "request means" is a function that allows a user to request detailed information about a customer based on a scenario or image.
[1790] The "emotion recognition means" is a function for recognizing the user's emotions.
[1791] The "adjustment means" is a function for adjusting the scenario and video based on the recognized emotion data.
[1792] The "registration means" is a function for registering basic information about a user.
[1793] The "storage means" is a function for storing the registered basic information.
[1794] "Authentication means" is a function that allows a user to access a system.
[1795] The "page providing means" is a function for providing a page exclusively for an authenticated user.
[1796] MODE FOR CARRYING OUT THE INVENTION
[1797] This invention is a system that allows users to virtually experience travel destinations that are difficult to visit in person. This system generates personalized travel experiences based on the user's individual requests, and by incorporating emotion recognition technology, it provides an optimal experience that responds to the user's real-time emotions.
[1798] System Components
[1799] 1. User terminal: A device (e.g., PC, tablet, smartphone) that provides an interface for users to select the destination and experience the virtual journey.
[1800] 2. Server: A back-end system that generates, manages, and provides user information, travel scenarios, and videos.
[1801] 3. Text Generation AI: An artificial intelligence technology that generates travel scenarios based on the locations selected by the user.
[1802] 4. Video Generation AI: Artificial intelligence technology that generates footage of a selected location based on a generated scenario.
[1803] 5. Emotion Recognition Engine: An engine that recognizes user emotions in real time and optimizes the travel experience based on the results.
[1804] 6. Database: A storage system that stores user information, travel scenarios, video, and emotion data.
[1805] System Operation Overview
[1806] When a user accesses the system and selects a place they want to go, the server first retrieves information about the selected place from the database and sends it to a text generation AI. The text generation AI generates a detailed travel scenario based on the selected place and sends the scenario back to the server. The server then sends the generated scenario to a video generation AI, which generates a video of the place based on it.
[1807] The generated scenarios and videos are stored on a server and provided to the user's device. The user can view the scenarios and videos through their device. The emotion recognition engine also uses the user's camera and microphone to analyze emotions from facial expressions and voice in real time and sends the data to the server.
[1808] Optimizing experiences through emotions
[1809] When the server receives the emotion data sent from the emotion recognition engine, it optimizes the travel experience based on that data. For example, if the user is moved, related videos and scenarios are automatically generated to further deepen the emotion, enriching the user's experience. This series of processes is repeated in real time, ensuring that the user always receives an optimized travel experience.
[1810] Specific examples
[1811] For example, if the user selects "The Eiffel Tower in France," the system will send the following prompt to the generative AI model:
[1812] "Generate a scenario in which the user visits the Eiffel Tower in France. The user is a woman in her 20s who likes sightseeing."
[1813] Based on this prompt, the text generation AI generates a detailed scenario about the Eiffel Tower (history, tourist attractions, etc.), and the video generation AI then generates a 360-degree panoramic video based on that scenario.
[1814] In this way, the present invention can reflect the user's emotions in real time and provide a personalized and enriched travel experience.
[1815] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1816] Step 1:
[1817] A user accesses the system using their own device (PC, tablet, smartphone). Using a web browser or a dedicated app, the user enters basic information such as name, age, health status, and travel preferences on the sign-up screen and clicks the "Register" button (input: user's basic information). This information is then sent to the server (output: sending user's basic information).
[1818] Step 2:
[1819] The server receives the basic information sent by the user and performs a validation check to ensure that the format and content are correct (input: basic user information). If validation is successful, this information is saved in the database (output: saved user information). If validation fails, a message is generated to prompt the user to re-enter the invalid information and sent to the user's terminal (output: message prompting re-entry).
[1820] Step 3:
[1821] The terminal receives the "Registration Complete" message from the server and displays it to the user (input: registration complete message). At the same time, it redirects the user to the login screen (output: display of login screen).
[1822] Step 4:
[1823] The user logs in using their account information. After logging in, the user accesses the destination selection screen and chooses the place they want to go (input: login information and destination selection).
[1824] Step 5:
[1825] The server authenticates the user's login information and receives the selected travel destination information (input: login information and travel destination selection). The server records the selected travel destination information in the database and prepares for the next scenario generation (output: saving travel destination information).
[1826] Step 6:
[1827] The server sends travel destination information to the text generation AI and requests it to generate a detailed travel scenario (input: travel destination information) (example prompt: "Please generate a scenario for visiting the Eiffel Tower in France. The user is a woman in her twenties who enjoys sightseeing."). The text generation AI generates a detailed scenario based on the prompt and sends the scenario back to the server (output: generated travel scenario).
[1828] Step 7:
[1829] The server receives the generated scenario and sends it to the video generation AI as a video generation request (input: travel scenario). The video generation AI generates a video of the travel destination based on the scenario and sends it back to the server (output: generated video).
[1830] Step 8:
[1831] The server stores the generated scenario and video in a database (input: generated scenario and video), and then provides the scenario and video to the user's device (output: providing scenario and video).
[1832] Step 9:
[1833] The device sends facial expressions and voice data collected through the user's camera and microphone to the emotion recognition engine in real time (input: facial and voice data). The emotion recognition engine analyzes the data, generates the user's emotion data, and sends it to the server (output: emotion data).
[1834] Step 10:
[1835] The server receives the emotion data and sends a request to the text generation AI and video generation AI again to adjust the optimal scenario and video based on that data (input: emotion data). The generated new scenario and video are sent back to the server (output: adjusted scenario and video).
[1836] Step 11:
[1837] The server provides the adjusted scenario and video back to the user's device (input: adjusted scenario and video), allowing the user to receive a more personalized and enriched travel experience (output: re-provision of scenario and video).
[1838] The above is the specific operation and data processing flow in each processing step of this system.
[1839] (Application example 2)
[1840] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1841] Conventional virtual shopping experience systems allow users to simply browse products in a virtual space, making it difficult to provide a personalized experience based on the user's emotions and preferences. Furthermore, they lack the ability to identify user emotions in real time and optimize the content provided based on those emotions. Furthermore, it is difficult for users to experience the same realistic shopping experience as in a physical store from the comfort of their own home.
[1842] The specification processing by the specification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes a selection means for the user to select a destination they wish to go to, a text generation means for generating a scenario related to the selected destination, a video generation means for generating a video of the destination based on the generated scenario, a provision means for providing the generated scenario and video to the user, a request means for the user to request detailed information about a specific place based on the scenario or video, a text generation means for generating detailed information based on the request, an emotion recognition means for identifying the user's emotion and optimizing the scenario and video to be provided based on the emotion, and an optimization means for regenerating the scenario and video to be provided based on the generated emotion data. This allows the user to have a realistic shopping experience as if they were visiting a physical store, even from the comfort of their own home, and further enables them to enjoy personalized services based on emotion recognition.
[1843] "User" refers to an individual who uses the system.
[1844] "Destination" refers to a virtual location selected by the user that the user wants to visit.
[1845] The "selection means" refers to a means for the user to select the place they want to go.
[1846] "Text generation means" refers to a means for automatically generating a scenario related to a selected destination.
[1847] "Video generation means" refers to means for generating video based on a generated scenario.
[1848] "Providing means" refers to a means for providing the generated scenario and video to the user.
[1849] The "request means" refers to a means for a user to request detailed information about a specific location based on a scenario or video.
[1850] The "emotion recognition means" refers to a means for identifying a user's emotions in real time and optimizing the scenarios and videos provided based on those emotions.
[1851] The "optimization means" refers to a means for regenerating the scenario and video to be provided based on the generated emotion data, and optimizing the user experience.
[1852] "Basic information" refers to personal information registered in the system, such as the user's health status and travel preferences.
[1853] "Registration means" refers to a means for registering basic information of a user in the system.
[1854] "Storage means" refers to the means for storing registered basic information.
[1855] "Authentication means" refers to a means for authenticating a user when logging into a system.
[1856] The "page providing means" refers to a means for providing a page exclusively for an authenticated user.
[1857] "Virtual visit means" refers to a means for allowing a user to progress through the experience as if they had visited the destination.
[1858] The "personalized provision means" refers to a means for providing more personalized products and information to a user based on the emotion identified by the emotion recognition means.
[1859] This invention is a system that allows users to virtually visit a virtual store and have a shopping experience. Furthermore, by combining it with an emotion recognition engine, it provides a personalized experience based on the user's emotions.
[1860] System Components
[1861] 1. User Device:
[1862] This is a device that allows users to select a virtual store and view scenarios and videos. When using a smartphone, a dedicated application is installed. This application is developed using React Native.
[1863] 2. Server:
[1864] This is the backend system that generates, manages, and provides user information, scenarios, and videos. The server uses AWS EC2 or GCP. The following main components are placed on the server:
[1865] Text generation AI (OpenAI GPT-4)
[1866] Video generation AI (DALL-E or other 3D video generation models)
[1867] Emotion Engine (Affectiva SDK)
[1868] Database (MySQL or PostgreSQL)
[1869] Program processing
[1870] 1. Initial settings and user information registration
[1871] A user installs the application and creates a new account, entering personal information such as name, age, preferences, etc., which is then saved to a database. This process is done on the front-end (React Native) and back-end (Node.js / Express).
[1872] 2. Select a store
[1873] After logging in, the user selects the virtual store they want to visit. The store list is retrieved from the server via API communication. The API uses REST or GraphQL.
[1874] 3. Scenario and video generation
[1875] Based on the information of the selected store, the text generation AI (GPT-4) generates a guidance scenario. The generated scenario is sent to the video generation AI (DALL-E), which generates a video of the virtual store. This processing is performed using the OpenAI API and serverless processing (AWS Lambda).
[1876] 4. Operation of the Emotion Engine
[1877] The camera and microphone on the user's device (smartphone) are used to transmit the user's facial expressions and voice to the emotion engine (Affectiva SDK). Emotional data is sent to the server in real time, and the scenario and video are optimized by text generation AI and video generation AI.
[1878] 5. Providing a user experience
[1879] The optimized scenarios and videos are displayed on the user's device, and based on emotion recognition, product listings and custom messages that pique the user's interest are dynamically updated to provide a personalized shopping experience.
[1880] Examples of concrete examples and prompts
[1881] Prompt for text generation AI:
[1882] Enter the user's age, preferences, and selected store information:
[1883] Age: 30
[1884] Interests: Fashion, accessories
[1885] Selected store: High-end fashion mall in the city
[1886] Use this information to generate product lists and scenarios that may be of interest to your users.
[1887] Prompts for video generation AI:
[1888] Generate a 3D image of the interior of a high-end fashion mall in the city, highlighting the product shelves that users are most interested in.
[1889] This system allows users to transcend physical constraints and enjoy a rich and personalized virtual shopping experience. The emotion recognition engine provides optimal content according to the user's emotions, ensuring a highly satisfying experience.
[1890] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1891] Step 1: Initial settings and user information registration
[1892] Input: After users install the application, they enter personal information such as their name, age, preferences, etc.
[1893] What happens: When a user fills out a form and clicks the "Submit" button, the information is sent from the frontend (React Native) to the backend (Node.js / Express), where it is validated and saved to a MySQL or PostgreSQL database.
[1894] Output: The user's personal information is saved in the database and a "Successful registration" message is displayed.
[1895] Step 2: Select a store
[1896] Input: After the user logs in, they select the virtual store they want to visit.
[1897] Specific operation: The store list screen is displayed on the front end, and the user selects the desired store. The selection information is sent to the server, and the information of the selected store is saved in the database.
[1898] Output: The store information selected by the user is saved on the server and used in the next step.
[1899] Step 3: Generate a scenario
[1900] Input: Selected store information.
[1901] Specific operation: The server generates a prompt sentence based on the selected store information and sends it to OpenAI's GPT-4 API. The text generation AI (GPT-4) generates a store guide scenario and returns the result to the server.
[1902] Output: The scenario text generated by GPT-4 is saved on the server.
[1903] Step 4: Generate footage
[1904] Input: Scenario text generated by GPT-4.
[1905] Specific operation: The server generates prompt sentences based on the scenario text and sends them to the video generation AI (DALL-E or other 3D video generation model). The video generation AI generates images of the virtual store based on the scenario and sends the results back to the server.
[1906] Output: The video generated by DALL-E is stored on the server.
[1907] Step 5: Emotion Recognition in Action
[1908] Input: User's facial and voice data.
[1909] Specific operation: The camera and microphone on the user device (smartphone) capture facial expressions and voice, and send them to the emotion recognition engine (Affectiva SDK). The emotion engine analyzes the data, generates emotion data, and sends it to the server.
[1910] Output: The emotion data generated by the emotion recognition engine is stored on the server.
[1911] Step 6: Scenario and footage optimization
[1912] Input: Emotion data, initial scenario and video.
[1913] Specific operation: The server sends a request to the text generation AI and video generation AI again to optimize the scenario and video generated based on the emotion data. As a result, the scenario and video are regenerated based on the user's emotions.
[1914] Output: The new optimized scenario and footage are saved on the server.
[1915] Step 7: Delivering the user experience
[1916] Input: Optimized scenario and footage.
[1917] Specific operation: The user device retrieves the optimized scenario and video from the server and provides it to the user. Product lists and custom messages based on the user's emotion are dynamically displayed.
[1918] Output: The user watches the optimized scenario and video, and enjoys a personalized shopping experience.
[1919] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1920] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1921] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1922] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1923] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1924] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1925] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1926] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1927] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1928] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1929] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1930] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1931] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1932] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1933] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1934] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1935] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1936] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1937] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1938] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1939] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1940] The following is further disclosed regarding the above embodiment.
[1941] (Claim 1)
[1942] A selection means for allowing the user to select a travel destination they wish to visit;
[1943] a text generation means for generating a scenario relating to a selected travel destination;
[1944] a video generation means for generating a video of the travel destination based on the generated scenario;
[1945] providing means for providing the generated scenario and video to a user;
[1946] a request means for a user to request detailed information about a specific location based on a scenario or a video;
[1947] The system includes a text generation means for generating detailed information based on the request.
[1948] (Claim 2)
[1949] a registration means for registering basic information such as the user's health status and travel preferences;
[1950] a storage means for storing the registered basic information;
[1951] The system of claim 1 further comprising:
[1952] (Claim 3)
[1953] an authentication means for users to log into the system;
[1954] a page providing means for providing a page dedicated to an authenticated user;
[1955] The system of claim 1 further comprising:
[1956] "Example 1"
[1957] (Claim 1)
[1958] A selection means for allowing the user to select a travel destination they wish to visit;
[1959] a text generation means for generating a scenario relating to a selected travel destination;
[1960] a video generation means for generating a video of the travel destination based on the generated scenario;
[1961] providing means for providing the generated scenario and video to a user;
[1962] a request means for a user to request detailed information about a specific location based on a scenario or a video;
[1963] a text generation means for generating detailed information based on the request;
[1964] a display means for displaying the detailed information generated in response to a user request on a user terminal;
[1965] A storage means for registering and storing user information at the initial setting;
[1966] A system including:
[1967] (Claim 2)
[1968] a registration means for registering basic information such as the user's health status and travel preferences;
[1969] a storage means for storing the registered basic information;
[1970] The system of claim 1 further comprising:
[1971] (Claim 3)
[1972] an authentication means for users to log into the system;
[1973] a page providing means for providing a page dedicated to an authenticated user;
[1974] The system of claim 1 further comprising:
[1975] "Application Example 1"
[1976] (Claim 1)
[1977] A selection means for allowing the user to select a travel destination they wish to visit;
[1978] a text generation means for generating a scenario relating to a selected travel destination;
[1979] a video generation means for generating a video of the travel destination based on the generated scenario;
[1980] providing means for providing the generated scenario and video to a user;
[1981] a request means for a user to request detailed information about a specific location based on a scenario or a video;
[1982] a text generation means for generating detailed information based on the request;
[1983] a virtual shopping means for a user to purchase goods within the virtual environment;
[1984] A delivery method for delivering items selected in a virtual environment in the real world
[1985] A system including:
[1986] (Claim 2)
[1987] a registration means for registering basic information such as the user's health status and travel preferences;
[1988] a storage means for storing the registered basic information;
[1989] The system of claim 1 further comprising:
[1990] (Claim 3)
[1991] an authentication means for users to log into the system;
[1992] a page providing means for providing a page dedicated to an authenticated user;
[1993] The system of claim 1 further comprising:
[1994] "Example 2: Combining Emotion Engines"
[1995] (Claim 1)
[1996] selection means for selecting a location the user wishes to visit;
[1997] a text generation means for generating a scenario relating to a selected location;
[1998] a video generation means for generating a video of a location based on the generated scenario;
[1999] providing means for providing the generated scenario and video to a user;
[2000] a request means for a user to request detailed information about a customer based on a scenario or a video;
[2001] a text generation means for generating detailed information based on the request;
[2002] emotion recognition means for recognizing an emotion of a user;
[2003] an adjustment means for adjusting the scenario and the video based on the recognized emotion data;
[2004] A system including:
[2005] (Claim 2)
[2006] a registration means for registering basic information of a user;
[2007] a storage means for storing the registered basic information;
[2008] The system of claim 1 further comprising:
[2009] (Claim 3)
[2010] a means of authentication for users to access the system;
[2011] a page providing means for providing a page dedicated to an authenticated user;
[2012] The system of claim 1 further comprising:
[2013] "Application example 2 when combining emotion engines"
[2014] (Claim 1)
[2015] selection means for selecting a destination the user wishes to visit;
[2016] a text generation means for generating a scenario relating to a selected destination;
[2017] a video generation means for generating a video of the destination based on the generated scenario;
[2018] providing means for providing the generated scenario and video to a user;
[2019] a request means for a user to request detailed information about a specific location based on a scenario or a video;
[2020] a text generation means for generating detailed information based on the request;
[2021] an emotion recognition means for identifying the emotion of a user and optimizing the scenarios and videos to be provided based on the emotion;
[2022] an optimization means for regenerating the provided scenario and video based on the generated emotion data;
[2023] A system including:
[2024] (Claim 2)
[2025] a registration means for registering basic information of a user;
[2026] a storage means for storing the registered basic information;
[2027] The system of claim 1 further comprising:
[2028] (Claim 3)
[2029] an authentication means for users to log into the system;
[2030] a page providing means for providing a page dedicated to an authenticated user;
[2031] A virtual visit means for allowing a user to proceed with the experience as if they had visited the destination;
[2032] a personalized provision means for providing more personalized products and information based on the emotion identified by the emotion recognition means;
[2033] The system of claim 1 further comprising: [Explanation of symbols]
[2034] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. A selection means for allowing the user to select a travel destination they wish to visit; a text generation means for generating a scenario relating to a selected travel destination; a video generation means for generating a video of the travel destination based on the generated scenario; providing means for providing the generated scenario and video to a user; a request means for a user to request detailed information about a specific location based on a scenario or a video; and a text generation means for generating detailed information based on the request.
2. a registration means for registering basic information such as the user's health status and travel preferences; a storage means for storing the registered basic information; The system of claim 1 further comprising:
3. an authentication means for users to log into the system; a page providing means for providing a page dedicated to an authenticated user; The system of claim 1 further comprising:
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A