System

The system addresses the limitations of current VR by converting voice requests to text, collecting data, generating scenarios with AI, rendering 3D environments in real time, and updating them based on user movements, providing a highly immersive and continuously improving user experience.

JP2026019150APending Publication Date: 2026-02-05SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024120559
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-25
Publication Date
2026-02-05

AI Technical Summary

Technical Problem

Current virtual reality technologies lack effective methods for generating detailed scenarios in real time and updating 3D environments in response to user movements, and they do not systematically collect user feedback to improve the user experience.

Method used

A system that converts user voice requests to text, collects relevant data from the Internet, generates scenarios using generative AI, renders 3D environments in real time using a virtual reality engine, and updates the environment based on user movements while collecting feedback for system improvement.

Benefits of technology

Enables users to experience specific historical moments or future worlds in a highly immersive and realistic manner, with continuous system enhancement based on user feedback.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026019150000001_ABST
    Figure 2026019150000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system comprising: means for converting a request of a user into text data based on the request; means for collecting relevant data from the Internet; means for generating a scenario using a generative artificial intelligence based on the collected data; means for rendering a 3D environment using a virtual reality engine based on the generated scenario; means for transmitting the generated 3D environment data to a user terminal; means for displaying the 3D environment on the user terminal and updating the 3D environment in real time according to a motion of the user; and means for collecting feedback from the user and improving the system.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] In the fields of modern education and entertainment, there is a lack of ways to experience historical moments in a realistic way. Traditional educational materials also struggle to deepen understanding of history and culture, making it difficult for users to have an immersive learning or entertainment experience. Furthermore, while current virtual reality technology utilizes a limited number of methods to generate scenarios in real time using generative artificial intelligence, a new approach is needed to improve the quality of the user experience. [Means for solving the problem]

[0005] To solve the above problems, the present invention provides the following means. First, it provides a means for converting voice into text data based on a user request. Next, it includes a means for collecting related data (text, images, audio, and video) from the Internet. Furthermore, it provides a means for analyzing the collected data and formatting it into an appropriate format for passing it to a generating AI, and a means for generating a scenario using the generating AI, thereby realizing real-time scenario generation according to the user's wishes. It also provides a means for rendering a 3D environment using a virtual reality engine based on the generated scenario, and transmitting the generated 3D environment data to a user terminal. It also provides a means for displaying the 3D environment on the user terminal and updating it in real time according to the user's movements. Finally, it includes a means for collecting user feedback and improving the system, thereby realizing a system that consistently provides a high-quality virtual experience.

[0006] A "request" is an instruction by a user requesting a specific experience.

[0007] "Text data" is data that has been converted from audio, images, or video into a character string format.

[0008] "Means of collecting relevant data from the Internet" refers to a function that automatically retrieves information such as text, images, audio, and video from websites and databases based on specific keywords.

[0009] "Generative AI" is a system that uses artificial intelligence technology to generate new scenarios and content based on input data.

[0010] A "scenario" is a set of settings that includes details of the environment and events that the user experiences.

[0011] "Virtual reality engine" is a general term for software and hardware used to create virtual reality environments.

[0012] A "3D environment" is a virtual world constructed in three-dimensional space.

[0013] "Real-time rendering" is a processing technology that instantly renders a 3D environment in response to the user's movements and line of sight.

[0014] "Feedback" refers to the opinions and comments provided by users after their experience.

[0015] "System improvement" refers to updating and adjusting the system to improve its performance and functionality based on collected feedback. [Brief explanation of the drawings]

[0016] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11]FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0017] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0018] First, the terms used in the following description will be explained.

[0019] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0020] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0021] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0022] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0023] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0024] [First embodiment]

[0025] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0026] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0027] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0028] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0029] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0030] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0031] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0032] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0033] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0034] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0035] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0036] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0037] This invention provides a system that allows users to experience specific historical moments in the past or future through virtual reality. Based on user requests, this system collects relevant data from the Internet, generates scenarios using generative artificial intelligence, and provides a mechanism for rendering 3D environments in real time using a virtual reality engine.

[0038] Program Overview

[0039] 1. User submits request:

[0040] The user speaks into the device to request an experience of a specific time in the past or future.

[0041] The device analyzes the voice and converts it into text data.

[0042] 2. Data Collection:

[0043] The terminal transmits the text data to the server.

[0044] The server collects relevant information from the Internet based on the request.

[0045] 3. Data analysis and scenario generation:

[0046] The server analyzes the collected data and generates scenarios using generative artificial intelligence.

[0047] A scenario details the buildings, people, and events associated with a specified time period and place.

[0048] 4. Real-time rendering:

[0049] The server uses the generated scenario to render a 3D environment in a virtual reality engine.

[0050] The 3D environment contains details and realism that match the user's requirements.

[0051] 5. Data transmission and display:

[0052] The server sends the rendered 3D environment data to the terminal.

[0053] The device displays the received data on a virtual reality headset.

[0054] 6. Immersive user experience:

[0055] The user wears a virtual reality headset and is immersed in a virtual environment.

[0056] The environment is updated in real time according to the user's movements and gaze.

[0057] 7. Feedback collection and system improvement:

[0058] After the user has completed the experience, the terminal displays a feedback form in which the user can enter comments about the experience and requests for improvements.

[0059] The device sends feedback to the server, which analyzes it and uses it to improve the system.

[0060] Specific examples

[0061] Case: Renaissance Florence

[0062] 1. Submit your request:

[0063] The user speaks to the device and says, "I want to experience Renaissance Florence."

[0064] The terminal converts this voice into text data and sends it to the server.

[0065] 2. Data Collection:

[0066] The server collects information related to "Renaissance Florence" from the Internet.

[0067] For example, data such as text, images, audio, and video related to buildings, people, and events from that time is obtained.

[0068] 3. Data analysis and scenario generation:

[0069] The server analyzes the collected data and uses generative artificial intelligence to generate a scenario of "Florence during the Renaissance."

[0070] Scenarios include Da Vinci painting and the construction site of Florence Cathedral.

[0071] 4. Real-time rendering:

[0072] Based on the generated scenario, the server uses a virtual reality engine to render a 3D environment of Florence in real time.

[0073] 5. Data transmission and display:

[0074] The server sends the rendered 3D environment data to the terminal.

[0075] The device receives the data and displays it on a virtual reality headset.

[0076] 6. Immersive Experience:

[0077] Users put on a headset and are immersed in a realistic Renaissance Florence.

[0078] Users can interact with Da Vinci and tour the inside of the cathedral.

[0079] 7. Feedback and System Improvement:

[0080] After completing the experience, users fill out a feedback form to indicate their satisfaction with the experience and suggestions for improvement.

[0081] The device sends this to the server, which analyzes the feedback and uses it to improve the system.

[0082] This allows users to experience historical moments and future worlds in a highly immersive way.

[0083] The processing flow will be explained below.

[0084] Step 1:

[0085] The user makes a request to experience a specific time in the past or future by voice into the device, which then recognizes the voice and converts the voice data into text data.

[0086] Step 2:

[0087] The device sends the converted text data to the server as an HTTP request, which includes the user's desired experience.

[0088] Step 3:

[0089] Based on the received request, the server launches a scraping tool to scrape relevant text, image, audio, and video data from the Internet.

[0090] Step 4:

[0091] The server uses a scraping tool to collect data from the Internet that matches the specified keywords, and the collected data is temporarily stored.

[0092] Step 5:

[0093] The server analyzes the collected data using natural language processing and image recognition technology and formats it appropriately to be passed to the generative artificial intelligence.

[0094] Step 6:

[0095] The server inputs the formatted data into a generative AI to generate a scenario for the specified time period and location, including detailed historical background and key buildings, people, and events.

[0096] Step 7:

[0097] The server uses an evaluation algorithm to check the quality of the generated scenario, correcting or regenerating it as necessary. Once the evaluation is complete, the scenario moves on to the next stage.

[0098] Step 8:

[0099] The server uses a virtual reality engine to render a 3D environment in real time based on the scenario, using the GPU to achieve high performance and detailed rendering.

[0100] Step 9:

[0101] The server compresses the rendered 3D environment data and transmits it to the device in an efficient binary format, using streaming techniques to minimize data latency.

[0102] Step 10:

[0103] The device then extracts the received 3D environment data and displays it on the virtual reality headset, which updates the displayed 3D environment in real time according to the user's movements and gaze.

[0104] Step 11:

[0105] Users wear a virtual reality headset and enter an immersive virtual environment, where they can freely move around and experience interactive elements (e.g., interact with people, manipulate objects).

[0106] Step 12:

[0107] When the user finishes their experience, the device displays a feedback form where the user can enter comments about their experience and suggestions for improvement. Collecting feedback contributes to improving user satisfaction.

[0108] Step 13:

[0109] The device sends the collected feedback data to the server, which analyzes the feedback and adjusts the algorithms of the generative AI and virtual reality engine to help improve the system, resulting in an even better user experience next time.

[0110] Example 1

[0111] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0112] Conventional virtual reality systems face challenges when it comes to quickly and accurately generating detailed scenarios and rendering 3D environments in real time when users want to experience specific historical moments or future worlds. Other challenges include updating the 3D environment in real time in response to user movements and systematically collecting user feedback to improve the system.

[0113] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0114] In this invention, the server includes means for converting user requests into text data based on the user requests, means for collecting related information from a communication network, means for generating a scenario using artificial intelligence based on the collected information, means for creating a 3D environment using virtual reality technology based on the generated scenario, means for transmitting the generated 3D environment data to a user device, means for displaying the 3D environment on the user device and updating it in real time according to the user's movements, and means for collecting user responses and improving the system, thereby enabling users to experience specific historical moments or future worlds in detail and realistically.

[0115] "User" refers to a person who uses the system to experience virtual reality.

[0116] A "request" refers to a request for a specific historical moment or future world that a user would like to experience.

[0117] "Text data" refers to data in which a user's request is converted into text form through voice recognition or other input methods.

[0118] "Communications network" means a wide-area computer network, including the Internet and other data communications infrastructures.

[0119] "Information" refers to data such as text, images, and videos collected based on user requests.

[0120] "Artificial intelligence" refers to machine learning algorithms and neural networks used to analyze collected information and generate specified scenarios.

[0121] A "scenario" refers to a set of information that contains a detailed description of a particular historical moment or future world that a user wants to experience.

[0122] "Virtual reality technology" refers to technology that uses computer technology to generate a 3D environment and provide an immersive experience to the user.

[0123] "3D environment" refers to a three-dimensional virtual space created using virtual reality technology based on a scenario.

[0124] "User device" refers to the terminal or headset that a user uses to have a virtual reality experience.

[0125] "Real-time updates" means that the 3D environment is instantly reflected and updated in response to inputs such as the user's movements and gaze.

[0126] "Reactions" refer to the feedback and opinions users provide through their virtual reality experience.

[0127] "System" refers to the comprehensive technological platform that includes all of the above means.

[0128] The present invention is a system that allows users to experience specific historical moments in the past or future through virtual reality. Based on a user's request, the system collects relevant information from the internet, generates a scenario using a generative AI model, and uses virtual reality technology to render a 3D environment in real time.

[0129] First, a user requests a historical moment or future world they would like to experience through a voice request. For example, a user might say, "I want to experience Tokyo of the future." This voice request is converted into text data by the device's voice recognition software (e.g., Google Speech-to-Text API).

[0130] The device sends the converted text data to a server, which then collects related information from the Internet via a communications network. Specifically, it uses web scraping tools (e.g., BeautifulSoup) or APIs (e.g., Wikipedia API) to obtain the necessary text, images, videos, etc.

[0131] The collected information is analyzed by the server and passed to a generative AI model (e.g., OpenAI's GPT-4). This generative AI model creates a detailed scenario based on the collected information. For example, it uses a prompt such as, "Generate a detailed scenario about Tokyo in the future. Include skyscrapers, a modern transportation system, and people's activities."

[0132] Based on the generated scenario, the server uses virtual reality technology to render a 3D environment in real time using a virtual reality engine (e.g., Unreal Engine or Unity). Specifically, the future Tokyo cityscape, transportation system, building interiors, streetscapes, and other details are depicted.

[0133] The generated 3D environment data is sent to the terminal and displayed on the user device (e.g., a virtual reality headset). The user wears the headset and is immersed in the virtual environment. The user's movements and gaze are detected by tracking sensors (e.g., the tracking system in the HTC Vive), and the 3D environment is updated in real time. This allows the user to freely explore the virtual reality and enjoy a detailed experience.

[0134] Finally, once the user has finished the experience, the device displays a feedback form. The user can enter comments about the experience or suggestions for improvement, which the device then sends to the server. The server analyzes the collected feedback and incorporates it into the next update or improvement. This feedback is then used with natural language processing tools and machine learning algorithms to continuously improve the system.

[0135] As described above, the present invention provides a system that provides a user with a high level of immersion and allows them to experience specific historical moments in the past or future in detail and realistically.

[0136] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0137] Step 1: User submits request

[0138] Users speak their requests into the device, for example, saying, "I want to experience Renaissance Florence."

[0139] The device receives this voice input and converts it into text using voice recognition software, specifically the Google Speech-to-Text API.

[0140] Input: User's voice request

[0141] Output: The request converted to character data

[0142] Step 2: Data collection

[0143] The terminal transmits the converted character data to the server.

[0144] Based on the received request, the server collects relevant information from the Internet via a communication network, for example, using a web scraping tool (BeautifulSoup) or an API (Wikipedia API).

[0145] Input: Request converted to character data

[0146] Output: Collected information (text, images, videos, etc.)

[0147] Step 3: Data analysis and scenario generation

[0148] The server analyzes the collected information and organizes it using a natural language processing tool (spaCy).

[0149] The server inputs a prompt into a generative AI model (e.g., OpenAI's GPT-4) to generate a detailed scenario. An example of the prompt is, "Generate a scenario set in Renaissance Florence, including a scene where Da Vinci is painting."

[0150] Input: Collected information

[0151] Output: Generated scenario

[0152] Step 4: Real-time rendering

[0153] Based on the generated scenario, the server uses virtual reality technology to render a 3D environment in real time, specifically using a virtual reality engine (Unreal Engine or Unity).

[0154] Input: Generated scenario

[0155] Output: Rendered 3D environment

[0156] Step 5: Send and display data

[0157] The server sends the rendered 3D environment data to the device, using WebSocket or HTTP / 2 as the communication protocol.

[0158] The device prepares the received data for display on a virtual reality headset.

[0159] Input: Rendered 3D environment data

[0160] Output: Display to a virtual reality headset

[0161] Step 6: User Immersion

[0162] The user wears a virtual reality headset and is immersed in a virtual environment.

[0163] The device uses sensors to track the user's movements and gaze, updating the 3D environment in real time, for example using the tracking system in the HTC Vive.

[0164] Input: User movement and gaze data

[0165] Output: Real-time updated 3D environment

[0166] Step 7: Gather feedback and improve the system

[0167] After the user has finished the experience, the terminal displays a feedback form.

[0168] Users enter comments about their experience and requests for improvement, and the device sends these to the server.

[0169] The server analyzes the collected feedback and uses natural language processing tools and machine learning algorithms to help improve the system.

[0170] Input: User feedback

[0171] Output: Analyzed feedback and system improvements

[0172] The above is the specific processing flow of this system.

[0173] (Application example 1)

[0174] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0175] Current virtual reality systems lack efficient means for users to experience specific historical moments or future events with a high level of realism. Furthermore, technologies for automatically generating scenarios in real time based on collected data and then generating and updating virtual reality environments in real time based on those scenarios remain a challenge. Therefore, systems that improve the quality of the user experience are needed.

[0176] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0177] In this invention, the server includes means for converting a user's request into text data based on the request, means for collecting relevant data from the Internet, means for generating a scenario using a generating artificial intelligence based on the collected data, means for rendering a three-dimensional environment using a virtual reality engine based on the generated scenario, means for transmitting the generated three-dimensional environment data to a user terminal, means for displaying the three-dimensional environment on the user terminal and updating it in real time according to the user's movements, means for collecting user feedback and improving the system, means for providing the generated three-dimensional environment to the user using a visual display device, means for analyzing the collected data and formatting it into an appropriate format for passing it to the generating artificial intelligence, means for ensuring the quality of the scenario using an algorithm for evaluating the quality of the generated scenario, means for rendering the generated three-dimensional environment in real time using a virtual reality engine and transmitting it to a visual display device, and means for displaying the generated three-dimensional environment in real time on the user's mobile device or head-mounted display and updating the environment in accordance with the user's line of sight and movements, thereby enabling users to experience historical moments and future events in real time with a high level of realism.

[0178] A "user request" refers to a user's desire to experience a particular historical moment or future event through audio or text.

[0179] "Text data" refers to data obtained by converting a user's request from voice into text information.

[0180] "Related Data" refers to information about a specified time and place that is collected from the Internet based on a user's request.

[0181] "Generative AI" refers to artificial intelligence technology that generates scenarios based on collected data.

[0182] A "scenario" is a plan or script that details the events and settings to be experienced within a virtual reality environment based on generated data.

[0183] "Virtual reality engine" refers to a software and hardware configuration for rendering three-dimensional environments in real time.

[0184] "Three-dimensional environment" refers to the virtual space that a user experiences through virtual reality.

[0185] "User terminal" refers to a device including a smartphone, a head-mounted display, or other display device.

[0186] "Visual display device" refers to a display device used by a user to visually experience a three-dimensional environment.

[0187] "Feedback" refers to comments, ratings, and requests for improvement regarding the experience provided by the user after the experience.

[0188] "Formatting" refers to the process of converting collected data into an appropriate form for use by the generative AI.

[0189] "Algorithm" refers to a computational procedure for assessing and ensuring the quality of generated scenarios.

[0190] The present invention provides a system that allows a user to experience a specific historical moment or a future event through virtual reality. The system comprises the following means:

[0191] First, the user inputs a request by voice or text using a device such as a smartphone or head-mounted display. For example, the user's request might be something like, "I want to experience a Renaissance city." This request is converted into text data using speech recognition technology (for example, Python's speech_recognition library).

[0192] The server then retrieves the converted text data and collects relevant data from the Internet. During this process, the required data is retrieved from the source using, for example, the requests library. The collected data is then analyzed using generative artificial intelligence (e.g., OpenAI's API) and formatted accordingly.

[0193] The server generates a scenario based on the formatted data. The scenario generated using the AI ​​generator includes detailed events and settings related to the specified time period and location. When generating the scenario, a prompt message corresponding to the user's request is passed to the AI ​​generator. For example, the following prompt message is used:

[0194] Example prompt:

[0195] A user has requested to experience a Renaissance city. Generate a scenario based on the following relevant data:

[0196] Renaissance buildings and streets

[0197] Costumes and culture of the time

[0198] Historical events and everyday scenes

[0199] Generate detailed scenarios and make them suitable for virtual reality experiences.

[0200] The generated scenario is rendered in real time as a 3D environment using a virtual reality engine (e.g., a 3D engine such as py3dengine). The rendered 3D environment data is sent from the server to the user's device, and the user is immersed in the virtual environment through a visual display device (e.g., a head-mounted display or a smartphone).

[0201] After the user has completed their experience in the virtual environment, the server collects feedback from the user by entering comments and ratings into a dedicated form, which is used to improve the system.

[0202] For example, if a user requests "I want to experience a Renaissance city," the server collects relevant data and uses generative artificial intelligence to generate a specific Renaissance scenario, including detailed descriptions of buildings, cityscapes, and costumed people in action. The generated scenario is then rendered as a three-dimensional environment by a virtual reality engine, allowing the user to use a visual display device to tour the city and interact with historical figures.

[0203] As described above, the present embodiment provides a series of processes for enhancing a user's immersive experience.

[0204] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0205] Step 1:

[0206] The user speaks a request into the device. This request describes a desire to experience a specific historical moment or future event. The device recognizes this speech as input data and converts it into text data using speech recognition technology (for example, Python's speech_recognition library). The output is the user's request converted into text data as character information.

[0207] Step 2:

[0208] The terminal sends the converted text data to the server. The server receives the text data as input and collects related data from the Internet. Here, it uses the requests library to obtain the necessary data from the target sources. The output is the collected related data.

[0209] Step 3:

[0210] The server takes the collected data as input, analyzes it using generative AI (for example, OpenAI's API), and formats it into an appropriate format. During this process, the data is structured and necessary information is filtered. The output is data formatted in an appropriate format for the generative AI to generate scenarios.

[0211] Step 4:

[0212] The server uses the formatted data to send prompts to the generation AI to generate a scenario. The prompt text is a specific scenario request based on the user's request and the collected data. The generation AI generates a scenario based on the prompt and provides the scenario in text format as output.

[0213] Step 5:

[0214] The server takes the generated scenario as input and uses a virtual reality engine (e.g., py3dengine) to render a 3D environment in real time. During this process, a 3D environment is generated according to the generated scenario. The output is 3D data rendered as a virtual reality environment.

[0215] Step 6:

[0216] The server transmits the rendered 3D environment data to the user's device, which receives this data as input and displays the virtual environment to the user through a visual display device (such as a head-mounted display or a smartphone). The user is immersed in this virtual environment and experiences it.

[0217] Step 7:

[0218] After the user has completed their experience in the virtual environment, the device displays a feedback form. The user enters comments, ratings, and suggestions for improvement about the experience. The device sends this as input data to the server. The server analyzes the feedback and uses it to improve the system. The output is improvements based on the feedback.

[0219] This series of processing steps allows users to experience historical moments and future events in real time with a high level of realism.

[0220] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0221] This invention is a system that recognizes a user's emotions in a virtual reality environment and dynamically changes the experience content based on those emotions. Based on the user's request, the system converts the request into text data, collects related data from the Internet, and generates a scenario using generative artificial intelligence. Based on the generated scenario, a 3D environment is rendered using a virtual reality engine and displayed on the user's device. Furthermore, by updating the environment in real time based on the user's movements and combining it with an emotion engine that recognizes the user's emotions, an experience tailored to the user's emotions is provided.

[0222] Program Overview

[0223] 1. User submits request:

[0224] The user requests an experience of a specific time in the past or future by speaking into the device, which then analyzes the voice and converts the voice data into text data.

[0225] 2. Data Collection:

[0226] The device sends text data to the server as an HTTP request, and the server collects related information from the Internet based on the request content.

[0227] 3. Data analysis and scenario generation:

[0228] The server analyzes the collected data and uses generative artificial intelligence to generate scenarios that include detailed descriptions of historical background, key buildings, people, and events.

[0229] 4. Real-time rendering:

[0230] The server uses the generated scenario to render a 3D environment in real time in a virtual reality engine, with details and realism that match the user's requirements.

[0231] 5. Data transmission and display:

[0232] The server sends the rendered 3D environment data to the device, which displays it on a virtual reality headset.

[0233] 6. Immersive user experience:

[0234] Users wear a virtual reality headset and are immersed in a virtual environment that updates in real time based on the user's movements and gaze.

[0235] 7. Emotion Recognition with Emotion Engine:

[0236] The server uses an emotion engine that analyzes the user's facial expressions and vocal tone to recognize the user's emotions in real time. The recognized emotion data is analyzed and reflected in the content of the virtual environment.

[0237] 8. Dynamic Environment Changes:

[0238] The server dynamically modifies the virtual reality environment based on the emotional data recognized by the emotion engine, for example by increasing the effects in the environment if the user is surprised, or by adding interactive elements if the user is engrossed.

[0239] 9. Feedback Collection and System Improvement:

[0240] After the user has finished the experience, the device displays a feedback form, where the user can enter comments about the experience and suggestions for improvement. The device then sends this feedback to the server, which then analyzes the feedback and emotional data to help improve the system.

[0241] Specific examples

[0242] Case: Renaissance Florence

[0243] 1. Submit your request:

[0244] The user speaks to the device, saying, "I want to experience Renaissance Florence." The device converts this speech into text data and sends it to the server.

[0245] 2. Data Collection:

[0246] The server collects information related to "Renaissance Florence" from the Internet, including text, images, audio, and video data related to buildings, people, and events from that time.

[0247] 3. Data analysis and scenario generation:

[0248] The server analyzes the collected data and uses generative artificial intelligence to generate a "Renaissance Florence" scenario, including scenes of Da Vinci painting and the construction site of Florence Cathedral.

[0249] 4. Real-time rendering:

[0250] Based on the generated scenario, the server uses a virtual reality engine to render a 3D environment of Florence in real time.

[0251] 5. Data transmission and display:

[0252] The server sends the rendered 3D environment data to the device, which receives the data and displays it on the virtual reality headset.

[0253] 6. Immersive Experience:

[0254] Users put on a headset and are immersed in a realistic Renaissance Florence, where they can interact with Da Vinci and tour the interior of the cathedral.

[0255] 7. Emotion Recognition with Emotion Engine:

[0256] The server analyzes the user's facial expressions and voice using an emotion engine and recognizes that the user is excited.

[0257] 8. Dynamic Environment Changes:

[0258] Based on the user's emotional data, the server adds special events (e.g., fireworks and musical performances) to the streets of Florence to make the experience even more engaging.

[0259] 9. Feedback and System Improvement:

[0260] After completing the experience, the user fills in a feedback form to indicate their satisfaction with the experience and any suggestions for improvement. The device then sends this information to the server, which then analyzes the feedback and emotional data and reflects it in improvements to the system.

[0261] In this way, the system of the present invention recognizes the user's emotions and dynamically modifies the virtual reality experience based on them, providing a richer and more engaging experience.

[0262] The processing flow will be explained below.

[0263] Step 1:

[0264] The user speaks into the device, saying, "I want to experience Renaissance Florence." The device uses voice recognition software to convert the voice data into text.

[0265] Step 2:

[0266] The device sends the converted text data to the server as an HTTP request, which includes the user's desired experience.

[0267] Step 3:

[0268] Based on the received request, the server uses a scraping tool to collect related information (text, image, audio, and video data) that matches the specified keywords from the Internet.

[0269] Step 4:

[0270] The server analyzes the scraped data using natural language processing and image recognition technology and formats it appropriately to be passed on to generative artificial intelligence (AI).

[0271] Step 5:

[0272] The server inputs the formatted data into a generative artificial intelligence (AI) that generates a scenario about "Renaissance Florence," including detailed descriptions of buildings, people, and events.

[0273] Step 6:

[0274] The server uses an evaluation algorithm to verify the quality of the generated scenarios, and corrects or regenerates them as necessary.

[0275] Step 7:

[0276] The server uses quality-tested scenarios to leverage a virtual reality engine to render 3D environments in real time.

[0277] Step 8:

[0278] The server compresses the rendered 3D environment data in binary format and transmits it to the terminal using efficient data transfer techniques.

[0279] Step 9:

[0280] The device receives 3D environment data from the server and displays it on the virtual reality headset. The device tracks the user's gaze and movements in real time and updates the environment.

[0281] Step 10:

[0282] Users put on a virtual reality headset and begin experiencing a realistic Renaissance Florence. They can freely explore the virtual environment and interact with the characters.

[0283] Step 11:

[0284] The server uses an emotion engine to collect the user's facial expressions and voice tone in real time from the camera and microphone installed in the headset and analyze the user's emotions.

[0285] Step 12:

[0286] The server dynamically modifies the virtual reality environment based on the analyzed emotional data, for example adding effects and interactive elements if the user is excited.

[0287] Step 13:

[0288] After the user has finished the experience, the terminal displays a feedback form, where the user can enter comments about the experience and requests for improvements. The terminal then sends the feedback to the server.

[0289] Step 14:

[0290] The server analyzes the feedback and emotional data, adjusts the algorithms of the generative AI and virtual reality engine, and reflects the results in system improvements, resulting in an even better user experience next time.

[0291] Example 2

[0292] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0293] Conventional virtual reality systems lack the flexibility to adapt to user requests and the ability to dynamically change the environment based on the user's emotions. These shortcomings limit the quality of the user experience and make it difficult to provide a richer and more engaging virtual reality experience. Therefore, there is a need for a system that dynamically changes the generation of scenarios and rendering of virtual reality environments based on user requests, depending on the user's emotions.

[0294] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0295] In this invention, the server includes: means for converting a user's request into text data based on the request; means for collecting related data from the Internet; means for generating a scenario using a generative model based on the collected data; means for rendering a three-dimensional environment using a virtual reality engine based on the generated scenario; means for transmitting the generated three-dimensional environment data to a user terminal; means for displaying the three-dimensional environment on the user terminal and updating it in real time according to the user's movements; means including an emotion analysis engine for recognizing the user's emotions; means for dynamically changing the virtual reality environment based on the recognized emotion data; and means for collecting user feedback and improving the system. This makes it possible to flexibly generate scenarios according to user requests and dynamically change the virtual reality environment according to the user's emotions.

[0296] "Means for converting a request into text data" refers to a function for analyzing a request input by a user in a voice or other format and converting it into text.

[0297] "Means of collecting relevant data from the Internet" refers to the function of obtaining necessary information from websites and databases via a network.

[0298] A "generative model" refers to an artificial intelligence algorithm that generates appropriate results for a given input (prompt).

[0299] "Means for generating a scenario" refers to the function of creating stories and events for a virtual reality experience that responds to a user's request based on collected data.

[0300] "Virtual reality engine" refers to a software platform for generating and rendering three-dimensional virtual environments in real time.

[0301] "Means for rendering a three-dimensional environment" refers to a function for depicting a three-dimensional visual environment based on a generated scenario.

[0302] "Means for transmitting three-dimensional environmental data to a user terminal" refers to a function for transferring three-dimensional visual data generated by the server to the terminal used by the user.

[0303] "Means for displaying a three-dimensional environment and updating it in real time in response to the user's movements" refers to a function that dynamically changes the three-dimensional virtual environment displayed on the user's terminal in synchronization with the user's movements and operations.

[0304] An "emotion analysis engine" refers to software that analyzes a user's facial expressions and vocal tone to identify emotions.

[0305] "Means for dynamically modifying the virtual reality environment based on recognized emotional data" refers to the ability to adjust the content and events of the virtual reality in real time based on data obtained through emotional analysis.

[0306] "Means for collecting feedback and improving the system" refers to a function for incorporating user evaluations and opinions of the experience and improving the experience in future.

[0307] The present invention is a system that generates a scenario based on a user's request in a virtual reality environment and provides the user with a three-dimensional environment corresponding to that scenario. This invention makes it possible to recognize the user's emotions in real time and dynamically change the virtual reality environment based on those emotions.

[0308] First, the user speaks into the device to request an experience of a specific time or place. The device then uses voice recognition software, such as the Google Speech-to-Text API, to convert this voice data into text. This text data then becomes the basis for interpreting the request.

[0309] The converted text data is then sent to the server as an HTTP request, which uses web crawling tools and APIs to gather relevant information from the internet, such as Beautiful Soup and Scrapy, to gather text, images, and audio data.

[0310] The collected data is then analyzed using natural language processing tools, such as spaCy and NLTK. Based on the results of this analysis, a generative AI model (e.g., GPT-4, ChatGPT) generates a scenario. Input prompts are crucial for scenario generation. For example, a prompt such as "Walk through the streets of Renaissance Florence" generates detailed historical background and character dialogue.

[0311] Based on the generated scenario, the server uses a virtual reality engine (e.g., Unity, Unreal Engine) to render a three-dimensional virtual environment. The rendered data is sent to the user's device via HTTP. The user's device receives this data and displays it on a virtual reality headset (e.g., Oculus Rift, HTC Vive). The user can then immerse themselves in and interact with this virtual environment.

[0312] During the user's experience, the device uses a camera and microphone to capture the user's facial expressions and voice. This data is analyzed in real time by an emotion analysis engine (e.g., Affectiva, Kairos). Based on the user's emotion data, the server dynamically modifies the virtual environment. For example, if the user is surprised, effects such as fireworks are added to the virtual environment.

[0313] After the experience is over, the device displays a feedback form to the user, allowing them to input their satisfaction with the experience and suggestions for improvement. This data is then sent to the server, which analyzes the feedback and emotional data to improve the experience for the next session.

[0314] A concrete example is a user request to "experience Renaissance Florence." In this case, the user speaks the request into the device, which converts this speech into text data and sends it to the server. The server collects relevant information from the Internet and generates a scenario using a generative AI model. It then renders a three-dimensional environment of Florence using a virtual reality engine and sends it to the user's device. The user puts on the virtual reality headset and explores the city of Florence. If the user is excited, the server adds special events to make the experience even more engaging.

[0315] This invention makes it possible to flexibly generate scenarios based on user requests and dynamically change the virtual reality environment in response to the user's emotions.

[0316] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0317] Step 1:

[0318] User request submission

[0319] Input: The user speaks into the device, "I want to experience Renaissance Florence."

[0320] How it works: The device uses speech recognition software (e.g., Google Speech-to-Text API) to convert speech into text. Voice input is text, and output is text.

[0321] Output: The converted text data. The text is "I want to experience Renaissance Florence."

[0322] Step 2:

[0323] Data collection

[0324] Input: The request as text data.

[0325] How it works: The device sends this text data to the server as an HTTP request. The server then uses a web crawling tool (e.g., Beautiful Soup, Scrapy) to collect relevant information from the Internet. This includes data such as text, images, and audio.

[0326] Output: A dataset containing relevant information.

[0327] Step 3:

[0328] Data analysis and scenario generation

[0329] Input: A dataset containing the relevant collected information.

[0330] How it works: The server uses natural language processing tools (e.g., spaCy, NLTK) to parse the data and format it into a suitable format for a generative AI model (e.g., GPT-4, ChatGPT). It then generates a scenario based on the generative AI model. For example, you can input the prompt "Florence during the Renaissance," and the scenario will be output.

[0331] Output: Generated scenario data.

[0332] Step 4:

[0333] Real-time rendering

[0334] Input: Generated scenario data.

[0335] How it works: The server uses a virtual reality engine (e.g., Unity, Unreal Engine) to render a 3D environment in real time. Based on the scenario, models of objects and characters in the environment are generated and rendered in real time.

[0336] Output: Rendered 3D environment data.

[0337] Step 5:

[0338] Sending and Displaying Data

[0339] Input: Rendered 3D environment data.

[0340] How it works: The server sends this data to the user's device, which then displays it on a virtual reality headset (e.g., Oculus Rift, HTC Vive).

[0341] Output: A three-dimensional environment displayed in a virtual reality headset.

[0342] Step 6:

[0343] Immersive User Experience

[0344] Input: A three-dimensional environment displayed on a virtual reality headset.

[0345] How it works: The user puts on a virtual reality headset and is immersed in a virtual environment. The environment updates in real time based on the user's movements and gaze. For example, the user can explore and interact with Renaissance Florence.

[0346] Output: Updated virtual environment and user experience data.

[0347] Step 7:

[0348] Emotion recognition by emotion engine

[0349] Input: User's facial expressions and voice.

[0350] How it works: The device captures the user's facial expressions and voice using a camera and microphone. The server analyzes this data in real time using an emotion analysis engine (e.g., Affectiva, Kairos) to recognize the user's emotions.

[0351] Output: Recognized emotion data.

[0352] Step 8:

[0353] Dynamic Environment Changes

[0354] Input: Recognized emotion data.

[0355] Action: The server dynamically changes the virtual reality environment based on the emotional data, for example adding special events or effects if the user is excited.

[0356] Output: A dynamically modified virtual reality environment.

[0357] Step 9:

[0358] Feedback collection and system improvement

[0359] Input: User feedback after the experience.

[0360] How it works: After the experience is over, the device displays a feedback form to the user. The user enters their satisfaction with the experience and suggestions for improvement, and this data is sent to the server. The server analyzes the feedback and emotion data to help improve the system.

[0361] Output: An improved system and updated user experience data.

[0362] (Application example 2)

[0363] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0364] Conventional virtual reality systems provide a fixed experience without taking the user's emotions into consideration. This means that the experience cannot be dynamically changed based on the user's emotions and reactions, resulting in a lack of immersion and personalization. Furthermore, it has not been possible for the system to recognize the user's emotions during the experience and provide an optimized experience based on those emotions. This has led to issues such as low user satisfaction and engagement.

[0365] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0366] In this invention, the server includes means for converting a user's request into text data based on the user's request, means for collecting related data from the Internet, means for generating a scenario based on the collected data using artificial intelligence, means for rendering a 3D environment based on the generated scenario using a virtual reality engine, means for transmitting the generated 3D environment data to a user terminal, means for displaying the 3D environment on the user terminal and updating it in real time according to the user's movements, means for recognizing the user's emotions and dynamically changing the virtual reality environment based on the recognized emotions, and means for collecting user feedback and improving the system. This enables dynamic changes to the virtual reality experience based on the user's emotions, providing a more immersive and personalized experience.

[0367] A "user request" is a request to the system for content or information that the user wants to experience.

[0368] A "means for converting to text data" is a process or device that analyzes and converts voice or other non-text data into written information.

[0369] "Means for collecting relevant data from the Internet" refers to a process or device that searches for and obtains specified information or data via the Internet.

[0370] "Generative AI" is an AI technology that generates scenarios and content for specific purposes based on collected data.

[0371] A "scenario generation means" is a process or device that analyzes collected data and arranges it into a specific story or scenario format.

[0372] A "virtual reality engine" is software or hardware that uses computer graphics technology to create and render virtual reality environments.

[0373] A "means for rendering a 3D environment" is a process or device that depicts a three-dimensional virtual space in real time based on a generated scenario.

[0374] The "means for transmitting the generated 3D environment data to the user terminal" is a process or device that transfers the rendered 3D environment data to the terminal used by the user.

[0375] "Means for displaying a 3D environment on a user device and updating it in real time in response to the user's movements" refers to a process or device that displays a 3D virtual space on a user device and dynamically changes that virtual space in response to the user's operations and movements.

[0376] "Means for recognizing a user's emotions and dynamically modifying the virtual reality environment based on the recognized emotions" refers to a process or device that analyzes a user's facial expressions and voice to determine their emotions and adjusts the VR experience to match those emotions.

[0377] "Means for collecting user feedback and improving the system" refers to a process or device that collects opinions and ratings provided by users after their experience and uses that data to update the system's functionality and performance.

[0378] The present invention is a system that recognizes a user's emotions and dynamically modifies a virtual reality experience based on the emotions. The system can generate experience content and render a virtual reality environment in real time based on the user's requests.

[0379] Hardware and software used

[0380] Hardware:

[0381] PC or smartphone with a camera

[0382] Virtual reality headsets (e.g., general VR devices)

[0383] software:

[0384] OpenCV library for emotion recognition

[0385] Hugging Face Transformers for Scenario Generation

[0386] Custom VR engine for real-time rendering (e.g. Unreal Engine Server)

[0387] API for collecting internet information

[0388] System Program Overview

[0389] Based on the user's request, the server converts the request into text data. Requests made by voice input are converted into text data using a highly accurate voice recognition library. Based on the text data, the server collects related data from the Internet. This is done using a specific search API to collect the required information.

[0390] The collected data is passed to a generative AI model, which generates scenarios using Hugging Face's Transformers. The generated scenarios are then rendered into a 3D environment using a virtual reality engine (e.g., an Unreal Engine server).

[0391] The rendered 3D environment data is sent to the device and displayed on the user's virtual reality headset. The 3D environment is updated in real time according to the user's movements. This updating allows the user to move freely around the VR space.

[0392] The system uses a camera and microphone to capture the user's facial expressions and voice, and then analyzes their emotions in real time using an emotion recognition engine (OpenCV). Based on the analyzed emotional data, the virtual reality environment is dynamically modified. For example, if the user is excited, special effects are added in the environment, and if the user is relaxed, the atmosphere of the environment is made calmer.

[0393] Finally, after the user has finished the experience, the terminal displays a feedback form where the user can enter comments and suggestions about the experience. This feedback information is sent to the server and used to continuously improve the system.

[0394] Specific examples

[0395] A user requests a virtual shopping store by speaking, "I want to buy new shoes." This request is converted into text data, and the server collects related information. Based on the collected data, an AI model generates a shopping scenario, and a VR engine renders a 3D environment. When the user puts on a virtual reality headset and enters the virtual store, a camera captures the user's facial expressions, and emotional data is analyzed. If the user finds a pair of shoes and expresses joy, the system displays a pop-up in the storefront offering a special discount.

[0396] Prompt Sentence Examples

[0397] User request: "I want to buy new shoes from a virtual shopping store."

[0398] AI-generated prompt: "Generate a scenario in which a user searches for shoes in a virtual shopping store."

[0399] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0400] Step 1:

[0401] The user speaks the request

[0402] Specific operation: The user says to the terminal, "I want to buy new shoes at a virtual shopping store."

[0403] Input: Audio data

[0404] Data calculation: Converts voice data into text data using a voice recognition library.

[0405] Output: Text data (e.g., "I want to buy new shoes at a virtual shopping store")

[0406] Step 2:

[0407] Sends requests to the server and collects relevant data

[0408] Specific operation: The terminal sends the converted text data to the server as an HTTP request, and the server searches the Internet for information related to the request content.

[0409] Input: Text data (request content)

[0410] Data Computing: Collecting relevant information from multiple databases and sources on the Internet.

[0411] Output: Collected relevant data (e.g. product information, images, reviews)

[0412] Step 3:

[0413] Analyze the collected data and format it appropriately

[0414] Specific operation: The server analyzes the collected data and formats it appropriately to be passed to the generative AI model.

[0415] Input: Relevant data collected

[0416] Data operations: Data analysis and formatting (e.g., text analysis, data structuring)

[0417] Output: Formatted data

[0418] Step 4:

[0419] Generate a scenario

[0420] Specific operation: Based on the collected data, the server sends prompts to the generative AI model (Hugging Face Transformers) to generate a shopping scenario.

[0421] Input: Formatted data, prompt (e.g., "Generate a scenario in which a user searches for shoes in a virtual shopping store.")

[0422] Data calculation: scenario generation using generative AI models

[0423] Output: Generated scenario

[0424] Step 5:

[0425] Rendering a Virtual Reality Environment

[0426] How it works: The server sends the generated scenario to a VR engine (e.g., Unreal Engine server), which renders the 3D virtual environment in real time.

[0427] Input: Generated scenario

[0428] Data Computing: 3D Modeling and Rendering with a VR Engine

[0429] Output: 3D environment data

[0430] Step 6:

[0431] 3D environment data is sent to the user's device and displayed

[0432] Specific operation: The server sends the rendered 3D environment data to the user's device, which then displays the data on the VR headset.

[0433] Input: 3D environment data

[0434] Data Computing: Communication and Data Transfer

[0435] Output: 3D virtual environment displayed in a VR headset

[0436] Step 7:

[0437] Update the 3D environment in real time according to the user's movements

[0438] How it works: As the user moves around and looks at objects in the VR space, the environment is updated in real time. Sensors capture movements, and the server calculates and transmits the corresponding environmental changes.

[0439] Input: User movement data (e.g., position change, gaze direction)

[0440] Data Computation: Analyzing motion data and re-rendering the environment

[0441] Output: Updated 3D environment data

[0442] Step 8:

[0443] Recognize user emotions and dynamically change the environment

[0444] How it works: The camera captures the user's facial expressions, and the emotion recognition engine (OpenCV) analyzes them to recognize their emotions. The server dynamically adjusts the virtual reality environment based on this emotional data.

[0445] Input: User's facial expression data, voice tone

[0446] Data Computing: Facial Expression Analysis and Emotion Recognition

[0447] Output: Change the environment based on the emotion (e.g., special discount popup, adding interactive elements)

[0448] Step 9:

[0449] Collect feedback after the experience and refine the system

[0450] Specific operation: After the user finishes the experience, the device displays a feedback form and sends the user's comments and requests for improvement to the server. The server analyzes this data and uses it to improve the system.

[0451] Input: User feedback data (e.g., satisfaction, areas for improvement)

[0452] Data Computing: Feedback Analysis and System Optimization

[0453] Output: Improved system functionality and performance

[0454] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0455] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0456] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0457] [Second embodiment]

[0458] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0459] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0460] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0461] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0462] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0463] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0464] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0465] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0466] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0467] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0468] In the smart glasses 214, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0469] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0470] This invention provides a system that allows users to experience specific historical moments in the past or future through virtual reality. Based on user requests, this system collects relevant data from the Internet, generates scenarios using generative artificial intelligence, and provides a mechanism for rendering 3D environments in real time using a virtual reality engine.

[0471] Program Overview

[0472] 1. User submits request:

[0473] The user speaks into the device to request an experience of a specific time in the past or future.

[0474] The device analyzes the voice and converts it into text data.

[0475] 2. Data Collection:

[0476] The terminal transmits the text data to the server.

[0477] The server collects relevant information from the Internet based on the request.

[0478] 3. Data analysis and scenario generation:

[0479] The server analyzes the collected data and generates scenarios using generative artificial intelligence.

[0480] A scenario details the buildings, people, and events associated with a specified time period and place.

[0481] 4. Real-time rendering:

[0482] The server uses the generated scenario to render a 3D environment in a virtual reality engine.

[0483] The 3D environment contains details and realism that match the user's requirements.

[0484] 5. Data transmission and display:

[0485] The server sends the rendered 3D environment data to the terminal.

[0486] The device displays the received data on a virtual reality headset.

[0487] 6. Immersive user experience:

[0488] The user wears a virtual reality headset and is immersed in a virtual environment.

[0489] The environment is updated in real time according to the user's movements and gaze.

[0490] 7. Feedback collection and system improvement:

[0491] After the user has completed the experience, the terminal displays a feedback form in which the user can enter comments about the experience and requests for improvements.

[0492] The device sends feedback to the server, which analyzes it and uses it to improve the system.

[0493] Specific examples

[0494] Case: Renaissance Florence

[0495] 1. Submit your request:

[0496] The user speaks to the device and says, "I want to experience Renaissance Florence."

[0497] The terminal converts this voice into text data and sends it to the server.

[0498] 2. Data Collection:

[0499] The server collects information related to "Renaissance Florence" from the Internet.

[0500] For example, data such as text, images, audio, and video related to buildings, people, and events from that time is obtained.

[0501] 3. Data analysis and scenario generation:

[0502] The server analyzes the collected data and uses generative artificial intelligence to generate a scenario of "Florence during the Renaissance."

[0503] Scenarios include Da Vinci painting and the construction site of Florence Cathedral.

[0504] 4. Real-time rendering:

[0505] Based on the generated scenario, the server uses a virtual reality engine to render a 3D environment of Florence in real time.

[0506] 5. Data transmission and display:

[0507] The server sends the rendered 3D environment data to the terminal.

[0508] The device receives the data and displays it on a virtual reality headset.

[0509] 6. Immersive Experience:

[0510] Users put on a headset and are immersed in a realistic Renaissance Florence.

[0511] Users can interact with Da Vinci and tour the inside of the cathedral.

[0512] 7. Feedback and System Improvement:

[0513] After completing the experience, users fill out a feedback form to indicate their satisfaction with the experience and suggestions for improvement.

[0514] The device sends this to the server, which analyzes the feedback and uses it to improve the system.

[0515] This allows users to experience historical moments and future worlds in a highly immersive way.

[0516] The processing flow will be explained below.

[0517] Step 1:

[0518] The user makes a request to experience a specific time period in the past or future by voice into the device, which then recognizes the voice and converts the voice data into text data.

[0519] Step 2:

[0520] The device sends the converted text data to the server as an HTTP request, which includes the user's desired experience.

[0521] Step 3:

[0522] Based on the received request, the server launches a scraping tool to scrape relevant text, image, audio, and video data from the Internet.

[0523] Step 4:

[0524] The server uses a scraping tool to collect data from the Internet that matches the specified keywords, and the collected data is temporarily stored.

[0525] Step 5:

[0526] The server analyzes the collected data using natural language processing and image recognition technology and formats it appropriately to be passed to the generative artificial intelligence.

[0527] Step 6:

[0528] The server inputs the formatted data into a generative AI to generate a scenario for the specified time period and location, including detailed historical background and key buildings, people, and events.

[0529] Step 7:

[0530] The server uses an evaluation algorithm to check the quality of the generated scenario, correcting or regenerating it as necessary. Once the evaluation is complete, the scenario moves on to the next stage.

[0531] Step 8:

[0532] The server uses a virtual reality engine to render a 3D environment in real time based on the scenario, using the GPU to achieve high performance and detailed rendering.

[0533] Step 9:

[0534] The server compresses the rendered 3D environment data and transmits it to the device in an efficient binary format, using streaming techniques to minimize data latency.

[0535] Step 10:

[0536] The device then extracts the received 3D environment data and displays it on the virtual reality headset, which updates the displayed 3D environment in real time according to the user's movements and gaze.

[0537] Step 11:

[0538] Users wear a virtual reality headset and enter an immersive virtual environment, where they can freely move around and experience interactive elements (e.g., interact with people, manipulate objects).

[0539] Step 12:

[0540] When the user finishes their experience, the device displays a feedback form where the user can enter comments about their experience and suggestions for improvement. Collecting feedback contributes to improving user satisfaction.

[0541] Step 13:

[0542] The device sends the collected feedback data to the server, which analyzes the feedback and adjusts the algorithms of the generative AI and virtual reality engine to help improve the system, resulting in an even better user experience next time.

[0543] Example 1

[0544] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0545] Conventional virtual reality systems face challenges when it comes to quickly and accurately generating detailed scenarios and rendering 3D environments in real time when users want to experience specific historical moments or future worlds. Other challenges include updating the 3D environment in real time in response to user movements and systematically collecting user feedback to improve the system.

[0546] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0547] In this invention, the server includes means for converting user requests into text data based on the user requests, means for collecting related information from a communication network, means for generating a scenario using artificial intelligence based on the collected information, means for creating a 3D environment using virtual reality technology based on the generated scenario, means for transmitting the generated 3D environment data to a user device, means for displaying the 3D environment on the user device and updating it in real time according to the user's movements, and means for collecting user responses and improving the system, thereby enabling users to experience specific historical moments or future worlds in detail and realistically.

[0548] "User" refers to a person who uses the system to experience virtual reality.

[0549] A "request" refers to a request for a specific historical moment or future world that a user would like to experience.

[0550] "Text data" refers to data in which a user's request is converted into text form through voice recognition or other input methods.

[0551] "Communications network" means a wide-area computer network, including the Internet and other data communications infrastructures.

[0552] "Information" refers to data such as text, images, and videos collected based on user requests.

[0553] "Artificial intelligence" refers to machine learning algorithms and neural networks used to analyze collected information and generate specified scenarios.

[0554] A "scenario" refers to a set of information that contains a detailed description of a particular historical moment or future world that a user wants to experience.

[0555] "Virtual reality technology" refers to technology that uses computer technology to generate a 3D environment and provide an immersive experience to the user.

[0556] "3D environment" refers to a three-dimensional virtual space created using virtual reality technology based on a scenario.

[0557] "User device" refers to the terminal or headset that a user uses to have a virtual reality experience.

[0558] "Real-time updates" means that the 3D environment is instantly reflected and updated in response to inputs such as the user's movements and gaze.

[0559] "Reactions" refer to the feedback and opinions users provide through their virtual reality experience.

[0560] "System" refers to the comprehensive technological platform that includes all of the above means.

[0561] The present invention is a system that allows users to experience specific historical moments in the past or future through virtual reality. Based on a user's request, the system collects relevant information from the internet, generates a scenario using a generative AI model, and uses virtual reality technology to render a 3D environment in real time.

[0562] First, a user requests a historical moment or future world they would like to experience through a voice request. For example, a user might say, "I want to experience Tokyo of the future." This voice request is converted into text data by the device's voice recognition software (e.g., Google Speech-to-Text API).

[0563] The device sends the converted text data to a server, which then collects related information from the Internet via a communications network. Specifically, it uses web scraping tools (e.g., BeautifulSoup) or APIs (e.g., Wikipedia API) to obtain the necessary text, images, videos, etc.

[0564] The collected information is analyzed by the server and passed to a generative AI model (e.g., OpenAI's GPT-4). This generative AI model creates a detailed scenario based on the collected information. For example, it uses a prompt such as, "Generate a detailed scenario about Tokyo in the future. Include skyscrapers, a modern transportation system, and people's activities."

[0565] Based on the generated scenario, the server uses virtual reality technology to render a 3D environment in real time using a virtual reality engine (e.g., Unreal Engine or Unity). Specifically, the future Tokyo cityscape, transportation system, building interiors, streetscapes, and other details are depicted.

[0566] The generated 3D environment data is sent to the terminal and displayed on the user device (e.g., a virtual reality headset). The user wears the headset and is immersed in the virtual environment. The user's movements and gaze are detected by tracking sensors (e.g., the tracking system in the HTC Vive), and the 3D environment is updated in real time. This allows the user to freely explore the virtual reality and enjoy a detailed experience.

[0567] Finally, once the user has finished the experience, the device displays a feedback form. The user can enter comments about the experience or suggestions for improvement, which the device then sends to the server. The server analyzes the collected feedback and incorporates it into the next update or improvement. This feedback is then used with natural language processing tools and machine learning algorithms to continuously improve the system.

[0568] As described above, the present invention provides a system that provides a user with a high level of immersion and allows them to experience specific historical moments in the past or future in detail and realistically.

[0569] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0570] Step 1: User submits request

[0571] Users speak their requests into the device, for example, saying, "I want to experience Renaissance Florence."

[0572] The device receives this voice input and converts it into text using voice recognition software, specifically the Google Speech-to-Text API.

[0573] Input: User's voice request

[0574] Output: The request converted to character data

[0575] Step 2: Data collection

[0576] The terminal transmits the converted character data to the server.

[0577] Based on the received request, the server collects relevant information from the Internet via a communication network, for example, using a web scraping tool (BeautifulSoup) or an API (Wikipedia API).

[0578] Input: Request converted to character data

[0579] Output: Collected information (text, images, videos, etc.)

[0580] Step 3: Data analysis and scenario generation

[0581] The server analyzes the collected information and organizes it using a natural language processing tool (spaCy).

[0582] The server inputs a prompt into a generative AI model (e.g., OpenAI's GPT-4) to generate a detailed scenario. An example of the prompt is, "Generate a scenario set in Renaissance Florence, including a scene where Da Vinci is painting."

[0583] Input: Collected information

[0584] Output: Generated scenario

[0585] Step 4: Real-time rendering

[0586] Based on the generated scenario, the server uses virtual reality technology to render a 3D environment in real time, specifically using a virtual reality engine (Unreal Engine or Unity).

[0587] Input: Generated scenario

[0588] Output: Rendered 3D environment

[0589] Step 5: Send and display data

[0590] The server sends the rendered 3D environment data to the device, using WebSocket or HTTP / 2 as the communication protocol.

[0591] The device prepares the received data for display on a virtual reality headset.

[0592] Input: Rendered 3D environment data

[0593] Output: Display to a virtual reality headset

[0594] Step 6: User Immersion

[0595] The user wears a virtual reality headset and is immersed in a virtual environment.

[0596] The device uses sensors to track the user's movements and gaze, updating the 3D environment in real time, for example using the tracking system in the HTC Vive.

[0597] Input: User movement and gaze data

[0598] Output: Real-time updated 3D environment

[0599] Step 7: Gather feedback and improve the system

[0600] After the user has finished the experience, the terminal displays a feedback form.

[0601] Users enter comments about their experience and requests for improvement, and the device sends these to the server.

[0602] The server analyzes the collected feedback and uses natural language processing tools and machine learning algorithms to help improve the system.

[0603] Input: User feedback

[0604] Output: Analyzed feedback and system improvements

[0605] The above is the specific processing flow of this system.

[0606] (Application example 1)

[0607] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0608] Current virtual reality systems lack efficient means for users to experience specific historical moments or future events with a high level of realism. Furthermore, technologies for automatically generating scenarios in real time based on collected data and then generating and updating virtual reality environments in real time based on those scenarios remain a challenge. Therefore, systems that improve the quality of the user experience are needed.

[0609] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0610] In this invention, the server includes means for converting a user's request into text data based on the request, means for collecting relevant data from the Internet, means for generating a scenario using a generating artificial intelligence based on the collected data, means for rendering a three-dimensional environment using a virtual reality engine based on the generated scenario, means for transmitting the generated three-dimensional environment data to a user terminal, means for displaying the three-dimensional environment on the user terminal and updating it in real time according to the user's movements, means for collecting user feedback and improving the system, means for providing the generated three-dimensional environment to the user using a visual display device, means for analyzing the collected data and formatting it into an appropriate format for passing it to the generating artificial intelligence, means for ensuring the quality of the scenario using an algorithm for evaluating the quality of the generated scenario, means for rendering the generated three-dimensional environment in real time using a virtual reality engine and transmitting it to a visual display device, and means for displaying the generated three-dimensional environment in real time on the user's mobile device or head-mounted display and updating the environment in accordance with the user's line of sight and movements, thereby enabling users to experience historical moments and future events in real time with a high level of realism.

[0611] A "user request" refers to a user's desire to experience a particular historical moment or future event through audio or text.

[0612] "Text data" refers to data obtained by converting a user's request from voice into text information.

[0613] "Related Data" refers to information about a specified time and place that is collected from the Internet based on a user's request.

[0614] "Generative AI" refers to artificial intelligence technology that generates scenarios based on collected data.

[0615] A "scenario" is a plan or script that details the events and settings to be experienced within a virtual reality environment based on generated data.

[0616] "Virtual reality engine" refers to a software and hardware configuration for rendering three-dimensional environments in real time.

[0617] "Three-dimensional environment" refers to the virtual space that a user experiences through virtual reality.

[0618] "User terminal" refers to a device including a smartphone, a head-mounted display, or other display device.

[0619] "Visual display device" refers to a display device used by a user to visually experience a three-dimensional environment.

[0620] "Feedback" refers to comments, ratings, and requests for improvement regarding the experience provided by the user after the experience.

[0621] "Formatting" refers to the process of converting collected data into an appropriate form for use by the generative AI.

[0622] "Algorithm" refers to a computational procedure for assessing and ensuring the quality of generated scenarios.

[0623] The present invention provides a system that allows a user to experience a specific historical moment or a future event through virtual reality. The system comprises the following means:

[0624] First, the user inputs a request by voice or text using a device such as a smartphone or head-mounted display. For example, the user's request might be something like, "I want to experience a Renaissance city." This request is converted into text data using speech recognition technology (for example, Python's speech_recognition library).

[0625] The server then retrieves the converted text data and collects relevant data from the Internet. During this process, the required data is retrieved from the source using, for example, the requests library. The collected data is then analyzed using generative artificial intelligence (e.g., OpenAI's API) and formatted accordingly.

[0626] The server generates a scenario based on the formatted data. The scenario generated using the AI ​​generator includes detailed events and settings related to the specified time period and location. When generating the scenario, a prompt message corresponding to the user's request is passed to the AI ​​generator. For example, the following prompt message is used:

[0627] Example prompt:

[0628] A user has requested to experience a Renaissance city. Generate a scenario based on the following relevant data:

[0629] Renaissance buildings and streets

[0630] Costumes and culture of the time

[0631] Historical events and everyday scenes

[0632] Generate detailed scenarios and make them suitable for virtual reality experiences.

[0633] The generated scenario is rendered in real time as a 3D environment using a virtual reality engine (e.g., a 3D engine such as py3dengine). The rendered 3D environment data is sent from the server to the user's device, and the user is immersed in the virtual environment through a visual display device (e.g., a head-mounted display or a smartphone).

[0634] After the user has completed their experience in the virtual environment, the server collects feedback from the user by entering comments and ratings into a dedicated form, which is used to improve the system.

[0635] For example, if a user requests "I want to experience a Renaissance city," the server collects relevant data and uses generative artificial intelligence to generate a specific Renaissance scenario, including detailed descriptions of buildings, cityscapes, and costumed people in action. The generated scenario is then rendered as a three-dimensional environment by a virtual reality engine, allowing the user to use a visual display device to tour the city and interact with historical figures.

[0636] As described above, the present embodiment provides a series of processes for enhancing a user's immersive experience.

[0637] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0638] Step 1:

[0639] The user speaks a request into the device. This request describes a desire to experience a specific historical moment or future event. The device recognizes this speech as input data and converts it into text data using speech recognition technology (for example, Python's speech_recognition library). The output is the user's request converted into text data as character information.

[0640] Step 2:

[0641] The terminal sends the converted text data to the server. The server receives the text data as input and collects related data from the Internet. Here, it uses the requests library to obtain the necessary data from the target sources. The output is the collected related data.

[0642] Step 3:

[0643] The server takes the collected data as input, analyzes it using generative AI (for example, OpenAI's API), and formats it into an appropriate format. During this process, the data is structured and necessary information is filtered. The output is data formatted in an appropriate format for the generative AI to generate scenarios.

[0644] Step 4:

[0645] The server uses the formatted data to send prompts to the generation AI to generate a scenario. The prompt text is a specific scenario request based on the user's request and the collected data. The generation AI generates a scenario based on the prompt and provides the scenario in text format as output.

[0646] Step 5:

[0647] The server takes the generated scenario as input and uses a virtual reality engine (e.g., py3dengine) to render a 3D environment in real time. During this process, a 3D environment is generated according to the generated scenario. The output is 3D data rendered as a virtual reality environment.

[0648] Step 6:

[0649] The server transmits the rendered 3D environment data to the user's device, which receives this data as input and displays the virtual environment to the user through a visual display device (such as a head-mounted display or a smartphone). The user is immersed in this virtual environment and experiences it.

[0650] Step 7:

[0651] After the user has completed their experience in the virtual environment, the device displays a feedback form. The user enters comments, ratings, and suggestions for improvement about the experience. The device sends this as input data to the server. The server analyzes the feedback and uses it to improve the system. The output is improvements based on the feedback.

[0652] This series of processing steps allows users to experience historical moments and future events in real time with a high level of realism.

[0653] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0654] This invention is a system that recognizes a user's emotions in a virtual reality environment and dynamically changes the experience content based on those emotions. Based on the user's request, the system converts the request into text data, collects related data from the Internet, and generates a scenario using generative artificial intelligence. Based on the generated scenario, a 3D environment is rendered using a virtual reality engine and displayed on the user's device. Furthermore, by updating the environment in real time based on the user's movements and combining it with an emotion engine that recognizes the user's emotions, an experience tailored to the user's emotions is provided.

[0655] Program Overview

[0656] 1. User submits request:

[0657] The user requests an experience of a specific time in the past or future by speaking into the device, which then analyzes the voice and converts the voice data into text data.

[0658] 2. Data Collection:

[0659] The device sends text data to the server as an HTTP request, and the server collects related information from the Internet based on the request content.

[0660] 3. Data analysis and scenario generation:

[0661] The server analyzes the collected data and uses generative artificial intelligence to generate scenarios that include detailed descriptions of historical background, key buildings, people, and events.

[0662] 4. Real-time rendering:

[0663] The server uses the generated scenario to render a 3D environment in real time in a virtual reality engine, with details and realism that match the user's requirements.

[0664] 5. Data transmission and display:

[0665] The server sends the rendered 3D environment data to the device, which displays it on a virtual reality headset.

[0666] 6. Immersive user experience:

[0667] Users wear a virtual reality headset and are immersed in a virtual environment that updates in real time based on the user's movements and gaze.

[0668] 7. Emotion Recognition with Emotion Engine:

[0669] The server uses an emotion engine that analyzes the user's facial expressions and vocal tone to recognize the user's emotions in real time. The recognized emotion data is analyzed and reflected in the content of the virtual environment.

[0670] 8. Dynamic Environment Changes:

[0671] The server dynamically modifies the virtual reality environment based on the emotional data recognized by the emotion engine, for example by increasing the effects in the environment if the user is surprised, or by adding interactive elements if the user is engrossed.

[0672] 9. Feedback Collection and System Improvement:

[0673] After the user has finished the experience, the device displays a feedback form, where the user can enter comments about the experience and suggestions for improvement. The device then sends this feedback to the server, which then analyzes the feedback and emotional data to help improve the system.

[0674] Specific examples

[0675] Case: Renaissance Florence

[0676] 1. Submit your request:

[0677] The user speaks to the device, saying, "I want to experience Renaissance Florence." The device converts this speech into text data and sends it to the server.

[0678] 2. Data Collection:

[0679] The server collects information related to "Renaissance Florence" from the Internet, including text, images, audio, and video data related to buildings, people, and events from that time.

[0680] 3. Data analysis and scenario generation:

[0681] The server analyzes the collected data and uses generative artificial intelligence to generate a "Renaissance Florence" scenario, including scenes of Da Vinci painting and the construction site of Florence Cathedral.

[0682] 4. Real-time rendering:

[0683] Based on the generated scenario, the server uses a virtual reality engine to render a 3D environment of Florence in real time.

[0684] 5. Data transmission and display:

[0685] The server sends the rendered 3D environment data to the device, which receives the data and displays it on the virtual reality headset.

[0686] 6. Immersive Experience:

[0687] Users put on a headset and are immersed in a realistic Renaissance Florence, where they can interact with Da Vinci and tour the interior of the cathedral.

[0688] 7. Emotion Recognition with Emotion Engine:

[0689] The server analyzes the user's facial expressions and voice using an emotion engine and recognizes that the user is excited.

[0690] 8. Dynamic Environment Changes:

[0691] Based on the user's emotional data, the server adds special events (e.g., fireworks and musical performances) to the streets of Florence to make the experience even more engaging.

[0692] 9. Feedback and System Improvement:

[0693] After completing the experience, the user fills in a feedback form to indicate their satisfaction with the experience and any suggestions for improvement. The device then sends this information to the server, which then analyzes the feedback and emotional data and reflects it in improvements to the system.

[0694] In this way, the system of the present invention recognizes the user's emotions and dynamically modifies the virtual reality experience based on them, providing a richer and more engaging experience.

[0695] The processing flow will be explained below.

[0696] Step 1:

[0697] The user speaks into the device, saying, "I want to experience Renaissance Florence." The device uses voice recognition software to convert the voice data into text.

[0698] Step 2:

[0699] The device sends the converted text data to the server as an HTTP request, which includes the user's desired experience.

[0700] Step 3:

[0701] Based on the received request, the server uses a scraping tool to collect related information (text, image, audio, and video data) that matches the specified keywords from the Internet.

[0702] Step 4:

[0703] The server analyzes the scraped data using natural language processing and image recognition technology and formats it appropriately to be passed on to generative artificial intelligence (AI).

[0704] Step 5:

[0705] The server inputs the formatted data into a generative artificial intelligence (AI) that generates a scenario about "Renaissance Florence," including detailed descriptions of buildings, people, and events.

[0706] Step 6:

[0707] The server uses an evaluation algorithm to verify the quality of the generated scenarios, and corrects or regenerates them as necessary.

[0708] Step 7:

[0709] The server uses quality-tested scenarios to leverage a virtual reality engine to render 3D environments in real time.

[0710] Step 8:

[0711] The server compresses the rendered 3D environment data in binary format and transmits it to the terminal using efficient data transfer techniques.

[0712] Step 9:

[0713] The device receives 3D environment data from the server and displays it on the virtual reality headset. The device tracks the user's gaze and movements in real time and updates the environment.

[0714] Step 10:

[0715] Users put on a virtual reality headset and begin experiencing a realistic Renaissance Florence. They can freely explore the virtual environment and interact with the characters.

[0716] Step 11:

[0717] The server uses an emotion engine to collect the user's facial expressions and voice tone in real time from the camera and microphone installed in the headset and analyze the user's emotions.

[0718] Step 12:

[0719] The server dynamically modifies the virtual reality environment based on the analyzed emotional data, for example adding effects and interactive elements if the user is excited.

[0720] Step 13:

[0721] After the user has finished the experience, the terminal displays a feedback form, where the user can enter comments about the experience and requests for improvements. The terminal then sends the feedback to the server.

[0722] Step 14:

[0723] The server analyzes the feedback and emotional data, adjusts the algorithms of the generative AI and virtual reality engine, and reflects the results in system improvements, resulting in an even better user experience next time.

[0724] Example 2

[0725] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0726] Conventional virtual reality systems lack the flexibility to adapt to user requests and the ability to dynamically change the environment based on the user's emotions. These shortcomings limit the quality of the user experience and make it difficult to provide a richer and more engaging virtual reality experience. Therefore, there is a need for a system that dynamically changes the generation of scenarios and rendering of virtual reality environments based on user requests, depending on the user's emotions.

[0727] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0728] In this invention, the server includes: means for converting a user's request into text data based on the request; means for collecting related data from the Internet; means for generating a scenario using a generative model based on the collected data; means for rendering a three-dimensional environment using a virtual reality engine based on the generated scenario; means for transmitting the generated three-dimensional environment data to a user terminal; means for displaying the three-dimensional environment on the user terminal and updating it in real time according to the user's movements; means including an emotion analysis engine for recognizing the user's emotions; means for dynamically changing the virtual reality environment based on the recognized emotion data; and means for collecting user feedback and improving the system. This makes it possible to flexibly generate scenarios according to user requests and dynamically change the virtual reality environment according to the user's emotions.

[0729] "Means for converting a request into text data" refers to a function for analyzing a request input by a user in a voice or other format and converting it into text.

[0730] "Means of collecting relevant data from the Internet" refers to the function of obtaining necessary information from websites and databases via a network.

[0731] A "generative model" refers to an artificial intelligence algorithm that generates appropriate results for a given input (prompt).

[0732] "Means for generating a scenario" refers to the function of creating stories and events for a virtual reality experience that responds to a user's request based on collected data.

[0733] "Virtual reality engine" refers to a software platform for generating and rendering three-dimensional virtual environments in real time.

[0734] "Means for rendering a three-dimensional environment" refers to a function for depicting a three-dimensional visual environment based on a generated scenario.

[0735] "Means for transmitting three-dimensional environmental data to a user terminal" refers to a function for transferring three-dimensional visual data generated by the server to the terminal used by the user.

[0736] "Means for displaying a three-dimensional environment and updating it in real time in response to the user's movements" refers to a function that dynamically changes the three-dimensional virtual environment displayed on the user's terminal in synchronization with the user's movements and operations.

[0737] An "emotion analysis engine" refers to software that analyzes a user's facial expressions and vocal tone to identify emotions.

[0738] "Means for dynamically modifying the virtual reality environment based on recognized emotional data" refers to the ability to adjust the content and events of the virtual reality in real time based on data obtained through emotional analysis.

[0739] "Means for collecting feedback and improving the system" refers to a function for incorporating user evaluations and opinions of the experience and improving the experience in future.

[0740] The present invention is a system that generates a scenario based on a user's request in a virtual reality environment and provides the user with a three-dimensional environment corresponding to that scenario. This invention makes it possible to recognize the user's emotions in real time and dynamically change the virtual reality environment based on those emotions.

[0741] First, the user speaks into the device to request an experience of a specific time or place. The device then uses voice recognition software, such as the Google Speech-to-Text API, to convert this voice data into text. This text data then becomes the basis for interpreting the request.

[0742] The converted text data is then sent to the server as an HTTP request, which uses web crawling tools and APIs to gather relevant information from the internet, such as Beautiful Soup and Scrapy, to gather text, images, and audio data.

[0743] The collected data is then analyzed using natural language processing tools, such as spaCy and NLTK. Based on the results of this analysis, a generative AI model (e.g., GPT-4, ChatGPT) generates a scenario. Input prompts are crucial for scenario generation. For example, a prompt such as "Walk through the streets of Renaissance Florence" generates detailed historical background and character dialogue.

[0744] Based on the generated scenario, the server uses a virtual reality engine (e.g., Unity, Unreal Engine) to render a three-dimensional virtual environment. The rendered data is sent to the user's device via HTTP. The user's device receives this data and displays it on a virtual reality headset (e.g., Oculus Rift, HTC Vive). The user can then immerse themselves in and interact with this virtual environment.

[0745] During the user's experience, the device uses a camera and microphone to capture the user's facial expressions and voice. This data is analyzed in real time by an emotion analysis engine (e.g., Affectiva, Kairos). Based on the user's emotion data, the server dynamically modifies the virtual environment. For example, if the user is surprised, effects such as fireworks are added to the virtual environment.

[0746] After the experience is over, the device displays a feedback form to the user, allowing them to input their satisfaction with the experience and suggestions for improvement. This data is then sent to the server, which analyzes the feedback and emotional data to improve the experience for the next session.

[0747] A concrete example is a user request to "experience Renaissance Florence." In this case, the user speaks the request into the device, which converts this speech into text data and sends it to the server. The server collects relevant information from the Internet and generates a scenario using a generative AI model. It then renders a three-dimensional environment of Florence using a virtual reality engine and sends it to the user's device. The user puts on the virtual reality headset and explores the city of Florence. If the user is excited, the server adds special events to make the experience even more engaging.

[0748] This invention makes it possible to flexibly generate scenarios based on user requests and dynamically change the virtual reality environment in response to the user's emotions.

[0749] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0750] Step 1:

[0751] User request submission

[0752] Input: The user speaks into the device, "I want to experience Renaissance Florence."

[0753] How it works: The device uses speech recognition software (e.g., Google Speech-to-Text API) to convert speech into text. Voice input is text, and output is text.

[0754] Output: The converted text data. The text is "I want to experience Renaissance Florence."

[0755] Step 2:

[0756] Data collection

[0757] Input: The request as text data.

[0758] How it works: The device sends this text data to the server as an HTTP request. The server then uses a web crawling tool (e.g., Beautiful Soup, Scrapy) to collect relevant information from the Internet. This includes data such as text, images, and audio.

[0759] Output: A dataset containing relevant information.

[0760] Step 3:

[0761] Data analysis and scenario generation

[0762] Input: A dataset containing the relevant collected information.

[0763] How it works: The server uses natural language processing tools (e.g., spaCy, NLTK) to parse the data and format it into a suitable format for a generative AI model (e.g., GPT-4, ChatGPT). It then generates a scenario based on the generative AI model. For example, you can input the prompt "Florence during the Renaissance," and the scenario will be output.

[0764] Output: Generated scenario data.

[0765] Step 4:

[0766] Real-time rendering

[0767] Input: Generated scenario data.

[0768] How it works: The server uses a virtual reality engine (e.g., Unity, Unreal Engine) to render a 3D environment in real time. Based on the scenario, models of objects and characters in the environment are generated and rendered in real time.

[0769] Output: Rendered 3D environment data.

[0770] Step 5:

[0771] Sending and Displaying Data

[0772] Input: Rendered 3D environment data.

[0773] How it works: The server sends this data to the user's device, which then displays it on a virtual reality headset (e.g., Oculus Rift, HTC Vive).

[0774] Output: A three-dimensional environment displayed in a virtual reality headset.

[0775] Step 6:

[0776] Immersive User Experience

[0777] Input: A three-dimensional environment displayed on a virtual reality headset.

[0778] How it works: The user puts on a virtual reality headset and is immersed in a virtual environment. The environment updates in real time based on the user's movements and gaze. For example, the user can explore and interact with Renaissance Florence.

[0779] Output: Updated virtual environment and user experience data.

[0780] Step 7:

[0781] Emotion recognition by emotion engine

[0782] Input: User's facial expressions and voice.

[0783] How it works: The device captures the user's facial expressions and voice using a camera and microphone. The server analyzes this data in real time using an emotion analysis engine (e.g., Affectiva, Kairos) to recognize the user's emotions.

[0784] Output: Recognized emotion data.

[0785] Step 8:

[0786] Dynamic Environment Changes

[0787] Input: Recognized emotion data.

[0788] Action: The server dynamically changes the virtual reality environment based on the emotional data, for example adding special events or effects if the user is excited.

[0789] Output: A dynamically modified virtual reality environment.

[0790] Step 9:

[0791] Feedback collection and system improvement

[0792] Input: User feedback after the experience.

[0793] How it works: After the experience is over, the device displays a feedback form to the user. The user enters their satisfaction with the experience and suggestions for improvement, and this data is sent to the server. The server analyzes the feedback and emotion data to help improve the system.

[0794] Output: An improved system and updated user experience data.

[0795] (Application example 2)

[0796] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0797] Conventional virtual reality systems provide a fixed experience without taking the user's emotions into consideration. This means that the experience cannot be dynamically changed based on the user's emotions and reactions, resulting in a lack of immersion and personalization. Furthermore, it has not been possible for the system to recognize the user's emotions during the experience and provide an optimized experience based on those emotions. This has led to issues such as low user satisfaction and engagement.

[0798] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0799] In this invention, the server includes means for converting a user's request into text data based on the user's request, means for collecting related data from the Internet, means for generating a scenario based on the collected data using artificial intelligence, means for rendering a 3D environment based on the generated scenario using a virtual reality engine, means for transmitting the generated 3D environment data to a user terminal, means for displaying the 3D environment on the user terminal and updating it in real time according to the user's movements, means for recognizing the user's emotions and dynamically changing the virtual reality environment based on the recognized emotions, and means for collecting user feedback and improving the system. This enables dynamic changes to the virtual reality experience based on the user's emotions, providing a more immersive and personalized experience.

[0800] A "user request" is a request to the system for content or information that the user wants to experience.

[0801] A "means for converting to text data" is a process or device that analyzes and converts voice or other non-text data into written information.

[0802] "Means for collecting relevant data from the Internet" refers to a process or device that searches for and obtains specified information or data via the Internet.

[0803] "Generative AI" is an AI technology that generates scenarios and content for specific purposes based on collected data.

[0804] A "scenario generation means" is a process or device that analyzes collected data and arranges it into a specific story or scenario format.

[0805] A "virtual reality engine" is software or hardware that uses computer graphics technology to create and render virtual reality environments.

[0806] A "means for rendering a 3D environment" is a process or device that depicts a three-dimensional virtual space in real time based on a generated scenario.

[0807] The "means for transmitting the generated 3D environment data to the user terminal" is a process or device that transfers the rendered 3D environment data to the terminal used by the user.

[0808] "Means for displaying a 3D environment on a user device and updating it in real time in response to the user's movements" refers to a process or device that displays a 3D virtual space on a user device and dynamically changes that virtual space in response to the user's operations and movements.

[0809] "Means for recognizing a user's emotions and dynamically modifying the virtual reality environment based on the recognized emotions" refers to a process or device that analyzes a user's facial expressions and voice to determine their emotions and adjusts the VR experience to match those emotions.

[0810] "Means for collecting user feedback and improving the system" refers to a process or device that collects opinions and ratings provided by users after their experience and uses that data to update the system's functionality and performance.

[0811] The present invention is a system that recognizes a user's emotions and dynamically modifies a virtual reality experience based on the emotions. The system can generate experience content and render a virtual reality environment in real time based on the user's requests.

[0812] Hardware and software used

[0813] Hardware:

[0814] PC or smartphone with a camera

[0815] Virtual reality headsets (e.g., general VR devices)

[0816] software:

[0817] OpenCV library for emotion recognition

[0818] Hugging Face Transformers for Scenario Generation

[0819] Custom VR engine for real-time rendering (e.g. Unreal Engine Server)

[0820] API for collecting internet information

[0821] System Program Overview

[0822] Based on the user's request, the server converts the request into text data. Requests made by voice input are converted into text data using a highly accurate voice recognition library. Based on the text data, the server collects related data from the Internet. This is done using a specific search API to collect the required information.

[0823] The collected data is passed to a generative AI model, which generates scenarios using Hugging Face's Transformers. The generated scenarios are then rendered into a 3D environment using a virtual reality engine (e.g., an Unreal Engine server).

[0824] The rendered 3D environment data is sent to the device and displayed on the user's virtual reality headset. The 3D environment is updated in real time according to the user's movements. This updating allows the user to move freely around the VR space.

[0825] The system uses a camera and microphone to capture the user's facial expressions and voice, and then analyzes their emotions in real time using an emotion recognition engine (OpenCV). Based on the analyzed emotional data, the virtual reality environment is dynamically modified. For example, if the user is excited, special effects are added in the environment, and if the user is relaxed, the atmosphere of the environment is made calmer.

[0826] Finally, after the user has finished the experience, the terminal displays a feedback form where the user can enter comments and suggestions about the experience. This feedback information is sent to the server and used to continuously improve the system.

[0827] Specific examples

[0828] A user requests a virtual shopping store by speaking, "I want to buy new shoes." This request is converted into text data, and the server collects related information. Based on the collected data, an AI model generates a shopping scenario, and a VR engine renders a 3D environment. When the user puts on a virtual reality headset and enters the virtual store, a camera captures the user's facial expressions, and emotional data is analyzed. If the user finds a pair of shoes and expresses joy, the system displays a pop-up in the storefront offering a special discount.

[0829] Prompt Sentence Examples

[0830] User request: "I want to buy new shoes from a virtual shopping store."

[0831] AI-generated prompt: "Generate a scenario in which a user searches for shoes in a virtual shopping store."

[0832] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0833] Step 1:

[0834] The user speaks the request

[0835] Specific operation: The user says to the terminal, "I want to buy new shoes at a virtual shopping store."

[0836] Input: Audio data

[0837] Data calculation: Converts voice data into text data using a voice recognition library.

[0838] Output: Text data (e.g., "I want to buy new shoes at a virtual shopping store")

[0839] Step 2:

[0840] Sends requests to the server and collects relevant data

[0841] Specific operation: The terminal sends the converted text data to the server as an HTTP request, and the server searches the Internet for information related to the request content.

[0842] Input: Text data (request content)

[0843] Data Computing: Collecting relevant information from multiple databases and sources on the Internet.

[0844] Output: Collected relevant data (e.g. product information, images, reviews)

[0845] Step 3:

[0846] Analyze the collected data and format it appropriately

[0847] Specific operation: The server analyzes the collected data and formats it appropriately to be passed to the generative AI model.

[0848] Input: Relevant data collected

[0849] Data operations: Data analysis and formatting (e.g., text analysis, data structuring)

[0850] Output: Formatted data

[0851] Step 4:

[0852] Generate a scenario

[0853] Specific operation: Based on the collected data, the server sends prompts to the generative AI model (Hugging Face Transformers) to generate a shopping scenario.

[0854] Input: Formatted data, prompt (e.g., "Generate a scenario in which a user searches for shoes in a virtual shopping store.")

[0855] Data calculation: scenario generation using generative AI models

[0856] Output: Generated scenario

[0857] Step 5:

[0858] Rendering a Virtual Reality Environment

[0859] How it works: The server sends the generated scenario to a VR engine (e.g., Unreal Engine server), which renders the 3D virtual environment in real time.

[0860] Input: Generated scenario

[0861] Data Computing: 3D Modeling and Rendering with a VR Engine

[0862] Output: 3D environment data

[0863] Step 6:

[0864] 3D environment data is sent to the user's device and displayed

[0865] Specific operation: The server sends the rendered 3D environment data to the user's device, which then displays the data on the VR headset.

[0866] Input: 3D environment data

[0867] Data Computing: Communication and Data Transfer

[0868] Output: 3D virtual environment displayed in a VR headset

[0869] Step 7:

[0870] Update the 3D environment in real time according to the user's movements

[0871] How it works: As the user moves around and looks at objects in the VR space, the environment is updated in real time. Sensors capture movements, and the server calculates and transmits the corresponding changes in the environment.

[0872] Input: User movement data (e.g., position change, gaze direction)

[0873] Data Computation: Analyzing motion data and re-rendering the environment

[0874] Output: Updated 3D environment data

[0875] Step 8:

[0876] Recognize user emotions and dynamically change the environment

[0877] How it works: The camera captures the user's facial expressions, and the emotion recognition engine (OpenCV) analyzes them to recognize their emotions. The server dynamically adjusts the virtual reality environment based on this emotional data.

[0878] Input: User's facial expression data, voice tone

[0879] Data Computing: Facial Expression Analysis and Emotion Recognition

[0880] Output: Change the environment based on the emotion (e.g., special discount popup, adding interactive elements)

[0881] Step 9:

[0882] Collect feedback after the experience and refine the system

[0883] Specific operation: After the user finishes the experience, the device displays a feedback form and sends the user's comments and requests for improvement to the server. The server analyzes this data and uses it to improve the system.

[0884] Input: User feedback data (e.g., satisfaction, areas for improvement)

[0885] Data Computing: Feedback Analysis and System Optimization

[0886] Output: Improved system functionality and performance

[0887] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0888] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0889] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0890] [Third embodiment]

[0891] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0892] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0893] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0894] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0895] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0896] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0897] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0898] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0899] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0900] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0901] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0902] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0903] This invention provides a system that allows users to experience specific historical moments in the past or future through virtual reality. Based on user requests, this system collects relevant data from the Internet, generates scenarios using generative artificial intelligence, and provides a mechanism for rendering 3D environments in real time using a virtual reality engine.

[0904] Program Overview

[0905] 1. User submits request:

[0906] The user speaks into the device to request an experience of a specific time in the past or future.

[0907] The device analyzes the voice and converts it into text data.

[0908] 2. Data Collection:

[0909] The terminal transmits the text data to the server.

[0910] The server collects relevant information from the Internet based on the request.

[0911] 3. Data analysis and scenario generation:

[0912] The server analyzes the collected data and generates scenarios using generative artificial intelligence.

[0913] A scenario details the buildings, people, and events associated with a specified time period and place.

[0914] 4. Real-time rendering:

[0915] The server uses the generated scenario to render a 3D environment in a virtual reality engine.

[0916] The 3D environment contains details and realism that match the user's requirements.

[0917] 5. Data transmission and display:

[0918] The server sends the rendered 3D environment data to the terminal.

[0919] The device displays the received data on a virtual reality headset.

[0920] 6. Immersive user experience:

[0921] The user wears a virtual reality headset and is immersed in a virtual environment.

[0922] The environment is updated in real time according to the user's movements and gaze.

[0923] 7. Feedback collection and system improvement:

[0924] After the user has completed the experience, the terminal displays a feedback form in which the user can enter comments about the experience and requests for improvements.

[0925] The device sends feedback to the server, which analyzes it and uses it to improve the system.

[0926] Specific examples

[0927] Case: Renaissance Florence

[0928] 1. Submit your request:

[0929] The user speaks to the device and says, "I want to experience Renaissance Florence."

[0930] The terminal converts this voice into text data and sends it to the server.

[0931] 2. Data Collection:

[0932] The server collects information related to "Renaissance Florence" from the Internet.

[0933] For example, data such as text, images, audio, and video related to buildings, people, and events from that time is obtained.

[0934] 3. Data analysis and scenario generation:

[0935] The server analyzes the collected data and uses generative artificial intelligence to generate a scenario of "Florence during the Renaissance."

[0936] Scenarios include Da Vinci painting and the construction site of Florence Cathedral.

[0937] 4. Real-time rendering:

[0938] Based on the generated scenario, the server uses a virtual reality engine to render a 3D environment of Florence in real time.

[0939] 5. Data transmission and display:

[0940] The server sends the rendered 3D environment data to the terminal.

[0941] The device receives the data and displays it on a virtual reality headset.

[0942] 6. Immersive Experience:

[0943] Users put on a headset and are immersed in a realistic Renaissance Florence.

[0944] Users can interact with Da Vinci and tour the inside of the cathedral.

[0945] 7. Feedback and System Improvement:

[0946] After completing the experience, users fill out a feedback form to indicate their satisfaction with the experience and suggestions for improvement.

[0947] The device sends this to the server, which analyzes the feedback and uses it to improve the system.

[0948] This allows users to experience historical moments and future worlds in a highly immersive way.

[0949] The processing flow will be explained below.

[0950] Step 1:

[0951] The user makes a request to experience a specific time in the past or future by voice into the device, which then recognizes the voice and converts the voice data into text data.

[0952] Step 2:

[0953] The device sends the converted text data to the server as an HTTP request, which includes the user's desired experience.

[0954] Step 3:

[0955] Based on the received request, the server launches a scraping tool to scrape relevant text, image, audio, and video data from the Internet.

[0956] Step 4:

[0957] The server uses a scraping tool to collect data from the Internet that matches the specified keywords, and the collected data is temporarily stored.

[0958] Step 5:

[0959] The server analyzes the collected data using natural language processing and image recognition technology and formats it appropriately to be passed to the generative artificial intelligence.

[0960] Step 6:

[0961] The server inputs the formatted data into a generative AI to generate a scenario for the specified time period and location, including detailed historical background and key buildings, people, and events.

[0962] Step 7:

[0963] The server uses an evaluation algorithm to check the quality of the generated scenario, correcting or regenerating it as necessary. Once the evaluation is complete, the scenario moves on to the next stage.

[0964] Step 8:

[0965] The server uses a virtual reality engine to render a 3D environment in real time based on the scenario, using the GPU to achieve high performance and detailed rendering.

[0966] Step 9:

[0967] The server compresses the rendered 3D environment data and transmits it to the device in an efficient binary format, using streaming techniques to minimize data latency.

[0968] Step 10:

[0969] The device then extracts the received 3D environment data and displays it on the virtual reality headset, which updates the displayed 3D environment in real time according to the user's movements and gaze.

[0970] Step 11:

[0971] Users wear a virtual reality headset and enter an immersive virtual environment, where they can freely move around and experience interactive elements (e.g., interact with people, manipulate objects).

[0972] Step 12:

[0973] When the user finishes their experience, the device displays a feedback form where the user can enter comments about their experience and suggestions for improvement. Collecting feedback contributes to improving user satisfaction.

[0974] Step 13:

[0975] The device sends the collected feedback data to the server, which analyzes the feedback and adjusts the algorithms of the generative AI and virtual reality engine to help improve the system, resulting in an even better user experience next time.

[0976] Example 1

[0977] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0978] Conventional virtual reality systems face challenges when it comes to quickly and accurately generating detailed scenarios and rendering 3D environments in real time when users want to experience specific historical moments or future worlds. Other challenges include updating the 3D environment in real time in response to user movements and systematically collecting user feedback to improve the system.

[0979] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0980] In this invention, the server includes means for converting user requests into text data based on the user requests, means for collecting related information from a communication network, means for generating a scenario using artificial intelligence based on the collected information, means for creating a 3D environment using virtual reality technology based on the generated scenario, means for transmitting the generated 3D environment data to a user device, means for displaying the 3D environment on the user device and updating it in real time according to the user's movements, and means for collecting user responses and improving the system, thereby enabling users to experience specific historical moments or future worlds in detail and realistically.

[0981] "User" refers to a person who uses the system to experience virtual reality.

[0982] A "request" refers to a request for a specific historical moment or future world that a user would like to experience.

[0983] "Text data" refers to data in which a user's request is converted into text form through voice recognition or other input methods.

[0984] "Communications network" means a wide-area computer network, including the Internet and other data communications infrastructures.

[0985] "Information" refers to data such as text, images, and videos collected based on user requests.

[0986] "Artificial intelligence" refers to machine learning algorithms and neural networks used to analyze collected information and generate specified scenarios.

[0987] A "scenario" refers to a set of information that contains a detailed description of a particular historical moment or future world that a user wants to experience.

[0988] "Virtual reality technology" refers to technology that uses computer technology to generate a 3D environment and provide an immersive experience to the user.

[0989] "3D environment" refers to a three-dimensional virtual space created using virtual reality technology based on a scenario.

[0990] "User device" refers to the terminal or headset that a user uses to have a virtual reality experience.

[0991] "Real-time updates" means that the 3D environment is instantly reflected and updated in response to inputs such as the user's movements and gaze.

[0992] "Reactions" refer to the feedback and opinions users provide through their virtual reality experience.

[0993] "System" refers to the comprehensive technological platform that includes all of the above means.

[0994] The present invention is a system that allows users to experience specific historical moments in the past or future through virtual reality. Based on a user's request, the system collects relevant information from the internet, generates a scenario using a generative AI model, and uses virtual reality technology to render a 3D environment in real time.

[0995] First, a user requests a historical moment or future world they would like to experience through a voice request. For example, a user might say, "I want to experience Tokyo of the future." This voice request is converted into text data by the device's voice recognition software (e.g., Google Speech-to-Text API).

[0996] The device sends the converted text data to a server, which then collects related information from the Internet via a communications network. Specifically, it uses web scraping tools (e.g., BeautifulSoup) or APIs (e.g., Wikipedia API) to obtain the necessary text, images, videos, etc.

[0997] The collected information is analyzed by the server and passed to a generative AI model (e.g., OpenAI's GPT-4). This generative AI model creates a detailed scenario based on the collected information. For example, it uses a prompt such as, "Generate a detailed scenario about Tokyo in the future. Include skyscrapers, a modern transportation system, and people's activities."

[0998] Based on the generated scenario, the server uses virtual reality technology to render a 3D environment in real time using a virtual reality engine (e.g., Unreal Engine or Unity). Specifically, the future Tokyo cityscape, transportation system, building interiors, streetscapes, and other details are depicted.

[0999] The generated 3D environment data is sent to the terminal and displayed on the user device (e.g., a virtual reality headset). The user wears the headset and is immersed in the virtual environment. The user's movements and gaze are detected by tracking sensors (e.g., the tracking system in the HTC Vive), and the 3D environment is updated in real time. This allows the user to freely explore the virtual reality and enjoy a detailed experience.

[1000] Finally, once the user has finished the experience, the device displays a feedback form. The user can enter comments about the experience or suggestions for improvement, which the device then sends to the server. The server analyzes the collected feedback and incorporates it into the next update or improvement. This feedback is then used with natural language processing tools and machine learning algorithms to continuously improve the system.

[1001] As described above, the present invention provides a system that provides a user with a high level of immersion and allows them to experience specific historical moments in the past or future in detail and realistically.

[1002] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1003] Step 1: User submits request

[1004] Users speak their requests into the device, for example, saying, "I want to experience Renaissance Florence."

[1005] The device receives this voice input and converts it into text using voice recognition software, specifically the Google Speech-to-Text API.

[1006] Input: User's voice request

[1007] Output: The request converted to character data

[1008] Step 2: Data collection

[1009] The terminal transmits the converted character data to the server.

[1010] Based on the received request, the server collects relevant information from the Internet via a communication network, for example, using a web scraping tool (BeautifulSoup) or an API (Wikipedia API).

[1011] Input: Request converted to character data

[1012] Output: Collected information (text, images, videos, etc.)

[1013] Step 3: Data analysis and scenario generation

[1014] The server analyzes the collected information and organizes it using a natural language processing tool (spaCy).

[1015] The server inputs a prompt into a generative AI model (e.g., OpenAI's GPT-4) to generate a detailed scenario. An example of the prompt is, "Generate a scenario set in Renaissance Florence, including a scene where Da Vinci is painting."

[1016] Input: Collected information

[1017] Output: Generated scenario

[1018] Step 4: Real-time rendering

[1019] Based on the generated scenario, the server uses virtual reality technology to render a 3D environment in real time, specifically using a virtual reality engine (Unreal Engine or Unity).

[1020] Input: Generated scenario

[1021] Output: Rendered 3D environment

[1022] Step 5: Send and display data

[1023] The server sends the rendered 3D environment data to the device, using WebSocket or HTTP / 2 as the communication protocol.

[1024] The device prepares the received data for display on a virtual reality headset.

[1025] Input: Rendered 3D environment data

[1026] Output: Display to a virtual reality headset

[1027] Step 6: User Immersion

[1028] The user wears a virtual reality headset and is immersed in a virtual environment.

[1029] The device uses sensors to track the user's movements and gaze, updating the 3D environment in real time, for example using the tracking system in the HTC Vive.

[1030] Input: User movement and gaze data

[1031] Output: Real-time updated 3D environment

[1032] Step 7: Gather feedback and improve the system

[1033] After the user has finished the experience, the terminal displays a feedback form.

[1034] Users enter comments about their experience and requests for improvement, and the device sends these to the server.

[1035] The server analyzes the collected feedback and uses natural language processing tools and machine learning algorithms to help improve the system.

[1036] Input: User feedback

[1037] Output: Analyzed feedback and system improvements

[1038] The above is the specific processing flow of this system.

[1039] (Application example 1)

[1040] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1041] Current virtual reality systems lack efficient means for users to experience specific historical moments or future events with a high level of realism. Furthermore, technologies for automatically generating scenarios in real time based on collected data and then generating and updating virtual reality environments in real time based on those scenarios remain a challenge. Therefore, systems that improve the quality of the user experience are needed.

[1042] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1043] In this invention, the server includes means for converting a user's request into text data based on the request, means for collecting relevant data from the Internet, means for generating a scenario using a generating artificial intelligence based on the collected data, means for rendering a three-dimensional environment using a virtual reality engine based on the generated scenario, means for transmitting the generated three-dimensional environment data to a user terminal, means for displaying the three-dimensional environment on the user terminal and updating it in real time according to the user's movements, means for collecting user feedback and improving the system, means for providing the generated three-dimensional environment to the user using a visual display device, means for analyzing the collected data and formatting it into an appropriate format for passing it to the generating artificial intelligence, means for ensuring the quality of the scenario using an algorithm for evaluating the quality of the generated scenario, means for rendering the generated three-dimensional environment in real time using a virtual reality engine and transmitting it to a visual display device, and means for displaying the generated three-dimensional environment in real time on the user's mobile device or head-mounted display and updating the environment in accordance with the user's line of sight and movements, thereby enabling users to experience historical moments and future events in real time with a high level of realism.

[1044] A "user request" refers to a user's desire to experience a particular historical moment or future event through audio or text.

[1045] "Text data" refers to data obtained by converting a user's request from voice into text information.

[1046] "Related Data" refers to information about a specified time and place that is collected from the Internet based on a user's request.

[1047] "Generative AI" refers to artificial intelligence technology that generates scenarios based on collected data.

[1048] A "scenario" is a plan or script that details the events and settings to be experienced within a virtual reality environment based on generated data.

[1049] "Virtual reality engine" refers to a software and hardware configuration for rendering three-dimensional environments in real time.

[1050] "Three-dimensional environment" refers to the virtual space that a user experiences through virtual reality.

[1051] "User terminal" refers to a device including a smartphone, a head-mounted display, or other display device.

[1052] "Visual display device" refers to a display device used by a user to visually experience a three-dimensional environment.

[1053] "Feedback" refers to comments, ratings, and requests for improvement regarding the experience provided by the user after the experience.

[1054] "Formatting" refers to the process of converting collected data into an appropriate form for use by the generative AI.

[1055] "Algorithm" refers to a computational procedure for assessing and ensuring the quality of generated scenarios.

[1056] The present invention provides a system that allows a user to experience a specific historical moment or a future event through virtual reality. The system comprises the following means:

[1057] First, the user inputs a request by voice or text using a device such as a smartphone or head-mounted display. For example, the user's request might be something like, "I want to experience a Renaissance city." This request is converted into text data using speech recognition technology (for example, Python's speech_recognition library).

[1058] The server then retrieves the converted text data and collects relevant data from the Internet. During this process, the required data is retrieved from the source using, for example, the requests library. The collected data is then analyzed using generative artificial intelligence (e.g., OpenAI's API) and formatted accordingly.

[1059] The server generates a scenario based on the formatted data. The scenario generated using the AI ​​generator includes detailed events and settings related to the specified time period and location. When generating the scenario, a prompt message corresponding to the user's request is passed to the AI ​​generator. For example, the following prompt message is used:

[1060] Example prompt:

[1061] A user has requested to experience a Renaissance city. Generate a scenario based on the following relevant data:

[1062] Renaissance buildings and streets

[1063] Costumes and culture of the time

[1064] Historical events and everyday scenes

[1065] Generate detailed scenarios and make them suitable for virtual reality experiences.

[1066] The generated scenario is rendered in real time as a 3D environment using a virtual reality engine (e.g., a 3D engine such as py3dengine). The rendered 3D environment data is sent from the server to the user's device, and the user is immersed in the virtual environment through a visual display device (e.g., a head-mounted display or a smartphone).

[1067] After the user has completed their experience in the virtual environment, the server collects feedback from the user by entering comments and ratings into a dedicated form, which is used to improve the system.

[1068] For example, if a user requests "I want to experience a Renaissance city," the server collects relevant data and uses generative artificial intelligence to generate a specific Renaissance scenario, including detailed descriptions of buildings, cityscapes, and costumed people in action. The generated scenario is then rendered as a three-dimensional environment by a virtual reality engine, allowing the user to use a visual display device to tour the city and interact with historical figures.

[1069] As described above, the present embodiment provides a series of processes for enhancing a user's immersive experience.

[1070] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1071] Step 1:

[1072] The user speaks a request into the device. This request describes a desire to experience a specific historical moment or future event. The device recognizes this speech as input data and converts it into text data using speech recognition technology (for example, Python's speech_recognition library). The output is the user's request converted into text data as character information.

[1073] Step 2:

[1074] The terminal sends the converted text data to the server. The server receives the text data as input and collects related data from the Internet. Here, it uses the requests library to obtain the necessary data from the target sources. The output is the collected related data.

[1075] Step 3:

[1076] The server takes the collected data as input, analyzes it using generative AI (for example, OpenAI's API), and formats it into an appropriate format. During this process, the data is structured and necessary information is filtered. The output is data formatted in an appropriate format for the generative AI to generate scenarios.

[1077] Step 4:

[1078] The server uses the formatted data to send prompts to the generation AI to generate a scenario. The prompt text is a specific scenario request based on the user's request and the collected data. The generation AI generates a scenario based on the prompt and provides the scenario in text format as output.

[1079] Step 5:

[1080] The server takes the generated scenario as input and uses a virtual reality engine (e.g., py3dengine) to render a 3D environment in real time. During this process, a 3D environment is generated according to the generated scenario. The output is 3D data rendered as a virtual reality environment.

[1081] Step 6:

[1082] The server transmits the rendered 3D environment data to the user's device, which receives this data as input and displays the virtual environment to the user through a visual display device (such as a head-mounted display or a smartphone). The user is immersed in this virtual environment and experiences it.

[1083] Step 7:

[1084] After the user has completed their experience in the virtual environment, the device displays a feedback form. The user enters comments, ratings, and suggestions for improvement about the experience. The device sends this as input data to the server. The server analyzes the feedback and uses it to improve the system. The output is improvements based on the feedback.

[1085] This series of processing steps allows users to experience historical moments and future events in real time with a high level of realism.

[1086] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1087] This invention is a system that recognizes a user's emotions in a virtual reality environment and dynamically changes the experience content based on those emotions. Based on the user's request, the system converts the request into text data, collects related data from the Internet, and generates a scenario using generative artificial intelligence. Based on the generated scenario, a 3D environment is rendered using a virtual reality engine and displayed on the user's device. Furthermore, by updating the environment in real time based on the user's movements and combining it with an emotion engine that recognizes the user's emotions, an experience tailored to the user's emotions is provided.

[1088] Program Overview

[1089] 1. User submits request:

[1090] The user requests an experience of a specific time in the past or future by speaking into the device, which then analyzes the voice and converts the voice data into text data.

[1091] 2. Data Collection:

[1092] The device sends text data to the server as an HTTP request, and the server collects related information from the Internet based on the request content.

[1093] 3. Data analysis and scenario generation:

[1094] The server analyzes the collected data and uses generative artificial intelligence to generate scenarios that include detailed descriptions of historical background, key buildings, people, and events.

[1095] 4. Real-time rendering:

[1096] The server uses the generated scenario to render a 3D environment in real time in a virtual reality engine, with details and realism that match the user's requirements.

[1097] 5. Data transmission and display:

[1098] The server sends the rendered 3D environment data to the device, which displays it on a virtual reality headset.

[1099] 6. Immersive user experience:

[1100] Users wear a virtual reality headset and are immersed in a virtual environment that updates in real time based on the user's movements and gaze.

[1101] 7. Emotion Recognition with Emotion Engine:

[1102] The server uses an emotion engine that analyzes the user's facial expressions and vocal tone to recognize the user's emotions in real time. The recognized emotion data is analyzed and reflected in the content of the virtual environment.

[1103] 8. Dynamic Environment Changes:

[1104] The server dynamically modifies the virtual reality environment based on the emotional data recognized by the emotion engine, for example by increasing the effects in the environment if the user is surprised, or by adding interactive elements if the user is engrossed.

[1105] 9. Feedback Collection and System Improvement:

[1106] After the user has finished the experience, the device displays a feedback form, where the user can enter comments about the experience and suggestions for improvement. The device then sends this feedback to the server, which then analyzes the feedback and emotional data to help improve the system.

[1107] Specific examples

[1108] Case: Renaissance Florence

[1109] 1. Submit your request:

[1110] The user speaks to the device, saying, "I want to experience Renaissance Florence." The device converts this speech into text data and sends it to the server.

[1111] 2. Data Collection:

[1112] The server collects information related to "Renaissance Florence" from the Internet, including text, images, audio, and video data related to buildings, people, and events from that time.

[1113] 3. Data analysis and scenario generation:

[1114] The server analyzes the collected data and uses generative artificial intelligence to generate a "Renaissance Florence" scenario, including scenes of Da Vinci painting and the construction site of Florence Cathedral.

[1115] 4. Real-time rendering:

[1116] Based on the generated scenario, the server uses a virtual reality engine to render a 3D environment of Florence in real time.

[1117] 5. Data transmission and display:

[1118] The server sends the rendered 3D environment data to the device, which receives the data and displays it on the virtual reality headset.

[1119] 6. Immersive Experience:

[1120] Users put on a headset and are immersed in a realistic Renaissance Florence, where they can interact with Da Vinci and tour the interior of the cathedral.

[1121] 7. Emotion Recognition with Emotion Engine:

[1122] The server analyzes the user's facial expressions and voice using an emotion engine and recognizes that the user is excited.

[1123] 8. Dynamic Environment Changes:

[1124] Based on the user's emotional data, the server adds special events (e.g., fireworks and musical performances) to the streets of Florence to make the experience even more engaging.

[1125] 9. Feedback and System Improvement:

[1126] After completing the experience, the user fills in a feedback form to indicate their satisfaction with the experience and any suggestions for improvement. The device then sends this information to the server, which then analyzes the feedback and emotional data and reflects it in improvements to the system.

[1127] In this way, the system of the present invention recognizes the user's emotions and dynamically modifies the virtual reality experience based on them, providing a richer and more engaging experience.

[1128] The processing flow will be explained below.

[1129] Step 1:

[1130] The user speaks into the device, saying, "I want to experience Renaissance Florence." The device uses voice recognition software to convert the voice data into text.

[1131] Step 2:

[1132] The device sends the converted text data to the server as an HTTP request, which includes the user's desired experience.

[1133] Step 3:

[1134] Based on the received request, the server uses a scraping tool to collect related information (text, image, audio, and video data) that matches the specified keywords from the Internet.

[1135] Step 4:

[1136] The server analyzes the scraped data using natural language processing and image recognition technology and formats it appropriately to be passed on to generative artificial intelligence (AI).

[1137] Step 5:

[1138] The server inputs the formatted data into a generative artificial intelligence (AI) that generates a scenario about "Renaissance Florence," including detailed descriptions of buildings, people, and events.

[1139] Step 6:

[1140] The server uses an evaluation algorithm to verify the quality of the generated scenarios, and corrects or regenerates them as necessary.

[1141] Step 7:

[1142] The server uses quality-tested scenarios to leverage a virtual reality engine to render 3D environments in real time.

[1143] Step 8:

[1144] The server compresses the rendered 3D environment data in binary format and transmits it to the terminal using efficient data transfer techniques.

[1145] Step 9:

[1146] The device receives 3D environment data from the server and displays it on the virtual reality headset. The device tracks the user's gaze and movements in real time and updates the environment.

[1147] Step 10:

[1148] Users put on a virtual reality headset and begin experiencing a realistic Renaissance Florence. They can freely explore the virtual environment and interact with the characters.

[1149] Step 11:

[1150] The server uses an emotion engine to collect the user's facial expressions and voice tone in real time from the camera and microphone installed in the headset and analyze the user's emotions.

[1151] Step 12:

[1152] The server dynamically modifies the virtual reality environment based on the analyzed emotional data, for example adding effects and interactive elements if the user is excited.

[1153] Step 13:

[1154] After the user has finished the experience, the terminal displays a feedback form, where the user can enter comments about the experience and requests for improvements. The terminal then sends the feedback to the server.

[1155] Step 14:

[1156] The server analyzes the feedback and emotional data, adjusts the algorithms of the generative AI and virtual reality engine, and reflects the results in system improvements, resulting in an even better user experience next time.

[1157] Example 2

[1158] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1159] Conventional virtual reality systems lack the flexibility to adapt to user requests and the ability to dynamically change the environment based on the user's emotions. These shortcomings limit the quality of the user experience and make it difficult to provide a richer and more engaging virtual reality experience. Therefore, there is a need for a system that dynamically changes the generation of scenarios and rendering of virtual reality environments based on user requests, depending on the user's emotions.

[1160] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1161] In this invention, the server includes: means for converting a user's request into text data based on the request; means for collecting related data from the Internet; means for generating a scenario using a generative model based on the collected data; means for rendering a three-dimensional environment using a virtual reality engine based on the generated scenario; means for transmitting the generated three-dimensional environment data to a user terminal; means for displaying the three-dimensional environment on the user terminal and updating it in real time according to the user's movements; means including an emotion analysis engine for recognizing the user's emotions; means for dynamically changing the virtual reality environment based on the recognized emotion data; and means for collecting user feedback and improving the system. This makes it possible to flexibly generate scenarios according to user requests and dynamically change the virtual reality environment according to the user's emotions.

[1162] "Means for converting a request into text data" refers to a function for analyzing a request input by a user in a voice or other format and converting it into text.

[1163] "Means of collecting relevant data from the Internet" refers to the function of obtaining necessary information from websites and databases via a network.

[1164] A "generative model" refers to an artificial intelligence algorithm that generates appropriate results for a given input (prompt).

[1165] "Means for generating a scenario" refers to the function of creating stories and events for a virtual reality experience that responds to a user's request based on collected data.

[1166] "Virtual reality engine" refers to a software platform for generating and rendering three-dimensional virtual environments in real time.

[1167] "Means for rendering a three-dimensional environment" refers to a function for depicting a three-dimensional visual environment based on a generated scenario.

[1168] "Means for transmitting three-dimensional environmental data to a user terminal" refers to a function for transferring three-dimensional visual data generated by the server to the terminal used by the user.

[1169] "Means for displaying a three-dimensional environment and updating it in real time in response to the user's movements" refers to a function that dynamically changes the three-dimensional virtual environment displayed on the user's terminal in synchronization with the user's movements and operations.

[1170] An "emotion analysis engine" refers to software that analyzes a user's facial expressions and vocal tone to identify emotions.

[1171] "Means for dynamically modifying the virtual reality environment based on recognized emotional data" refers to the ability to adjust the content and events of the virtual reality in real time based on data obtained through emotional analysis.

[1172] "Means for collecting feedback and improving the system" refers to a function for incorporating user evaluations and opinions of the experience and improving the experience in future.

[1173] The present invention is a system that generates a scenario based on a user's request in a virtual reality environment and provides the user with a three-dimensional environment corresponding to that scenario. This invention makes it possible to recognize the user's emotions in real time and dynamically change the virtual reality environment based on those emotions.

[1174] First, the user speaks into the device to request an experience of a specific time or place. The device then uses voice recognition software, such as the Google Speech-to-Text API, to convert this voice data into text. This text data then becomes the basis for interpreting the request.

[1175] The converted text data is then sent to the server as an HTTP request, which uses web crawling tools and APIs to gather relevant information from the internet, such as Beautiful Soup and Scrapy, to gather text, images, and audio data.

[1176] The collected data is then analyzed using natural language processing tools, such as spaCy and NLTK. Based on the results of this analysis, a generative AI model (e.g., GPT-4, ChatGPT) generates a scenario. Input prompts are crucial for scenario generation. For example, a prompt such as "Walk through the streets of Renaissance Florence" generates detailed historical background and character dialogue.

[1177] Based on the generated scenario, the server uses a virtual reality engine (e.g., Unity, Unreal Engine) to render a three-dimensional virtual environment. The rendered data is sent to the user's device via HTTP. The user's device receives this data and displays it on a virtual reality headset (e.g., Oculus Rift, HTC Vive). The user can then immerse themselves in and interact with this virtual environment.

[1178] During the user's experience, the device uses a camera and microphone to capture the user's facial expressions and voice. This data is analyzed in real time by an emotion analysis engine (e.g., Affectiva, Kairos). Based on the user's emotion data, the server dynamically modifies the virtual environment. For example, if the user is surprised, effects such as fireworks are added to the virtual environment.

[1179] After the experience is over, the device displays a feedback form to the user, allowing them to input their satisfaction with the experience and suggestions for improvement. This data is then sent to the server, which analyzes the feedback and emotional data to improve the experience for the next session.

[1180] A concrete example is a user request to "experience Renaissance Florence." In this case, the user speaks the request into the device, which converts this speech into text data and sends it to the server. The server collects relevant information from the Internet and generates a scenario using a generative AI model. It then renders a three-dimensional environment of Florence using a virtual reality engine and sends it to the user's device. The user puts on the virtual reality headset and explores the city of Florence. If the user is excited, the server adds special events to make the experience even more engaging.

[1181] This invention makes it possible to flexibly generate scenarios based on user requests and dynamically change the virtual reality environment in response to the user's emotions.

[1182] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1183] Step 1:

[1184] User request submission

[1185] Input: The user speaks into the device, "I want to experience Renaissance Florence."

[1186] How it works: The device uses speech recognition software (e.g., Google Speech-to-Text API) to convert speech into text. Voice input is text, and output is text.

[1187] Output: The converted text data. The text is "I want to experience Renaissance Florence."

[1188] Step 2:

[1189] Data collection

[1190] Input: The request as text data.

[1191] How it works: The device sends this text data to the server as an HTTP request. The server then uses a web crawling tool (e.g., Beautiful Soup, Scrapy) to collect relevant information from the Internet. This includes data such as text, images, and audio.

[1192] Output: A dataset containing relevant information.

[1193] Step 3:

[1194] Data analysis and scenario generation

[1195] Input: A dataset containing the relevant collected information.

[1196] How it works: The server uses natural language processing tools (e.g., spaCy, NLTK) to parse the data and format it into a suitable format for a generative AI model (e.g., GPT-4, ChatGPT). It then generates a scenario based on the generative AI model. For example, you can input the prompt "Florence during the Renaissance," and the scenario will be output.

[1197] Output: Generated scenario data.

[1198] Step 4:

[1199] Real-time rendering

[1200] Input: Generated scenario data.

[1201] How it works: The server uses a virtual reality engine (e.g., Unity, Unreal Engine) to render a 3D environment in real time. Based on the scenario, models of objects and characters in the environment are generated and rendered in real time.

[1202] Output: Rendered 3D environment data.

[1203] Step 5:

[1204] Sending and Displaying Data

[1205] Input: Rendered 3D environment data.

[1206] How it works: The server sends this data to the user's device, which then displays it on a virtual reality headset (e.g., Oculus Rift, HTC Vive).

[1207] Output: A three-dimensional environment displayed in a virtual reality headset.

[1208] Step 6:

[1209] Immersive User Experience

[1210] Input: A three-dimensional environment displayed on a virtual reality headset.

[1211] How it works: The user puts on a virtual reality headset and is immersed in a virtual environment. The environment updates in real time based on the user's movements and gaze. For example, the user can explore and interact with Renaissance Florence.

[1212] Output: Updated virtual environment and user experience data.

[1213] Step 7:

[1214] Emotion recognition by emotion engine

[1215] Input: User's facial expressions and voice.

[1216] How it works: The device captures the user's facial expressions and voice using a camera and microphone. The server analyzes this data in real time using an emotion analysis engine (e.g., Affectiva, Kairos) to recognize the user's emotions.

[1217] Output: Recognized emotion data.

[1218] Step 8:

[1219] Dynamic Environment Changes

[1220] Input: Recognized emotion data.

[1221] Action: The server dynamically changes the virtual reality environment based on the emotional data, for example adding special events or effects if the user is excited.

[1222] Output: A dynamically modified virtual reality environment.

[1223] Step 9:

[1224] Feedback collection and system improvement

[1225] Input: User feedback after the experience.

[1226] How it works: After the experience is over, the device displays a feedback form to the user. The user enters their satisfaction with the experience and suggestions for improvement, and this data is sent to the server. The server analyzes the feedback and emotion data to help improve the system.

[1227] Output: An improved system and updated user experience data.

[1228] (Application example 2)

[1229] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1230] Conventional virtual reality systems provide a fixed experience without taking the user's emotions into consideration. This means that the experience cannot be dynamically changed based on the user's emotions and reactions, resulting in a lack of immersion and personalization. Furthermore, it has not been possible for the system to recognize the user's emotions during the experience and provide an optimized experience based on those emotions. This has led to issues such as low user satisfaction and engagement.

[1231] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1232] In this invention, the server includes means for converting a user's request into text data based on the user's request, means for collecting related data from the Internet, means for generating a scenario based on the collected data using artificial intelligence, means for rendering a 3D environment based on the generated scenario using a virtual reality engine, means for transmitting the generated 3D environment data to a user terminal, means for displaying the 3D environment on the user terminal and updating it in real time according to the user's movements, means for recognizing the user's emotions and dynamically changing the virtual reality environment based on the recognized emotions, and means for collecting user feedback and improving the system. This enables dynamic changes to the virtual reality experience based on the user's emotions, providing a more immersive and personalized experience.

[1233] A "user request" is a request to the system for content or information that the user wants to experience.

[1234] A "means for converting to text data" is a process or device that analyzes and converts voice or other non-text data into written information.

[1235] "Means for collecting relevant data from the Internet" refers to a process or device that searches for and obtains specified information or data via the Internet.

[1236] "Generative AI" is an AI technology that generates scenarios and content for specific purposes based on collected data.

[1237] A "scenario generation means" is a process or device that analyzes collected data and arranges it into a specific story or scenario format.

[1238] A "virtual reality engine" is software or hardware that uses computer graphics technology to create and render virtual reality environments.

[1239] A "means for rendering a 3D environment" is a process or device that depicts a three-dimensional virtual space in real time based on a generated scenario.

[1240] The "means for transmitting the generated 3D environment data to the user terminal" is a process or device that transfers the rendered 3D environment data to the terminal used by the user.

[1241] "Means for displaying a 3D environment on a user device and updating it in real time in response to the user's movements" refers to a process or device that displays a 3D virtual space on a user device and dynamically changes that virtual space in response to the user's operations and movements.

[1242] "Means for recognizing a user's emotions and dynamically modifying the virtual reality environment based on the recognized emotions" refers to a process or device that analyzes a user's facial expressions and voice to determine their emotions and adjusts the VR experience to match those emotions.

[1243] "Means for collecting user feedback and improving the system" refers to a process or device that collects opinions and ratings provided by users after their experience and uses that data to update the system's functionality and performance.

[1244] The present invention is a system that recognizes a user's emotions and dynamically modifies a virtual reality experience based on the emotions. The system can generate experience content and render a virtual reality environment in real time based on the user's requests.

[1245] Hardware and software used

[1246] Hardware:

[1247] PC or smartphone with a camera

[1248] Virtual reality headsets (e.g., general VR devices)

[1249] software:

[1250] OpenCV library for emotion recognition

[1251] Hugging Face Transformers for Scenario Generation

[1252] Custom VR engine for real-time rendering (e.g. Unreal Engine Server)

[1253] API for collecting internet information

[1254] System Program Overview

[1255] Based on the user's request, the server converts the request into text data. Requests made by voice input are converted into text data using a highly accurate voice recognition library. Based on the text data, the server collects related data from the Internet. This is done using a specific search API to collect the required information.

[1256] The collected data is passed to a generative AI model, which generates scenarios using Hugging Face's Transformers. The generated scenarios are then rendered into a 3D environment using a virtual reality engine (e.g., an Unreal Engine server).

[1257] The rendered 3D environment data is sent to the device and displayed on the user's virtual reality headset. The 3D environment is updated in real time according to the user's movements. This updating allows the user to move freely around the VR space.

[1258] The system uses a camera and microphone to capture the user's facial expressions and voice, and then analyzes their emotions in real time using an emotion recognition engine (OpenCV). Based on the analyzed emotional data, the virtual reality environment is dynamically modified. For example, if the user is excited, special effects are added in the environment, and if the user is relaxed, the atmosphere of the environment is made calmer.

[1259] Finally, after the user has finished the experience, the terminal displays a feedback form where the user can enter comments and suggestions about the experience. This feedback information is sent to the server and used to continuously improve the system.

[1260] Specific examples

[1261] A user requests a virtual shopping store by speaking, "I want to buy new shoes." This request is converted into text data, and the server collects related information. Based on the collected data, an AI model generates a shopping scenario, and a VR engine renders a 3D environment. When the user puts on a virtual reality headset and enters the virtual store, a camera captures the user's facial expressions, and emotional data is analyzed. If the user finds a pair of shoes and expresses joy, the system displays a pop-up in the storefront offering a special discount.

[1262] Prompt Sentence Examples

[1263] User request: "I want to buy new shoes from a virtual shopping store."

[1264] AI-generated prompt: "Generate a scenario in which a user searches for shoes in a virtual shopping store."

[1265] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1266] Step 1:

[1267] The user speaks the request

[1268] Specific operation: The user says to the terminal, "I want to buy new shoes at a virtual shopping store."

[1269] Input: Audio data

[1270] Data calculation: Converts voice data into text data using a voice recognition library.

[1271] Output: Text data (e.g., "I want to buy new shoes at a virtual shopping store")

[1272] Step 2:

[1273] Sends requests to the server and collects relevant data

[1274] Specific operation: The terminal sends the converted text data to the server as an HTTP request, and the server searches the Internet for information related to the request content.

[1275] Input: Text data (request content)

[1276] Data Computing: Collecting relevant information from multiple databases and sources on the Internet.

[1277] Output: Collected relevant data (e.g. product information, images, reviews)

[1278] Step 3:

[1279] Analyze the collected data and format it appropriately

[1280] Specific operation: The server analyzes the collected data and formats it appropriately to be passed to the generative AI model.

[1281] Input: Relevant data collected

[1282] Data operations: Data analysis and formatting (e.g., text analysis, data structuring)

[1283] Output: Formatted data

[1284] Step 4:

[1285] Generate a scenario

[1286] Specific operation: Based on the collected data, the server sends prompts to the generative AI model (Hugging Face Transformers) to generate a shopping scenario.

[1287] Input: Formatted data, prompt (e.g., "Generate a scenario in which a user searches for shoes in a virtual shopping store.")

[1288] Data calculation: scenario generation using generative AI models

[1289] Output: Generated scenario

[1290] Step 5:

[1291] Rendering a Virtual Reality Environment

[1292] How it works: The server sends the generated scenario to a VR engine (e.g., Unreal Engine server), which renders the 3D virtual environment in real time.

[1293] Input: Generated scenario

[1294] Data Computing: 3D Modeling and Rendering with a VR Engine

[1295] Output: 3D environment data

[1296] Step 6:

[1297] 3D environment data is sent to the user's device and displayed

[1298] Specific operation: The server sends the rendered 3D environment data to the user's device, which then displays the data on the VR headset.

[1299] Input: 3D environment data

[1300] Data Computing: Communication and Data Transfer

[1301] Output: 3D virtual environment displayed in a VR headset

[1302] Step 7:

[1303] Update the 3D environment in real time according to the user's movements

[1304] How it works: As the user moves around and looks at objects in the VR space, the environment is updated in real time. Sensors capture movements, and the server calculates and transmits the corresponding environmental changes.

[1305] Input: User movement data (e.g., position change, gaze direction)

[1306] Data Computation: Analyzing motion data and re-rendering the environment

[1307] Output: Updated 3D environment data

[1308] Step 8:

[1309] Recognize user emotions and dynamically change the environment

[1310] How it works: The camera captures the user's facial expressions, and the emotion recognition engine (OpenCV) analyzes them to recognize their emotions. The server dynamically adjusts the virtual reality environment based on this emotional data.

[1311] Input: User's facial expression data, voice tone

[1312] Data Computing: Facial Expression Analysis and Emotion Recognition

[1313] Output: Change the environment based on the emotion (e.g., special discount popup, adding interactive elements)

[1314] Step 9:

[1315] Collect feedback after the experience and refine the system

[1316] Specific operation: After the user finishes the experience, the device displays a feedback form and sends the user's comments and requests for improvement to the server. The server analyzes this data and uses it to improve the system.

[1317] Input: User feedback data (e.g., satisfaction, areas for improvement)

[1318] Data Computing: Feedback Analysis and System Optimization

[1319] Output: Improved system functionality and performance

[1320] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1321] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1322] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1323] [Fourth embodiment]

[1324] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1325] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1326] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1327] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1328] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1329] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1330] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1331] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1332] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1333] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1334] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1335] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1336] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1337] This invention provides a system that allows users to experience specific historical moments in the past or future through virtual reality. Based on user requests, this system collects relevant data from the Internet, generates scenarios using generative artificial intelligence, and provides a mechanism for rendering 3D environments in real time using a virtual reality engine.

[1338] Program Overview

[1339] 1. User submits request:

[1340] The user speaks into the device to request an experience of a specific time in the past or future.

[1341] The device analyzes the voice and converts it into text data.

[1342] 2. Data Collection:

[1343] The terminal transmits the text data to the server.

[1344] The server collects relevant information from the Internet based on the request.

[1345] 3. Data analysis and scenario generation:

[1346] The server analyzes the collected data and generates scenarios using generative artificial intelligence.

[1347] A scenario details the buildings, people, and events associated with a specified time period and place.

[1348] 4. Real-time rendering:

[1349] The server uses the generated scenario to render a 3D environment in a virtual reality engine.

[1350] The 3D environment contains details and realism that match the user's requirements.

[1351] 5. Data transmission and display:

[1352] The server sends the rendered 3D environment data to the terminal.

[1353] The device displays the received data on a virtual reality headset.

[1354] 6. Immersive user experience:

[1355] The user wears a virtual reality headset and is immersed in a virtual environment.

[1356] The environment is updated in real time according to the user's movements and gaze.

[1357] 7. Feedback collection and system improvement:

[1358] After the user has completed the experience, the terminal displays a feedback form in which the user can enter comments about the experience and requests for improvements.

[1359] The device sends feedback to the server, which analyzes it and uses it to improve the system.

[1360] Specific examples

[1361] Case: Renaissance Florence

[1362] 1. Submit your request:

[1363] The user speaks to the device and says, "I want to experience Renaissance Florence."

[1364] The terminal converts this voice into text data and sends it to the server.

[1365] 2. Data Collection:

[1366] The server collects information related to "Renaissance Florence" from the Internet.

[1367] For example, data such as text, images, audio, and video related to buildings, people, and events from that time is obtained.

[1368] 3. Data analysis and scenario generation:

[1369] The server analyzes the collected data and uses generative artificial intelligence to generate a scenario of "Florence during the Renaissance."

[1370] Scenarios include Da Vinci painting and the construction site of Florence Cathedral.

[1371] 4. Real-time rendering:

[1372] Based on the generated scenario, the server uses a virtual reality engine to render a 3D environment of Florence in real time.

[1373] 5. Data transmission and display:

[1374] The server sends the rendered 3D environment data to the terminal.

[1375] The device receives the data and displays it on a virtual reality headset.

[1376] 6. Immersive Experience:

[1377] Users put on a headset and are immersed in a realistic Renaissance Florence.

[1378] Users can interact with Da Vinci and tour the inside of the cathedral.

[1379] 7. Feedback and System Improvement:

[1380] After completing the experience, users fill out a feedback form to indicate their satisfaction with the experience and suggestions for improvement.

[1381] The device sends this to the server, which analyzes the feedback and uses it to improve the system.

[1382] This allows users to experience historical moments and future worlds in a highly immersive way.

[1383] The processing flow will be explained below.

[1384] Step 1:

[1385] The user makes a request to experience a specific time in the past or future by voice into the device, which then recognizes the voice and converts the voice data into text data.

[1386] Step 2:

[1387] The device sends the converted text data to the server as an HTTP request, which includes the user's desired experience.

[1388] Step 3:

[1389] Based on the received request, the server launches a scraping tool to scrape relevant text, image, audio, and video data from the Internet.

[1390] Step 4:

[1391] The server uses a scraping tool to collect data from the Internet that matches the specified keywords, and the collected data is temporarily stored.

[1392] Step 5:

[1393] The server analyzes the collected data using natural language processing and image recognition technology and formats it appropriately to be passed to the generative artificial intelligence.

[1394] Step 6:

[1395] The server inputs the formatted data into a generative AI to generate a scenario for the specified time period and location, including detailed historical background and key buildings, people, and events.

[1396] Step 7:

[1397] The server uses an evaluation algorithm to check the quality of the generated scenario, correcting or regenerating it as necessary. Once the evaluation is complete, the scenario moves on to the next stage.

[1398] Step 8:

[1399] The server uses a virtual reality engine to render a 3D environment in real time based on the scenario, using the GPU to achieve high performance and detailed rendering.

[1400] Step 9:

[1401] The server compresses the rendered 3D environment data and transmits it to the device in an efficient binary format, using streaming techniques to minimize data latency.

[1402] Step 10:

[1403] The device then extracts the received 3D environment data and displays it on the virtual reality headset, which updates the displayed 3D environment in real time according to the user's movements and gaze.

[1404] Step 11:

[1405] Users wear a virtual reality headset and enter an immersive virtual environment, where they can freely move around and experience interactive elements (e.g., interact with people, manipulate objects).

[1406] Step 12:

[1407] When the user finishes their experience, the device displays a feedback form where the user can enter comments about their experience and suggestions for improvement. Collecting feedback contributes to improving user satisfaction.

[1408] Step 13:

[1409] The device sends the collected feedback data to the server, which analyzes the feedback and adjusts the algorithms of the generative AI and virtual reality engine to help improve the system, resulting in an even better user experience next time.

[1410] Example 1

[1411] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1412] Conventional virtual reality systems face challenges when it comes to quickly and accurately generating detailed scenarios and rendering 3D environments in real time when users want to experience specific historical moments or future worlds. Other challenges include updating the 3D environment in real time in response to user movements and systematically collecting user feedback to improve the system.

[1413] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1414] In this invention, the server includes means for converting user requests into text data based on the user requests, means for collecting related information from a communication network, means for generating a scenario using artificial intelligence based on the collected information, means for creating a 3D environment using virtual reality technology based on the generated scenario, means for transmitting the generated 3D environment data to a user device, means for displaying the 3D environment on the user device and updating it in real time according to the user's movements, and means for collecting user responses and improving the system, thereby enabling users to experience specific historical moments or future worlds in detail and realistically.

[1415] "User" refers to a person who uses the system to experience virtual reality.

[1416] A "request" refers to a request for a specific historical moment or future world that a user would like to experience.

[1417] "Text data" refers to data in which a user's request is converted into text form through voice recognition or other input methods.

[1418] "Communications network" means a wide-area computer network, including the Internet and other data communications infrastructures.

[1419] "Information" refers to data such as text, images, and videos collected based on user requests.

[1420] "Artificial intelligence" refers to machine learning algorithms and neural networks used to analyze collected information and generate specified scenarios.

[1421] A "scenario" refers to a set of information that contains a detailed description of a particular historical moment or future world that a user wants to experience.

[1422] "Virtual reality technology" refers to technology that uses computer technology to generate a 3D environment and provide an immersive experience to the user.

[1423] "3D environment" refers to a three-dimensional virtual space created using virtual reality technology based on a scenario.

[1424] "User device" refers to the terminal or headset that a user uses to have a virtual reality experience.

[1425] "Real-time updates" means that the 3D environment is instantly reflected and updated in response to inputs such as the user's movements and gaze.

[1426] "Reactions" refer to the feedback and opinions users provide through their virtual reality experience.

[1427] "System" refers to the comprehensive technological platform that includes all of the above means.

[1428] The present invention is a system that allows users to experience specific historical moments in the past or future through virtual reality. Based on a user's request, the system collects relevant information from the internet, generates a scenario using a generative AI model, and uses virtual reality technology to render a 3D environment in real time.

[1429] First, a user requests a historical moment or future world they would like to experience through a voice request. For example, a user might say, "I want to experience Tokyo of the future." This voice request is converted into text data by the device's voice recognition software (e.g., Google Speech-to-Text API).

[1430] The device sends the converted text data to a server, which then collects related information from the Internet via a communications network. Specifically, it uses web scraping tools (e.g., BeautifulSoup) or APIs (e.g., Wikipedia API) to obtain the necessary text, images, videos, etc.

[1431] The collected information is analyzed by the server and passed to a generative AI model (e.g., OpenAI's GPT-4). This generative AI model creates a detailed scenario based on the collected information. For example, it uses a prompt such as, "Generate a detailed scenario about Tokyo in the future. Include skyscrapers, a modern transportation system, and people's activities."

[1432] Based on the generated scenario, the server uses virtual reality technology to render a 3D environment in real time using a virtual reality engine (e.g., Unreal Engine or Unity). Specifically, the future Tokyo cityscape, transportation system, building interiors, streetscapes, and other details are depicted.

[1433] The generated 3D environment data is sent to the terminal and displayed on the user device (e.g., a virtual reality headset). The user wears the headset and is immersed in the virtual environment. The user's movements and gaze are detected by tracking sensors (e.g., the tracking system in the HTC Vive), and the 3D environment is updated in real time. This allows the user to freely explore the virtual reality and enjoy a detailed experience.

[1434] Finally, once the user has finished the experience, the device displays a feedback form. The user can enter comments about the experience or suggestions for improvement, which the device then sends to the server. The server analyzes the collected feedback and incorporates it into the next update or improvement. This feedback is then used with natural language processing tools and machine learning algorithms to continuously improve the system.

[1435] As described above, the present invention provides a system that provides a user with a high level of immersion and allows them to experience specific historical moments in the past or future in detail and realistically.

[1436] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1437] Step 1: User submits request

[1438] Users speak their requests into the device, for example, saying, "I want to experience Renaissance Florence."

[1439] The device receives this voice input and converts it into text using voice recognition software, specifically the Google Speech-to-Text API.

[1440] Input: User's voice request

[1441] Output: The request converted to character data

[1442] Step 2: Data collection

[1443] The terminal transmits the converted character data to the server.

[1444] Based on the received request, the server collects relevant information from the Internet via a communication network, for example, using a web scraping tool (BeautifulSoup) or an API (Wikipedia API).

[1445] Input: Request converted to character data

[1446] Output: Collected information (text, images, videos, etc.)

[1447] Step 3: Data analysis and scenario generation

[1448] The server analyzes the collected information and organizes it using a natural language processing tool (spaCy).

[1449] The server inputs a prompt into a generative AI model (e.g., OpenAI's GPT-4) to generate a detailed scenario. An example of the prompt is, "Generate a scenario set in Renaissance Florence, including a scene where Da Vinci is painting."

[1450] Input: Collected information

[1451] Output: Generated scenario

[1452] Step 4: Real-time rendering

[1453] Based on the generated scenario, the server uses virtual reality technology to render a 3D environment in real time, specifically using a virtual reality engine (Unreal Engine or Unity).

[1454] Input: Generated scenario

[1455] Output: Rendered 3D environment

[1456] Step 5: Send and display data

[1457] The server sends the rendered 3D environment data to the device, using WebSocket or HTTP / 2 as the communication protocol.

[1458] The device prepares the received data for display on a virtual reality headset.

[1459] Input: Rendered 3D environment data

[1460] Output: Display to a virtual reality headset

[1461] Step 6: User Immersion

[1462] The user wears a virtual reality headset and is immersed in a virtual environment.

[1463] The device uses sensors to track the user's movements and gaze, updating the 3D environment in real time, for example using the tracking system in the HTC Vive.

[1464] Input: User movement and gaze data

[1465] Output: Real-time updated 3D environment

[1466] Step 7: Gather feedback and improve the system

[1467] After the user has finished the experience, the terminal displays a feedback form.

[1468] Users enter comments about their experience and requests for improvement, and the device sends these to the server.

[1469] The server analyzes the collected feedback and uses natural language processing tools and machine learning algorithms to help improve the system.

[1470] Input: User feedback

[1471] Output: Analyzed feedback and system improvements

[1472] The above is the specific processing flow of this system.

[1473] (Application example 1)

[1474] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1475] Current virtual reality systems lack efficient means for users to experience specific historical moments or future events with a high level of realism. Furthermore, technologies for automatically generating scenarios in real time based on collected data and then generating and updating virtual reality environments in real time based on those scenarios remain a challenge. Therefore, systems that improve the quality of the user experience are needed.

[1476] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1477] In this invention, the server includes means for converting a user's request into text data based on the request, means for collecting relevant data from the Internet, means for generating a scenario using a generating artificial intelligence based on the collected data, means for rendering a three-dimensional environment using a virtual reality engine based on the generated scenario, means for transmitting the generated three-dimensional environment data to a user terminal, means for displaying the three-dimensional environment on the user terminal and updating it in real time according to the user's movements, means for collecting user feedback and improving the system, means for providing the generated three-dimensional environment to the user using a visual display device, means for analyzing the collected data and formatting it into an appropriate format for passing it to the generating artificial intelligence, means for ensuring the quality of the scenario using an algorithm for evaluating the quality of the generated scenario, means for rendering the generated three-dimensional environment in real time using a virtual reality engine and transmitting it to a visual display device, and means for displaying the generated three-dimensional environment in real time on the user's mobile device or head-mounted display and updating the environment in accordance with the user's line of sight and movements, thereby enabling users to experience historical moments and future events in real time with a high level of realism.

[1478] A "user request" refers to a user's desire to experience a particular historical moment or future event through audio or text.

[1479] "Text data" refers to data obtained by converting a user's request from voice into text information.

[1480] "Related Data" refers to information about a specified time and place that is collected from the Internet based on a user's request.

[1481] "Generative AI" refers to artificial intelligence technology that generates scenarios based on collected data.

[1482] A "scenario" is a plan or script that details the events and settings to be experienced within a virtual reality environment based on generated data.

[1483] "Virtual reality engine" refers to a software and hardware configuration for rendering three-dimensional environments in real time.

[1484] "Three-dimensional environment" refers to the virtual space that a user experiences through virtual reality.

[1485] "User terminal" refers to a device including a smartphone, a head-mounted display, or other display device.

[1486] "Visual display device" refers to a display device used by a user to visually experience a three-dimensional environment.

[1487] "Feedback" refers to comments, ratings, and requests for improvement regarding the experience provided by the user after the experience.

[1488] "Formatting" refers to the process of converting collected data into an appropriate form for use by the generative AI.

[1489] "Algorithm" refers to a computational procedure for assessing and ensuring the quality of generated scenarios.

[1490] The present invention provides a system that allows a user to experience a specific historical moment or a future event through virtual reality. The system comprises the following means:

[1491] First, the user inputs a request by voice or text using a device such as a smartphone or head-mounted display. For example, the user's request might be something like, "I want to experience a Renaissance city." This request is converted into text data using speech recognition technology (for example, Python's speech_recognition library).

[1492] The server then retrieves the converted text data and collects relevant data from the Internet. During this process, the required data is retrieved from the source using, for example, the requests library. The collected data is then analyzed using generative artificial intelligence (e.g., OpenAI's API) and formatted accordingly.

[1493] The server generates a scenario based on the formatted data. The scenario generated using the AI ​​generator includes detailed events and settings related to the specified time period and location. When generating the scenario, a prompt message corresponding to the user's request is passed to the AI ​​generator. For example, the following prompt message is used:

[1494] Example prompt:

[1495] A user has requested to experience a Renaissance city. Generate a scenario based on the following relevant data:

[1496] Renaissance buildings and streets

[1497] Costumes and culture of the time

[1498] Historical events and everyday scenes

[1499] Generate detailed scenarios and make them suitable for virtual reality experiences.

[1500] The generated scenario is rendered in real time as a 3D environment using a virtual reality engine (e.g., a 3D engine such as py3dengine). The rendered 3D environment data is sent from the server to the user's device, and the user is immersed in the virtual environment through a visual display device (e.g., a head-mounted display or a smartphone).

[1501] After the user has completed their experience in the virtual environment, the server collects feedback from the user by entering comments and ratings into a dedicated form, which is used to improve the system.

[1502] For example, if a user requests "I want to experience a Renaissance city," the server collects relevant data and uses generative artificial intelligence to generate a specific Renaissance scenario, including detailed descriptions of buildings, cityscapes, and costumed people in action. The generated scenario is then rendered as a three-dimensional environment by a virtual reality engine, allowing the user to use a visual display device to tour the city and interact with historical figures.

[1503] As described above, the present embodiment provides a series of processes for enhancing a user's immersive experience.

[1504] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1505] Step 1:

[1506] The user speaks a request into the device. This request describes a desire to experience a specific historical moment or future event. The device recognizes this speech as input data and converts it into text data using speech recognition technology (for example, Python's speech_recognition library). The output is the user's request converted into text data as character information.

[1507] Step 2:

[1508] The terminal sends the converted text data to the server. The server receives the text data as input and collects related data from the Internet. Here, it uses the requests library to obtain the necessary data from the target sources. The output is the collected related data.

[1509] Step 3:

[1510] The server takes the collected data as input, analyzes it using generative AI (for example, OpenAI's API), and formats it into an appropriate format. During this process, the data is structured and necessary information is filtered. The output is data formatted in an appropriate format for the generative AI to generate scenarios.

[1511] Step 4:

[1512] The server uses the formatted data to send prompts to the generation AI to generate a scenario. The prompt text is a specific scenario request based on the user's request and the collected data. The generation AI generates a scenario based on the prompt and provides the scenario in text format as output.

[1513] Step 5:

[1514] The server takes the generated scenario as input and uses a virtual reality engine (e.g., py3dengine) to render a 3D environment in real time. During this process, a 3D environment is generated according to the generated scenario. The output is 3D data rendered as a virtual reality environment.

[1515] Step 6:

[1516] The server transmits the rendered 3D environment data to the user's device, which receives this data as input and displays the virtual environment to the user through a visual display device (such as a head-mounted display or a smartphone). The user is immersed in this virtual environment and experiences it.

[1517] Step 7:

[1518] After the user has completed their experience in the virtual environment, the device displays a feedback form. The user enters comments, ratings, and suggestions for improvement about the experience. The device sends this as input data to the server. The server analyzes the feedback and uses it to improve the system. The output is improvements based on the feedback.

[1519] This series of processing steps allows users to experience historical moments and future events in real time with a high level of realism.

[1520] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1521] This invention is a system that recognizes a user's emotions in a virtual reality environment and dynamically changes the experience content based on those emotions. Based on the user's request, the system converts the request into text data, collects related data from the Internet, and generates a scenario using generative artificial intelligence. Based on the generated scenario, a 3D environment is rendered using a virtual reality engine and displayed on the user's device. Furthermore, by updating the environment in real time based on the user's movements and combining it with an emotion engine that recognizes the user's emotions, an experience tailored to the user's emotions is provided.

[1522] Program Overview

[1523] 1. User submits request:

[1524] The user requests an experience of a specific time in the past or future by speaking into the device, which then analyzes the voice and converts the voice data into text data.

[1525] 2. Data Collection:

[1526] The device sends text data to the server as an HTTP request, and the server collects related information from the Internet based on the request content.

[1527] 3. Data analysis and scenario generation:

[1528] The server analyzes the collected data and uses generative artificial intelligence to generate scenarios that include detailed descriptions of historical background, key buildings, people, and events.

[1529] 4. Real-time rendering:

[1530] The server uses the generated scenario to render a 3D environment in real time in a virtual reality engine, with details and realism that match the user's requirements.

[1531] 5. Data transmission and display:

[1532] The server sends the rendered 3D environment data to the device, which displays it on a virtual reality headset.

[1533] 6. Immersive user experience:

[1534] Users wear a virtual reality headset and are immersed in a virtual environment that updates in real time based on the user's movements and gaze.

[1535] 7. Emotion Recognition with Emotion Engine:

[1536] The server uses an emotion engine that analyzes the user's facial expressions and vocal tone to recognize the user's emotions in real time. The recognized emotion data is analyzed and reflected in the content of the virtual environment.

[1537] 8. Dynamic Environment Changes:

[1538] The server dynamically modifies the virtual reality environment based on the emotional data recognized by the emotion engine, for example by increasing the effects in the environment if the user is surprised, or by adding interactive elements if the user is engrossed.

[1539] 9. Feedback Collection and System Improvement:

[1540] After the user has finished the experience, the device displays a feedback form, where the user can enter comments about the experience and suggestions for improvement. The device then sends this feedback to the server, which then analyzes the feedback and emotional data to help improve the system.

[1541] Specific examples

[1542] Case: Renaissance Florence

[1543] 1. Submit your request:

[1544] The user speaks to the device, saying, "I want to experience Renaissance Florence." The device converts this speech into text data and sends it to the server.

[1545] 2. Data Collection:

[1546] The server collects information related to "Renaissance Florence" from the Internet, including text, images, audio, and video data related to buildings, people, and events from that time.

[1547] 3. Data analysis and scenario generation:

[1548] The server analyzes the collected data and uses generative artificial intelligence to generate a "Renaissance Florence" scenario, including scenes of Da Vinci painting and the construction site of Florence Cathedral.

[1549] 4. Real-time rendering:

[1550] Based on the generated scenario, the server uses a virtual reality engine to render a 3D environment of Florence in real time.

[1551] 5. Data transmission and display:

[1552] The server sends the rendered 3D environment data to the device, which receives the data and displays it on the virtual reality headset.

[1553] 6. Immersive Experience:

[1554] Users put on a headset and are immersed in a realistic Renaissance Florence, where they can interact with Da Vinci and tour the interior of the cathedral.

[1555] 7. Emotion Recognition with Emotion Engine:

[1556] The server analyzes the user's facial expressions and voice using an emotion engine and recognizes that the user is excited.

[1557] 8. Dynamic Environment Changes:

[1558] Based on the user's emotional data, the server adds special events (e.g., fireworks and musical performances) to the streets of Florence to make the experience even more engaging.

[1559] 9. Feedback and System Improvement:

[1560] After completing the experience, the user fills in a feedback form to indicate their satisfaction with the experience and any suggestions for improvement. The device then sends this information to the server, which then analyzes the feedback and emotional data and reflects it in improvements to the system.

[1561] In this way, the system of the present invention recognizes the user's emotions and dynamically modifies the virtual reality experience based on them, providing a richer and more engaging experience.

[1562] The processing flow will be explained below.

[1563] Step 1:

[1564] The user speaks into the device, saying, "I want to experience Renaissance Florence." The device uses voice recognition software to convert the voice data into text.

[1565] Step 2:

[1566] The device sends the converted text data to the server as an HTTP request, which includes the user's desired experience.

[1567] Step 3:

[1568] Based on the received request, the server uses a scraping tool to collect related information (text, image, audio, and video data) that matches the specified keywords from the Internet.

[1569] Step 4:

[1570] The server analyzes the scraped data using natural language processing and image recognition technology and formats it appropriately to be passed on to generative artificial intelligence (AI).

[1571] Step 5:

[1572] The server inputs the formatted data into a generative artificial intelligence (AI) that generates a scenario about "Renaissance Florence," including detailed descriptions of buildings, people, and events.

[1573] Step 6:

[1574] The server uses an evaluation algorithm to verify the quality of the generated scenarios, and corrects or regenerates them as necessary.

[1575] Step 7:

[1576] The server uses quality-tested scenarios to leverage a virtual reality engine to render 3D environments in real time.

[1577] Step 8:

[1578] The server compresses the rendered 3D environment data in binary format and transmits it to the terminal using efficient data transfer techniques.

[1579] Step 9:

[1580] The device receives 3D environment data from the server and displays it on the virtual reality headset. The device tracks the user's gaze and movements in real time and updates the environment.

[1581] Step 10:

[1582] Users put on a virtual reality headset and begin experiencing a realistic Renaissance Florence. They can freely explore the virtual environment and interact with the characters.

[1583] Step 11:

[1584] The server uses an emotion engine to collect the user's facial expressions and voice tone in real time from the camera and microphone installed in the headset and analyze the user's emotions.

[1585] Step 12:

[1586] The server dynamically modifies the virtual reality environment based on the analyzed emotional data, for example adding effects and interactive elements if the user is excited.

[1587] Step 13:

[1588] After the user has finished the experience, the terminal displays a feedback form, where the user can enter comments about the experience and requests for improvements. The terminal then sends the feedback to the server.

[1589] Step 14:

[1590] The server analyzes the feedback and emotional data, adjusts the algorithms of the generative AI and virtual reality engine, and reflects the results in system improvements, resulting in an even better user experience next time.

[1591] Example 2

[1592] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1593] Conventional virtual reality systems lack the flexibility to adapt to user requests and the ability to dynamically change the environment based on the user's emotions. These shortcomings limit the quality of the user experience and make it difficult to provide a richer and more engaging virtual reality experience. Therefore, there is a need for a system that dynamically changes the generation of scenarios and rendering of virtual reality environments based on user requests, depending on the user's emotions.

[1594] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1595] In this invention, the server includes: means for converting a user's request into text data based on the request; means for collecting related data from the Internet; means for generating a scenario using a generative model based on the collected data; means for rendering a three-dimensional environment using a virtual reality engine based on the generated scenario; means for transmitting the generated three-dimensional environment data to a user terminal; means for displaying the three-dimensional environment on the user terminal and updating it in real time according to the user's movements; means including an emotion analysis engine for recognizing the user's emotions; means for dynamically changing the virtual reality environment based on the recognized emotion data; and means for collecting user feedback and improving the system. This makes it possible to flexibly generate scenarios according to user requests and dynamically change the virtual reality environment according to the user's emotions.

[1596] "Means for converting a request into text data" refers to a function for analyzing a request input by a user in a voice or other format and converting it into text.

[1597] "Means of collecting relevant data from the Internet" refers to the function of obtaining necessary information from websites and databases via a network.

[1598] A "generative model" refers to an artificial intelligence algorithm that generates appropriate results for a given input (prompt).

[1599] "Means for generating a scenario" refers to the function of creating stories and events for a virtual reality experience that responds to a user's request based on collected data.

[1600] "Virtual reality engine" refers to a software platform for generating and rendering three-dimensional virtual environments in real time.

[1601] "Means for rendering a three-dimensional environment" refers to a function for depicting a three-dimensional visual environment based on a generated scenario.

[1602] "Means for transmitting three-dimensional environmental data to a user terminal" refers to a function for transferring three-dimensional visual data generated by the server to the terminal used by the user.

[1603] "Means for displaying a three-dimensional environment and updating it in real time in response to the user's movements" refers to a function that dynamically changes the three-dimensional virtual environment displayed on the user's terminal in synchronization with the user's movements and operations.

[1604] An "emotion analysis engine" refers to software that analyzes a user's facial expressions and vocal tone to identify emotions.

[1605] "Means for dynamically modifying the virtual reality environment based on recognized emotional data" refers to the ability to adjust the content and events of the virtual reality in real time based on data obtained through emotional analysis.

[1606] "Means for collecting feedback and improving the system" refers to a function for incorporating user evaluations and opinions of the experience and improving the experience in future.

[1607] The present invention is a system that generates a scenario based on a user's request in a virtual reality environment and provides the user with a three-dimensional environment corresponding to that scenario. This invention makes it possible to recognize the user's emotions in real time and dynamically change the virtual reality environment based on those emotions.

[1608] First, the user speaks into the device to request an experience of a specific time or place. The device then uses voice recognition software, such as the Google Speech-to-Text API, to convert this voice data into text. This text data then becomes the basis for interpreting the request.

[1609] The converted text data is then sent to the server as an HTTP request, which uses web crawling tools and APIs to gather relevant information from the internet, such as Beautiful Soup and Scrapy, to gather text, images, and audio data.

[1610] The collected data is then analyzed using natural language processing tools, such as spaCy and NLTK. Based on the results of this analysis, a generative AI model (e.g., GPT-4, ChatGPT) generates a scenario. Input prompts are crucial for scenario generation. For example, a prompt such as "Walk through the streets of Renaissance Florence" generates detailed historical background and character dialogue.

[1611] Based on the generated scenario, the server uses a virtual reality engine (e.g., Unity, Unreal Engine) to render a three-dimensional virtual environment. The rendered data is sent to the user's device via HTTP. The user's device receives this data and displays it on a virtual reality headset (e.g., Oculus Rift, HTC Vive). The user can then immerse themselves in and interact with this virtual environment.

[1612] During the user's experience, the device uses a camera and microphone to capture the user's facial expressions and voice. This data is analyzed in real time by an emotion analysis engine (e.g., Affectiva, Kairos). Based on the user's emotion data, the server dynamically modifies the virtual environment. For example, if the user is surprised, effects such as fireworks are added to the virtual environment.

[1613] After the experience is over, the device displays a feedback form to the user, allowing them to input their satisfaction with the experience and suggestions for improvement. This data is then sent to the server, which analyzes the feedback and emotional data to improve the experience for the next session.

[1614] A concrete example is a user request to "experience Renaissance Florence." In this case, the user speaks the request into the device, which converts this speech into text data and sends it to the server. The server collects relevant information from the Internet and generates a scenario using a generative AI model. It then renders a three-dimensional environment of Florence using a virtual reality engine and sends it to the user's device. The user puts on the virtual reality headset and explores the city of Florence. If the user is excited, the server adds special events to make the experience even more engaging.

[1615] This invention makes it possible to flexibly generate scenarios based on user requests and dynamically change the virtual reality environment in response to the user's emotions.

[1616] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1617] Step 1:

[1618] User request submission

[1619] Input: The user speaks into the device, "I want to experience Renaissance Florence."

[1620] How it works: The device uses speech recognition software (e.g., Google Speech-to-Text API) to convert speech into text. Voice input is text, and output is text.

[1621] Output: The converted text data. The text is "I want to experience Renaissance Florence."

[1622] Step 2:

[1623] Data collection

[1624] Input: The request as text data.

[1625] How it works: The device sends this text data to the server as an HTTP request. The server then uses a web crawling tool (e.g., Beautiful Soup, Scrapy) to collect relevant information from the Internet. This includes data such as text, images, and audio.

[1626] Output: A dataset containing relevant information.

[1627] Step 3:

[1628] Data analysis and scenario generation

[1629] Input: A dataset containing the relevant collected information.

[1630] How it works: The server uses natural language processing tools (e.g., spaCy, NLTK) to parse the data and format it into a suitable format for a generative AI model (e.g., GPT-4, ChatGPT). It then generates a scenario based on the generative AI model. For example, you can input the prompt "Florence during the Renaissance," and the scenario will be output.

[1631] Output: Generated scenario data.

[1632] Step 4:

[1633] Real-time rendering

[1634] Input: Generated scenario data.

[1635] How it works: The server uses a virtual reality engine (e.g., Unity, Unreal Engine) to render a 3D environment in real time. Based on the scenario, models of objects and characters in the environment are generated and rendered in real time.

[1636] Output: Rendered 3D environment data.

[1637] Step 5:

[1638] Sending and Displaying Data

[1639] Input: Rendered 3D environment data.

[1640] How it works: The server sends this data to the user's device, which then displays it on a virtual reality headset (e.g., Oculus Rift, HTC Vive).

[1641] Output: A three-dimensional environment displayed in a virtual reality headset.

[1642] Step 6:

[1643] Immersive User Experience

[1644] Input: A three-dimensional environment displayed on a virtual reality headset.

[1645] How it works: The user puts on a virtual reality headset and is immersed in a virtual environment. The environment updates in real time based on the user's movements and gaze. For example, the user can explore and interact with Renaissance Florence.

[1646] Output: Updated virtual environment and user experience data.

[1647] Step 7:

[1648] Emotion recognition by emotion engine

[1649] Input: User's facial expressions and voice.

[1650] How it works: The device captures the user's facial expressions and voice using a camera and microphone. The server analyzes this data in real time using an emotion analysis engine (e.g., Affectiva, Kairos) to recognize the user's emotions.

[1651] Output: Recognized emotion data.

[1652] Step 8:

[1653] Dynamic Environment Changes

[1654] Input: Recognized emotion data.

[1655] Action: The server dynamically changes the virtual reality environment based on the emotional data, for example adding special events or effects if the user is excited.

[1656] Output: A dynamically modified virtual reality environment.

[1657] Step 9:

[1658] Feedback collection and system improvement

[1659] Input: User feedback after the experience.

[1660] How it works: After the experience is over, the device displays a feedback form to the user. The user enters their satisfaction with the experience and suggestions for improvement, and this data is sent to the server. The server analyzes the feedback and emotion data to help improve the system.

[1661] Output: An improved system and updated user experience data.

[1662] (Application example 2)

[1663] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1664] Conventional virtual reality systems provide a fixed experience without taking the user's emotions into consideration. This means that the experience cannot be dynamically changed based on the user's emotions and reactions, resulting in a lack of immersion and personalization. Furthermore, it has not been possible for the system to recognize the user's emotions during the experience and provide an optimized experience based on those emotions. This has led to issues such as low user satisfaction and engagement.

[1665] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1666] In this invention, the server includes means for converting a user's request into text data based on the user's request, means for collecting related data from the Internet, means for generating a scenario based on the collected data using artificial intelligence, means for rendering a 3D environment based on the generated scenario using a virtual reality engine, means for transmitting the generated 3D environment data to a user terminal, means for displaying the 3D environment on the user terminal and updating it in real time according to the user's movements, means for recognizing the user's emotions and dynamically changing the virtual reality environment based on the recognized emotions, and means for collecting user feedback and improving the system. This enables dynamic changes to the virtual reality experience based on the user's emotions, providing a more immersive and personalized experience.

[1667] A "user request" is a request to the system for content or information that the user wants to experience.

[1668] A "means for converting to text data" is a process or device that analyzes and converts voice or other non-text data into written information.

[1669] "Means for collecting relevant data from the Internet" refers to a process or device that searches for and obtains specified information or data via the Internet.

[1670] "Generative AI" is an AI technology that generates scenarios and content for specific purposes based on collected data.

[1671] A "scenario generation means" is a process or device that analyzes collected data and arranges it into a specific story or scenario format.

[1672] A "virtual reality engine" is software or hardware that uses computer graphics technology to create and render virtual reality environments.

[1673] A "means for rendering a 3D environment" is a process or device that depicts a three-dimensional virtual space in real time based on a generated scenario.

[1674] The "means for transmitting the generated 3D environment data to the user terminal" is a process or device that transfers the rendered 3D environment data to the terminal used by the user.

[1675] "Means for displaying a 3D environment on a user device and updating it in real time in response to the user's movements" refers to a process or device that displays a 3D virtual space on a user device and dynamically changes that virtual space in response to the user's operations and movements.

[1676] "Means for recognizing a user's emotions and dynamically modifying the virtual reality environment based on the recognized emotions" refers to a process or device that analyzes a user's facial expressions and voice to determine their emotions and adjusts the VR experience to match those emotions.

[1677] "Means for collecting user feedback and improving the system" refers to a process or device that collects opinions and ratings provided by users after their experience and uses that data to update the system's functionality and performance.

[1678] The present invention is a system that recognizes a user's emotions and dynamically modifies a virtual reality experience based on the emotions. The system can generate experience content and render a virtual reality environment in real time based on the user's requests.

[1679] Hardware and software used

[1680] Hardware:

[1681] PC or smartphone with a camera

[1682] Virtual reality headsets (e.g., general VR devices)

[1683] software:

[1684] OpenCV library for emotion recognition

[1685] Hugging Face Transformers for Scenario Generation

[1686] Custom VR engine for real-time rendering (e.g. Unreal Engine Server)

[1687] API for collecting internet information

[1688] System Program Overview

[1689] Based on the user's request, the server converts the request into text data. Requests made by voice input are converted into text data using a highly accurate voice recognition library. Based on the text data, the server collects related data from the Internet. This is done using a specific search API to collect the required information.

[1690] The collected data is passed to a generative AI model, which generates scenarios using Hugging Face's Transformers. The generated scenarios are then rendered into a 3D environment using a virtual reality engine (e.g., an Unreal Engine server).

[1691] The rendered 3D environment data is sent to the device and displayed on the user's virtual reality headset. The 3D environment is updated in real time according to the user's movements. This updating allows the user to move freely around the VR space.

[1692] The system uses a camera and microphone to capture the user's facial expressions and voice, and then analyzes their emotions in real time using an emotion recognition engine (OpenCV). Based on the analyzed emotional data, the virtual reality environment is dynamically modified. For example, if the user is excited, special effects are added in the environment, and if the user is relaxed, the atmosphere of the environment is made calmer.

[1693] Finally, after the user has finished the experience, the terminal displays a feedback form where the user can enter comments and suggestions about the experience. This feedback information is sent to the server and used to continuously improve the system.

[1694] Specific examples

[1695] A user requests a virtual shopping store by speaking, "I want to buy new shoes." This request is converted into text data, and the server collects related information. Based on the collected data, an AI model generates a shopping scenario, and a VR engine renders a 3D environment. When the user puts on a virtual reality headset and enters the virtual store, a camera captures the user's facial expressions, and emotional data is analyzed. If the user finds a pair of shoes and expresses joy, the system displays a pop-up in the storefront offering a special discount.

[1696] Prompt Sentence Examples

[1697] User request: "I want to buy new shoes from a virtual shopping store."

[1698] AI-generated prompt: "Generate a scenario in which a user searches for shoes in a virtual shopping store."

[1699] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1700] Step 1:

[1701] The user speaks the request

[1702] Specific operation: The user says to the terminal, "I want to buy new shoes at a virtual shopping store."

[1703] Input: Audio data

[1704] Data calculation: Converts voice data into text data using a voice recognition library.

[1705] Output: Text data (e.g., "I want to buy new shoes at a virtual shopping store")

[1706] Step 2:

[1707] Sends requests to the server and collects relevant data

[1708] Specific operation: The terminal sends the converted text data to the server as an HTTP request, and the server searches the Internet for information related to the request content.

[1709] Input: Text data (request content)

[1710] Data Computing: Collecting relevant information from multiple databases and sources on the Internet.

[1711] Output: Collected relevant data (e.g. product information, images, reviews)

[1712] Step 3:

[1713] Analyze the collected data and format it appropriately

[1714] Specific operation: The server analyzes the collected data and formats it appropriately to be passed to the generative AI model.

[1715] Input: Relevant data collected

[1716] Data operations: Data analysis and formatting (e.g., text analysis, data structuring)

[1717] Output: Formatted data

[1718] Step 4:

[1719] Generate a scenario

[1720] Specific operation: Based on the collected data, the server sends prompts to the generative AI model (Hugging Face Transformers) to generate a shopping scenario.

[1721] Input: Formatted data, prompt (e.g., "Generate a scenario in which a user searches for shoes in a virtual shopping store.")

[1722] Data calculation: scenario generation using generative AI models

[1723] Output: Generated scenario

[1724] Step 5:

[1725] Rendering a Virtual Reality Environment

[1726] How it works: The server sends the generated scenario to a VR engine (e.g., Unreal Engine server), which renders the 3D virtual environment in real time.

[1727] Input: Generated scenario

[1728] Data Computing: 3D Modeling and Rendering with a VR Engine

[1729] Output: 3D environment data

[1730] Step 6:

[1731] 3D environment data is sent to the user's device and displayed

[1732] Specific operation: The server sends the rendered 3D environment data to the user's device, which then displays the data on the VR headset.

[1733] Input: 3D environment data

[1734] Data Computing: Communication and Data Transfer

[1735] Output: 3D virtual environment displayed in a VR headset

[1736] Step 7:

[1737] Update the 3D environment in real time according to the user's movements

[1738] How it works: As the user moves around and looks at objects in the VR space, the environment is updated in real time. Sensors capture movements, and the server calculates and transmits the corresponding environmental changes.

[1739] Input: User movement data (e.g., position change, gaze direction)

[1740] Data Computation: Analyzing motion data and re-rendering the environment

[1741] Output: Updated 3D environment data

[1742] Step 8:

[1743] Recognize user emotions and dynamically change the environment

[1744] How it works: The camera captures the user's facial expressions, and the emotion recognition engine (OpenCV) analyzes them to recognize their emotions. The server dynamically adjusts the virtual reality environment based on this emotional data.

[1745] Input: User's facial expression data, voice tone

[1746] Data Computing: Facial Expression Analysis and Emotion Recognition

[1747] Output: Change the environment based on the emotion (e.g., special discount popup, adding interactive elements)

[1748] Step 9:

[1749] Collect feedback after the experience and refine the system

[1750] Specific operation: After the user finishes the experience, the device displays a feedback form and sends the user's comments and requests for improvement to the server. The server analyzes this data and uses it to improve the system.

[1751] Input: User feedback data (e.g., satisfaction, areas for improvement)

[1752] Data Computing: Feedback Analysis and System Optimization

[1753] Output: Improved system functionality and performance

[1754] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1755] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1756] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1757] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1758] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1759] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1760] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1761] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1762] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1763] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1764] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1765] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1766] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1767] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1768] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1769] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1770] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1771] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1772] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1773] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1774] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1775] The following is further disclosed regarding the above embodiment.

[1776] (Claim 1)

[1777] means for converting a request into text data based on the request of a user;

[1778] means for collecting relevant data from the internet;

[1779] A means for generating a scenario using a generative artificial intelligence based on the collected data;

[1780] A means for rendering a 3D environment using a virtual reality engine based on the generated scenario; and

[1781] means for transmitting the generated 3D environment data to a user terminal;

[1782] A means for displaying a 3D environment on a user device and updating it in real time according to the user's movements;

[1783] a means of collecting user feedback and improving the system;

[1784] A system including:

[1785] (Claim 2)

[1786] Provide a means for analyzing the collected data and formatting it into an appropriate format for passing to the generating artificial intelligence;

[1787] 10. The system of claim 1.

[1788] (Claim 3)

[1789] having means for using evaluation algorithms to evaluate the generated scenarios and ensure their quality;

[1790] 10. The system of claim 1.

[1791] "Example 1"

[1792] (Claim 1)

[1793] means for converting a request into character data based on the user's request;

[1794] means for collecting relevant information from a communications network;

[1795] a means for generating scenarios using artificial intelligence based on the collected information;

[1796] A means for creating a 3D environment using virtual reality technology based on the generated scenario;

[1797] means for transmitting the generated 3D environment data to a user device;

[1798] means for displaying a 3D environment on a user device and updating the 3D environment in real time in response to user movements;

[1799] a means of collecting user feedback and improving the system;

[1800] A system including:

[1801] (Claim 2)

[1802] Provide a means to analyze the collected information and convert it into an appropriate format for passing to artificial intelligence;

[1803] 10. The system of claim 1.

[1804] (Claim 3)

[1805] have a means to evaluate the generated scenarios and use evaluation methods to ensure their quality;

[1806] 10. The system of claim 1.

[1807] "Application Example 1"

[1808] (Claim 1)

[1809] means for converting a request into text data based on the request of a user;

[1810] means for collecting relevant data from the internet;

[1811] A means for generating a scenario using a generative artificial intelligence based on the collected data;

[1812] a means for rendering a three-dimensional environment using a virtual reality engine based on the generated scenario;

[1813] means for transmitting the generated three-dimensional environment data to a user terminal;

[1814] means for displaying a three-dimensional environment on a user terminal and updating it in real time in response to the user's movements;

[1815] a means of collecting user feedback and improving the system;

[1816] means for presenting the generated three-dimensional environment to a user using a visual display device;

[1817] A system including:

[1818] (Claim 2)

[1819] A means for analyzing the collected data and formatting it into an appropriate format for passing to the generative artificial intelligence;

[1820] A means for ensuring the quality of the generated scenarios using an algorithm for evaluating the quality of the scenarios;

[1821] means for rendering the generated three-dimensional environment in real time using a virtual reality engine and transmitting it to a visual display device;

[1822] 10. The system of claim 1.

[1823] (Claim 3)

[1824] 10. The system of claim 1, further comprising means for displaying the generated three-dimensional environment in real time on a user's mobile device or head-mounted display, and updating the environment according to the user's gaze and movements.

[1825] "Example 2: Combining Emotion Engines"

[1826] (Claim 1)

[1827] means for converting a request into text data based on the request of a user;

[1828] means for collecting relevant data from the internet;

[1829] A means for generating scenarios using a generative model based on the collected data;

[1830] a means for rendering a three-dimensional environment using a virtual reality engine based on the generated scenario;

[1831] means for transmitting the generated three-dimensional environment data to a user terminal;

[1832] means for displaying a three-dimensional environment on a user terminal and updating it in real time in response to the user's movements;

[1833] means including an emotion analysis engine for recognizing an emotion of a user;

[1834] means for dynamically modifying the virtual reality environment based on the recognized emotion data;

[1835] a means of collecting user feedback and improving the system;

[1836] A system including:

[1837] (Claim 2)

[1838] Providing a means to parse the collected data and format it appropriately for passing to the generative model;

[1839] 10. The system of claim 1.

[1840] (Claim 3)

[1841] having means for using evaluation algorithms to evaluate the generated scenarios and ensure their quality;

[1842] 10. The system of claim 1.

[1843] "Application example 2 when combining emotion engines"

[1844] (Claim 1)

[1845] means for converting a request into text data based on the request of a user;

[1846] means for collecting relevant data from the internet;

[1847] A means for generating a scenario using a generative artificial intelligence based on the collected data;

[1848] A means for rendering a 3D environment using a virtual reality engine based on the generated scenario; and

[1849] means for transmitting the generated 3D environment data to a user terminal;

[1850] A means for displaying a 3D environment on a user device and updating it in real time according to the user's movements;

[1851] means for recognizing a user's emotion and dynamically modifying the virtual reality environment based on the recognized emotion;

[1852] a means of collecting user feedback and improving the system;

[1853] A system including:

[1854] (Claim 2)

[1855] Provide a means for analyzing the collected data and formatting it into an appropriate format for passing to the generating artificial intelligence;

[1856] 10. The system of claim 1.

[1857] (Claim 3)

[1858] a means for using an evaluation algorithm to evaluate the generated scenarios and ensure their quality;

[1859] Equipped with a means to analyze user emotional data and reflect it in the experience content.

[1860] 10. The system of claim 1. [Explanation of symbols]

[1861] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. means for converting a request into text data based on the request of a user; means for collecting relevant data from the internet; A means for generating a scenario using a generative artificial intelligence based on the collected data; A means for rendering a 3D environment using a virtual reality engine based on the generated scenario; and means for transmitting the generated 3D environment data to a user terminal; A means for displaying a 3D environment on a user device and updating it in real time according to the user's movements; a means of collecting user feedback and improving the system; A system including:

2. Provide a means for analyzing the collected data and formatting it into an appropriate format for passing to the generating artificial intelligence; The system of claim 1 .

3. having means for using evaluation algorithms to evaluate the generated scenarios and ensure their quality; The system of claim 1 .

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A