system
The system addresses travel barriers by delivering destination content to terminals for augmented or virtual reality, providing immersive experiences and continuous improvement through user interaction and feedback.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-10-21
- Publication Date
- 2026-05-07
AI Technical Summary
Existing travel experiences are limited by physical movement, language barriers, time constraints, and health issues, making it difficult for individuals to enjoy diverse destinations easily.
A system that provides destination-related content to a user's terminal, generating an augmented or virtual reality environment based on user input, allowing interaction and feedback for a realistic experience without actual travel.
Enables immersive travel experiences from home, overcoming physical limitations and enhancing user engagement through dynamic interaction and continuous system improvement.
Smart Images

Figure 2026074929000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of the chatbot's character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] To enjoy traveling, it often involves physical movement, language barriers, time and economic constraints. There are also cases where it is difficult to go out due to health reasons. A method that can solve these problems and allow anyone to easily experience various places is required.
Means for Solving the Problems
[0005] The present invention provides means for acquiring content related to a destination based on user input information and transferring the data to the user's terminal. Furthermore, by providing means for generating an extended reality environment using the content received by the terminal and controlling interactions within that environment based on the user's actions, it becomes possible to provide a realistic travel experience without actually traveling.
[0006] "User input information" refers to data such as destinations, preferred activities, and languages used, which users provide through the system.
[0007] "Destination-related content" refers to the collective digital data, including images, 3D models, audio data, and related information concerning the selected travel destination.
[0008] "Transferring means" refers to the technical methods and processes for sending data from a server to a user's terminal.
[0009] An "augmented reality environment" refers to a virtual visual environment that overlays digital information onto the real world.
[0010] "Means of controlling interaction" refers to technical elements that adjust the display of objects and information within the augmented reality environment in response to user input and actions. [Brief explanation of the drawing]
[0011] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which a plurality of emotions are mapped. [Figure 10] It shows an emotion map to which a plurality of emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Example 2 when an emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when an emotion engine is combined.
Mode for Carrying Out the Invention
[0012] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.
[0013] First, the terms used in the following description will be explained.
[0014] In the following embodiments, the signed processor (hereinafter simply referred to as "processor") may be one arithmetic unit or a combination of a plurality of arithmetic units. Further, the processor may be one type of arithmetic unit or a combination of a plurality of types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like. ]>
[0015] In the following embodiments, the tagged RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by a processor.
[0016] In the following embodiments, the tagged storage is one or more non-volatile storage devices that store various programs, various parameters, and the like. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.
[0017] In the following embodiments, the tagged communication I / F (Interface) is an interface including a communication processor, an antenna, and the like. The communication I / F controls communication between a plurality of computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0018] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B". That is, "A and / or B" means that it may be only A, only B, or a combination of A and B. Also, in this specification, when three or more matters are connected and expressed by "and / or", the same concept as "A and / or B" is applied.
[0019] [First Embodiment]
[0020] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0021] As shown in FIG. 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0022] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0023] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0024] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0025] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0026] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0027] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0028] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0029] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0030] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0031] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0032] This invention provides a system that enables users to experience travel from home through their own devices. In this system, the user selects a travel destination using an application. Information based on the selection is sent to a server, which retrieves the corresponding content from a database. The content includes high-resolution images, 3D models, and audio data of the destination. The server prepares this content for streaming to the user's device and begins the transfer.
[0033] The user's device processes the received data and generates an augmented reality environment through smart glasses or a headset. This environment is built based on the received data and provides realistic images in the user's field of view. When the user wears the device, they can enjoy a 360-degree virtual journey that responds to head movements and settings.
[0034] For example, if a user chooses to experience "Virtual Kyoto," the server prepares detailed 3D models and audio guides of Kyoto's famous landmarks. The user's device receives this information and uses AR technology to provide an experience of visiting various shrines and temples from the comfort of their home. During this process, the user can use hand movements and voice commands to ask for more detailed information or switch viewpoints.
[0035] Furthermore, users who have completed the experience can provide feedback on their satisfaction level and areas for improvement. This feedback is collected by the server and used to improve the system. In this way, the present invention enables diverse travel experiences that transcend physical limitations and provides new value to users.
[0036] The following describes the processing flow.
[0037] Step 1:
[0038] The user opens the application, selects their desired travel destination, and specifies related activities and preferred language. Once this information is entered, the device sends the request to the server.
[0039] Step 2:
[0040] The server parses the received request and retrieves content related to the destination from the database, such as high-resolution images, 3D models, and audio data. It then organizes this information appropriately and prepares it for data transfer.
[0041] Step 3:
[0042] The server begins streaming the organized content to the user's device. The transfer takes place in real time, and compression techniques are used as needed to optimize the process for smooth data reception.
[0043] Step 4:
[0044] The terminal receives data from the server and decompresses and decodes the received content. This prepares it for building the augmented reality environment. The terminal's processor performs the necessary processing to ensure the system operates smoothly.
[0045] Step 5:
[0046] The device generates an augmented reality environment through smart glasses or a headset. When the user wears the device, 3D objects are overlaid on their field of vision and displayed superimposed on the real world. This allows the user to feel as if they are right in front of their destination.
[0047] Step 6:
[0048] Users can interact within the augmented reality environment using voice commands and hand gestures. For example, they can select specific objects to display detailed information or converse with virtual guides.
[0049] Step 7:
[0050] Once a user finishes their experience, the device collects feedback from the user regarding their satisfaction level and areas for improvement, and sends it to the server. This allows information about system improvements to be accumulated on the server, enriching future experiences.
[0051] (Example 1)
[0052] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0053] There is a growing demand for virtual environments that allow users to experience the real world with greater realism without physical travel. However, current technology fails to deliver a satisfactory experience due to insufficient collection and real-time display of location-related data, as well as inadequate user interaction. Furthermore, it lacks the functionality to effectively utilize user feedback and continuously improve the virtual environment.
[0054] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0055] In this invention, the server includes means for searching for location-related data based on user selection information, means for transmitting the collected data to the user's information processing device, and means for constructing a virtual reality environment using the data received by the information processing device. This enables the user to have a virtual experience with a sense of presence.
[0056] A "user" refers to an individual or group that operates the system and experiences the virtual environment.
[0057] "Selection information" refers to information provided to users to determine the virtual location or environment they wish to visit.
[0058] "Location-related data" refers to a collection of information such as high-resolution images, 3D models, and audio guides about the location selected by the user.
[0059] An "information processing device" refers to an electronic device, such as a computer or smart device owned by a user, that has the function of receiving and processing data.
[0060] "Transmission" refers to the act or means of sending data from a server to an information processing device.
[0061] A "virtual reality environment" refers to a digitally constructed space that does not exist in the real world and can be experienced through user interaction.
[0062] "Interaction" refers to the two-way exchanges, such as operations and responses, that take place between the user and the virtual reality environment.
[0063] "Opinions" refer to feedback from users regarding their experiences and suggestions for improvement.
[0064] A "storage device" is an electronic storage medium used to store data so that it can be retrieved later.
[0065] "Eye movement" refers to actions related to changes in a user's visual focus or eye movements.
[0066] "Instant adjustment" refers to changing the displayed content of the virtual reality environment in real time in response to the user's actions and requests.
[0067] This invention relates to a system that realizes a virtual reality environment using a user's computer or smart device (information processing device). The user installs an application on their device and launches it to use the system. Based on the user's selection information, the user determines the virtual location they wish to visit. For example, the user can choose to experience "Virtual Kyoto."
[0068] The server retrieves data related to the selected location from its internal data store based on the customer's selection. This data includes high-resolution images, 3D models, and audio guides. The server then transmits this data to the user's device. This transmission requires efficient data transfer and typically uses real-time communication protocols over the internet.
[0069] The device receives streamed data and uses AR technologies such as Unity and Unreal Engine to construct a virtual reality environment. This constructed environment is displayed and experienced by the user through smart glasses or a headset. The device instantly adjusts the display within the virtual environment in response to the user's gaze and movements. This dynamic interaction enhances immersion and provides a more realistic experience for the user.
[0070] Users can manipulate the environment using hand movements and voice commands to acquire information and switch viewpoints. User feedback is collected by computer after the experience ends, sent to a server, and stored in memory. This feedback data is used to improve the virtual environment.
[0071] As an example of a generated AI model prompt, "I want to visualize a 3D model of Kyoto and play an audio guide about the temples" provides parameters useful for data retrieval and display. Users can have diverse and immersive travel experiences that go beyond physical limitations.
[0072] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0073] Step 1:
[0074] The user launches an application on their device and selects a virtual location they wish to visit. As input, the user provides destination selection information. This selection information serves as a basis for the system to identify the location data needed in the next step. As output, the selection information is sent to the server.
[0075] Step 2:
[0076] The server searches the data store for data related to the corresponding location based on the selection information it receives. It receives user selection information as input. Data processing involves executing database queries to extract datasets containing high-resolution images, 3D models, audio guides, etc. The collected data is then prepared on the server as output.
[0077] Step 3:
[0078] The server prepares the collected data for transmission to the user's terminal. The retrieved data is used as input. Data processing involves compression and encoding, converting the data into an optimized format for efficient data transfer. As output, the data is ready and transmission begins.
[0079] Step 4:
[0080] The terminal receives data from the server and constructs the virtual reality environment. It receives transmitted images, 3D models, and audio data as input. It uses Unity or Unreal Engine for data processing to generate the virtual reality environment. The generated AR environment is displayed on the display device as output.
[0081] Step 5:
[0082] Users experience the AR environment through smart glasses or a headset. They manipulate the environment through their gaze and body movements, and interact with it through voice and motion. The device receives user gaze and motion information as input, and adjusts the virtual environment display in real time based on this information. The output is an interactive experience that responds to movement.
[0083] Step 6:
[0084] After a user completes an experience, feedback is provided within the system. As input, user opinions and satisfaction data are entered into the device. This data is sent to a server and stored in memory. As output, the feedback data is used for improvement.
[0085] (Application Example 1)
[0086] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0087] Conventional travel experience systems make it difficult to realistically experience diverse travel destinations from the comfort of one's home. Furthermore, they lack features that allow for interaction that reflects user instructions and effectively incorporate post-experience feedback into the system. This makes it difficult to provide truly immersive remote experiences, and also hinders the continuous improvement of the system based on user input.
[0088] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0089] In this invention, the server includes means for acquiring a set of information related to a destination based on user input information, means for transferring the acquired set of information to the user's information processing device, means for generating an augmented reality environment using the information received by the information processing device, and means for operating the functions of the information processing device in response to the user's voice instructions using voice recognition technology. This makes it possible for the user to realistically experience a variety of travel destinations from the comfort of their home, and also enables flexible operation of the augmented reality environment based on the user's voice input and improvement of the system based on the feedback received.
[0090] "User input information" refers to instructions and information related to the travel destination or experience selected by the user through the information processing device.
[0091] "Destination-related information" refers to a collection of digital information, including geographical information, visual data, and audio guides for the selected travel destination.
[0092] An "information processing device" is a terminal used by users to experience virtual reality, and refers to devices such as smartphones and headsets.
[0093] "Transferring means" refers to the processes and technologies used to send data from a server to an information processing device.
[0094] An "augmented reality environment" is a computer-generated image environment that overlays digital information onto real-world environmental information for display.
[0095] "Speech recognition technology" is a technology that interprets the voice spoken by a user and allows a computer to respond appropriately based on that interpretation.
[0096] "Voice commands" are commands or requests that a user makes to a system or device using their voice.
[0097] A system implementing this invention begins by installing an application equipped with a user interface for the user to select a travel destination on an information processing device. When the user selects a travel destination, that information is sent to a server. The server retrieves a set of information related to the destination selected by the user from a database and transfers the information, including geographic information, visual data, and audio guides, that constitutes that set of information to the information processing device.
[0098] The information processing device generates an augmented reality environment using the information received from the server. By utilizing software such as ARKit (iOS) and ARCore (Android®), users can visually experience digital information superimposed onto the real world. For example, if a user selects "Virtual Paris Trip," detailed 3D models of the Eiffel Tower and the Louvre Museum are delivered and displayed in the augmented reality environment.
[0099] Furthermore, the information processing device interprets user voice commands using speech recognition technology (such as Google® Speech-to-Text). This allows users to interact and operate within the environment through voice. For example, if a user asks, "What is the height of the Eiffel Tower?", the information will be provided immediately.
[0100] This system allows users to experience visiting tourist destinations around the world from the comfort of their homes, and to provide feedback based on their experiences, contributing to system improvements. By inputting prompts such as, "Provide detailed information about the travel destination chosen by the user as a virtual experience with a visually represented 3D model and audio support," into the generating AI model, a more realistic and immersive experience can be achieved.
[0101] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0102] Step 1:
[0103] The user selects a travel destination through the application interface of the information processing device. The input is the name of the travel destination selected by the user. Based on the user's selection, the information is sent to the server.
[0104] Step 2:
[0105] The server receives the name of the travel destination sent by the user, and retrieves a set of information related to the destination, such as visual data, geographical information, and audio guides, by referring to the database. The input is the travel destination selected by the user, and the output is a set of various information related to that travel destination.
[0106] Step 3:
[0107] The server encodes the acquired information and streams or downloads it to the information processing device. The input is the information acquired by the server, and the output is the data transferred to the information processing device.
[0108] Step 4:
[0109] The information processing device uses ARKit or ARCore to generate an augmented reality environment based on the information received from the server. The input is the received information, and the output is the augmented reality environment presented to the user.
[0110] Step 5:
[0111] Users use voice commands to ask questions of or operate the information processing device. The input is the voice commands spoken by the user.
[0112] Step 6:
[0113] The information processing device uses speech recognition technology to convert user voice commands into text data and displays appropriate information in an augmented reality environment based on that input. The input is the user's voice command, and the output is the corresponding information or operation result.
[0114] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0115] This invention provides a personalized experience tailored to the user's emotional state by combining an emotion engine with a travel experience system using an augmented reality environment. This system operates by integrating the user's terminal, server, and emotion engine.
[0116] First, the user selects a destination they wish to visit through the application and enters relevant information. Once this information is sent to the server via the device, the server retrieves content related to the destination from its database and streams it to the user's device.
[0117] The user's device generates an augmented reality environment based on the received content and provides visual information to the user through smart glasses or a headset. Furthermore, the device uses a built-in emotion engine to infer emotions from the user's facial expressions, tone of voice, and other biometric information.
[0118] The emotion engine analyzes the user's emotions in real time and adjusts the content and interactions within the augmented reality environment based on this analysis. For example, if the user is relaxed, it provides a calm environment; if they are excited, it provides a more stimulating experience.
[0119] As a concrete example, suppose a user chooses to experience "Virtual New York." The server retrieves a 3D model of Times Square and event information and streams it. The device then constructs this as an augmented reality environment and presents it to the user's field of view. The emotion engine analyzes the user's reactions, and if the user is surprised or delighted, it suggests more fun attractions or displays information tailored to their interests.
[0120] After the experience ends, emotional data and feedback collected from users are sent to a server and used to improve future experiences. This allows for the provision of travel experiences optimized for each individual user. This invention enables dynamic analysis of emotions and contributes to further personalizing the user experience.
[0121] The following describes the processing flow.
[0122] Step 1:
[0123] The user launches the application and selects a destination they wish to visit. They then enter information such as related activities and preferred language, and send this information to the server via their device.
[0124] Step 2:
[0125] The server analyzes the received request and retrieves high-resolution images, 3D models, and audio data related to the destination from the database. It then packages this data and prepares it for streaming to the user's device.
[0126] Step 3:
[0127] The device receives content sent from the server. It then decompresses and decodes the content, preparing to generate the augmented reality environment.
[0128] Step 4:
[0129] The device's built-in emotion engine collects user facial expression data and voice tone. Based on this, it analyzes the user's emotional state in real time.
[0130] Step 5:
[0131] The device generates an augmented reality environment and displays it in the user's field of view using smart glasses or a headset. The content and audio guidance presented are dynamically adjusted according to the user's emotional state.
[0132] Step 6:
[0133] Users enjoy an interactive experience based on analysis by an emotion engine. For example, if a user is relaxed, calming music and scenery are emphasized, while if they are excited, active activities are presented.
[0134] Step 7:
[0135] When a user finishes an experience, the device sends emotional data and user feedback to the server. This information is stored in a database for system improvement and further personalization.
[0136] (Example 2)
[0137] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0138] Traditional augmented reality systems have struggled to dynamically adjust content based on user emotions and individual actions, making it difficult to provide an experience optimized for each user. Therefore, there is a growing need to provide travel experiences that are more adaptable to users' emotions and interests.
[0139] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0140] In this invention, the server includes means for acquiring destination-related information based on user input information, means for transferring the acquired information to the user's device, and means for generating an augmented reality environment using the information received by the device. This enables dynamic adjustment of display content based on the user's biometric information, analysis of emotions, and an optimized travel experience.
[0141] "User input information" refers to information that includes destinations the user wants to visit, activities of interest, and other personal preferences.
[0142] "Destination-related information" refers to information including geographical information, historical background, and data on tourist attractions related to the destination selected by the user.
[0143] "User device" refers to a smartphone, tablet, or dedicated device that a user uses to receive information and content and visualize the augmented reality environment.
[0144] An "augmented reality environment" refers to a technological environment that provides a richer and more interactive visual experience by overlaying digital information onto the physical reality environment.
[0145] "Biometric information" refers to data that includes a user's physiological and behavioral characteristics, such as facial expressions, tone of voice, and other physical responses.
[0146] "Analyzing and optimizing emotions" refers to the process of analyzing a user's biometric information and real-time emotional state, and then individually adjusting the user's experience based on the results.
[0147] "Collecting feedback and making improvements" refers to the process of gathering evaluations and comments from users after their experience and using that feedback to improve the system's content and functionality.
[0148] This invention is an augmented reality system designed to provide a personalized travel experience based on the user's emotions. Specific embodiments are described below.
[0149] The user begins a virtual journey using an application on their mobile device. First, the user selects a destination they wish to visit via the application and enters the necessary details. This information is sent from the user's device to the server. The server searches a database to retrieve various information related to the destination. This process uses web server software (e.g., generic name) and a database management system (e.g., generic name).
[0150] The acquired information is streamed to the user's device in real time, and the device uses this information to generate an augmented reality environment. The generation of augmented reality utilizes a corresponding AR platform (e.g., general name) and is visually presented to the user through smart glasses or compatible devices. This environment includes 3D models, guide information, and real-time information on local events.
[0151] Furthermore, the device is equipped with a built-in emotion engine. It collects the user's facial expressions and voice via the device's camera and microphone, and uses machine learning libraries (e.g., generic names) to analyze the user's emotions in real time. The emotion-based analysis results are used to adjust the augmented reality environment, optimizing the user experience.
[0152] For example, if a user selects "Virtual Times Square," the server provides relevant 3D models and information. If the user's emotions indicate excitement or surprise, the system adds more detailed information and interesting attractions.
[0153] After the experience ends, emotional data and feedback collected from users are stored on the server and used to improve future services. This allows for the creation of optimal travel experiences tailored to each user.
[0154] As an example of input to the generative AI model, prompt sentences such as "Explain how this system optimizes the augmented reality experience based on the user's emotions" can be used.
[0155] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0156] Step 1:
[0157] The user launches an application on their device, selects a destination they wish to visit and activities of interest, and enters detailed information. Based on this, the device sends destination information to the server in text format. The input data includes place names, planned activities, and the desired style of experience. The server receives the selected place names and associated user parameters as output.
[0158] Step 2:
[0159] The server searches the database based on the received destination information to retrieve relevant data. Using specific queries, it retrieves 3D model data and related tourist attractions and event information from the database management system. Input data consists of user selections sent to the server, while output data is content sent to the terminal in list format. Specifically, the server executes a backend program to extract information according to the user's preferences.
[0160] Step 3:
[0161] The terminal generates an augmented reality environment based on information received from the server. Using an AR-enabled device, 3D models and related information are integrated into the user's field of view. Using an AR development platform such as Unity, virtual objects are overlaid onto the real environment. This operation is performed by processing information packets sent from the server as input and providing the visualized AR content to the user as output.
[0162] Step 4:
[0163] The device activates a built-in emotion engine to collect the user's biometric information in real time. It analyzes facial expressions and voice tone captured from the camera and microphone, and uses machine learning algorithms to classify emotions. The input data consists of video and audio, and the output is the user's emotional state (e.g., surprise, joy). This process provides the capability to detect the user's reactions in real time.
[0164] Step 5:
[0165] The server dynamically adjusts the augmented reality content based on the output from the emotion engine. For example, if the user is surprised, it will present additional information or attractions that are more exciting. The input data is emotion information sent from the device, and the output data is the updated AR content. In this step, the server uses a generative AI model to create prompts and generate new information as content.
[0166] Step 6:
[0167] After the user experience ends, the device sends emotional data and feedback to the server. This information includes a log of the user's emotional changes and responses to a feedback form. As output, this data is stored in the server's records and used to improve future experiences. The data serves as crucial foundational information for improving the quality of the user experience.
[0168] (Application Example 2)
[0169] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".
[0170] Traditional virtual shopping systems have faced challenges in providing an optimal shopping experience tailored to individual users, as they often fail to adequately personalize the user experience based on their emotional state. Furthermore, the lack of real-time suggestions for relevant products based on user interests and responses has limited the user experience.
[0171] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0172] In this invention, the server includes means for acquiring digital information related to the destination based on user input information, means for transferring the acquired digital information to the user's information processing device, and means for analyzing the user's emotional state and adjusting the components within the augmented reality environment accordingly. This enables the provision of a personalized shopping experience that responds to the user's emotions and the real-time presentation of relevant product information based on the user's emotional analysis results.
[0173] "User input information" refers to information about destinations and interests that users provide to the system.
[0174] "Digital information" refers to data and content expressed in a format that can be processed by a computer.
[0175] An "information processing device" refers to an electronic device that has the function of receiving, processing, and displaying digital information.
[0176] An "augmented reality environment" refers to a virtual environment where digital information is overlaid onto images and information from the real world and displayed to the user.
[0177] "User behavior" refers to the physical actions of the user, including body movements, gestures, and gaze.
[0178] "Interaction" refers to the exchange of operations and responses that takes place between a system and a user.
[0179] "Emotional state" refers to the mental reactions and facial expressions a user displays in response to a particular stimulus.
[0180] "Analysis" refers to the process of understanding the content of collected data and information in detail.
[0181] "Components" refer to the individual elements that make up the functions, devices, and information parts of a system.
[0182] To realize this application, it is necessary to build a system involving three main elements: user, server, and terminal. First, the user wears an information processing device such as smart glasses or a headset to access the system. When the user inputs information about the destination they wish to visit or their interests, that information is sent to the server. Based on the user's input, the server retrieves digital information related to the destination from a database and transfers that digital information to the user's terminal.
[0183] The device utilizes received digital information to generate an augmented reality environment and displays this environment to the user through the information processing device to which the device is attached. Here, high-precision microphones and cameras are used to analyze the user's actions and emotions. For example, based on video data of the user's face acquired by the camera and audio data collected by the microphone, the user's emotional state is analyzed in real time using an emotion recognition API (e.g., Microsoft® Azure® Face API).
[0184] Based on these analysis results, the server dynamically adjusts the components within the augmented reality environment. If a user shows interest in a particular product, it can suggest related information and similar items to enhance the user's shopping experience. The software used includes a framework for augmented reality development (e.g., Unity).
[0185] For example, if a user becomes interested in a raincoat in a virtual shopping mall, the device that detects this interest will then present the user with related new collections and recommended items.
[0186] Furthermore, this system utilizes a generative AI model to improve services based on user responses. For example, input to the AI model can be provided through prompts like the following.
[0187] "Develop methods to predict how users will react to a particular product and provide them with the most relevant additional information."
[0188] "Analyze emotional data and propose the optimal algorithm for providing personalized experiences to users."
[0189] In this way, the present invention can provide a real-time, adaptive shopping experience based on the user's emotional state.
[0190] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0191] Step 1:
[0192] The user enters information about the destinations they wish to visit and their interests. This generates user input information, which is then sent from the user's device to the server.
[0193] Step 2:
[0194] The server retrieves digital information related to the destination from the database based on the received input information. The database stores information about pre-registered destinations. The retrieved digital information becomes the server's output, and this information is transferred to the user's terminal.
[0195] Step 3:
[0196] The terminal generates an augmented reality environment using digital information received from the server. This digital information includes visuals and content related to the destination. The generated augmented reality environment is the terminal's output, and this environment is presented to the user through an information processing device.
[0197] Step 4:
[0198] The user experiences an augmented reality environment through an information processing device. The user's movements, facial expressions, and voice are monitored by the device. As a result, user movement data is input into the terminal.
[0199] Step 5:
[0200] The device collects user behavior data using a high-precision microphone and camera, and analyzes this data using an emotion recognition API. The analyzed data generates an output indicating the user's emotional state. This output is used by the system to personalize the user experience.
[0201] Step 6:
[0202] The server dynamically adjusts the components of the augmented reality environment based on the user's emotional state. For example, if the server determines that the user is excited, it will present new event information or related products. The adjusted environment is the server's output, leading to an optimized user experience.
[0203] Step 7:
[0204] Using a generative AI model, information is generated based on user feedback and sentiment data to help further improve the service. This prompt message includes content such as, "Based on user responses, please consider what additional information should be provided." This output information will be used for future system improvements.
[0205] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0206] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include those described above. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions shown by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0207] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0208] [Second Embodiment]
[0209] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0210] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0211] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0212] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0213] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0214] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0215] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0216] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0217] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0218] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0219] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0220] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0221] This invention provides a system that enables users to experience travel from home through their own devices. In this system, the user selects a travel destination using an application. Information based on the selection is sent to a server, which retrieves the corresponding content from a database. The content includes high-resolution images, 3D models, and audio data of the destination. The server prepares this content for streaming to the user's device and begins the transfer.
[0222] The user's device processes the received data and generates an augmented reality environment through smart glasses or a headset. This environment is built based on the received data and provides realistic images in the user's field of view. When the user wears the device, they can enjoy a 360-degree virtual journey that responds to head movements and settings.
[0223] For example, if a user chooses to experience "Virtual Kyoto," the server prepares detailed 3D models and audio guides of Kyoto's famous landmarks. The user's device receives this information and uses AR technology to provide an experience of visiting various shrines and temples from the comfort of their home. During this process, the user can use hand movements and voice commands to ask for more detailed information or switch viewpoints.
[0224] Furthermore, users who have completed the experience can provide feedback on their satisfaction level and areas for improvement. This feedback is collected by the server and used to improve the system. In this way, the present invention enables diverse travel experiences that transcend physical limitations and provides new value to users.
[0225] The following describes the processing flow.
[0226] Step 1:
[0227] The user opens the application, selects their desired travel destination, and specifies related activities and preferred language. Once this information is entered, the device sends the request to the server.
[0228] Step 2:
[0229] The server parses the received request and retrieves content related to the destination from the database, such as high-resolution images, 3D models, and audio data. It then organizes this information appropriately and prepares it for data transfer.
[0230] Step 3:
[0231] The server begins streaming the organized content to the user's device. The transfer takes place in real time, and compression techniques are used as needed to optimize the process for smooth data reception.
[0232] Step 4:
[0233] The terminal receives data from the server and decompresses and decodes the received content. This prepares it for building the augmented reality environment. The terminal's processor performs the necessary processing to ensure the system operates smoothly.
[0234] Step 5:
[0235] The device generates an augmented reality environment through smart glasses or a headset. When the user wears the device, 3D objects are overlaid on their field of vision and displayed superimposed on the real world. This allows the user to feel as if they are right in front of their destination.
[0236] Step 6:
[0237] Users can interact within the augmented reality environment using voice commands and hand gestures. For example, they can select specific objects to display detailed information or converse with virtual guides.
[0238] Step 7:
[0239] Once a user finishes their experience, the device collects feedback from the user regarding their satisfaction level and areas for improvement, and sends it to the server. This allows information about system improvements to be accumulated on the server, enriching future experiences.
[0240] (Example 1)
[0241] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0242] There is a growing demand for virtual environments that allow users to experience the real world with greater realism without physical travel. However, current technology fails to deliver a satisfactory experience due to insufficient collection and real-time display of location-related data, as well as inadequate user interaction. Furthermore, it lacks the functionality to effectively utilize user feedback and continuously improve the virtual environment.
[0243] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0244] In this invention, the server includes means for searching for location-related data based on user selection information, means for transmitting the collected data to the user's information processing device, and means for constructing a virtual reality environment using the data received by the information processing device. This enables the user to have a virtual experience with a sense of presence.
[0245] A "user" refers to an individual or group that operates the system and experiences the virtual environment.
[0246] "Selection information" refers to information provided to users to determine the virtual location or environment they wish to visit.
[0247] "Location-related data" refers to a collection of information such as high-resolution images, 3D models, and audio guides about the location selected by the user.
[0248] An "information processing device" refers to an electronic device, such as a computer or smart device owned by a user, that has the function of receiving and processing data.
[0249] "Transmission" refers to the act or means of sending data from a server to an information processing device.
[0250] A "virtual reality environment" refers to a digitally constructed space that does not exist in the real world and can be experienced through user interaction.
[0251] "Interaction" refers to the two-way exchanges, such as operations and responses, that take place between the user and the virtual reality environment.
[0252] "Opinions" refer to feedback from users regarding their experiences and suggestions for improvement.
[0253] A "storage device" is an electronic storage medium used to store data so that it can be retrieved later.
[0254] "Eye movement" refers to actions related to changes in a user's visual focus or eye movements.
[0255] "Instant adjustment" refers to changing the displayed content of the virtual reality environment in real time in response to the user's actions and requests.
[0256] This invention relates to a system that realizes a virtual reality environment using a user's computer or smart device (information processing device). The user installs an application on their device and launches it to use the system. Based on the user's selection information, the user determines the virtual location they wish to visit. For example, the user can choose to experience "Virtual Kyoto."
[0257] The server retrieves data related to the selected location from its internal data store based on the customer's selection. This data includes high-resolution images, 3D models, and audio guides. The server then transmits this data to the user's device. This transmission requires efficient data transfer and typically uses real-time communication protocols over the internet.
[0258] The device receives streamed data and uses AR technologies such as Unity and Unreal Engine to construct a virtual reality environment. This constructed environment is displayed and experienced by the user through smart glasses or a headset. The device instantly adjusts the display within the virtual environment in response to the user's gaze and movements. This dynamic interaction enhances immersion and provides a more realistic experience for the user.
[0259] Users can manipulate the environment using hand movements and voice commands to acquire information and switch viewpoints. User feedback is collected by computer after the experience ends, sent to a server, and stored in memory. This feedback data is used to improve the virtual environment.
[0260] As an example of a generated AI model prompt, "I want to visualize a 3D model of Kyoto and play an audio guide about the temples" provides parameters useful for data retrieval and display. Users can have diverse and immersive travel experiences that go beyond physical limitations.
[0261] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0262] Step 1:
[0263] The user launches an application on their device and selects a virtual location they wish to visit. As input, the user provides destination selection information. This selection information serves as a basis for the system to identify the location data needed in the next step. As output, the selection information is sent to the server.
[0264] Step 2:
[0265] The server searches the data store for data related to the corresponding location based on the selection information it receives. It receives user selection information as input. Data processing involves executing database queries to extract datasets containing high-resolution images, 3D models, audio guides, etc. The collected data is then prepared on the server as output.
[0266] Step 3:
[0267] The server prepares the collected data for transmission to the user's terminal. The retrieved data is used as input. Data processing involves compression and encoding, converting the data into an optimized format for efficient data transfer. As output, the data is ready and transmission begins.
[0268] Step 4:
[0269] The terminal receives data from the server and constructs the virtual reality environment. It receives transmitted images, 3D models, and audio data as input. It uses Unity or Unreal Engine for data processing to generate the virtual reality environment. The generated AR environment is displayed on the display device as output.
[0270] Step 5:
[0271] Users experience the AR environment through smart glasses or a headset. They manipulate the environment through their gaze and body movements, and interact with it through voice and motion. The device receives user gaze and motion information as input, and adjusts the virtual environment display in real time based on this information. The output is an interactive experience that responds to movement.
[0272] Step 6:
[0273] After a user completes an experience, feedback is provided within the system. As input, user opinions and satisfaction data are entered into the device. This data is sent to a server and stored in memory. As output, the feedback data is used for improvement.
[0274] (Application Example 1)
[0275] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0276] Conventional travel experience systems make it difficult to realistically experience diverse travel destinations from the comfort of one's home. Furthermore, they lack features that allow for interaction that reflects user instructions and effectively incorporate post-experience feedback into the system. This makes it difficult to provide truly immersive remote experiences, and also hinders the continuous improvement of the system based on user input.
[0277] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0278] In this invention, the server includes means for acquiring a set of information related to a destination based on user input information, means for transferring the acquired set of information to the user's information processing device, means for generating an augmented reality environment using the information received by the information processing device, and means for operating the functions of the information processing device in response to the user's voice instructions using voice recognition technology. This makes it possible for the user to realistically experience a variety of travel destinations from the comfort of their home, and also enables flexible operation of the augmented reality environment based on the user's voice input and improvement of the system based on the feedback received.
[0279] "User input information" refers to instructions and information related to the travel destination or experience selected by the user through the information processing device.
[0280] The "information group related to the destination" is an aggregate of digital information including geographical information, visual data, audio guides, etc. at the selected travel destination.
[0281] The "information processing device" is a terminal used by the user to perform a virtual reality experience, referring to devices such as smartphones and headsets.
[0282] The "means for transferring" is a process or technology for transmitting data from a server to an information processing device.
[0283] The "augmented reality environment" is an environment of computer-generated images for overlaying digital information on actual environmental information for display.
[0284] The "speech recognition technology" is a technology for interpreting the speech uttered by the user and enabling the computer to react appropriately based on it.
[0285] The "voice instruction" is an instruction or request made by the user to the system or device through voice.
[0286] The system for implementing this invention begins by installing an application with a user interface for the user to select a travel destination on the information processing device. When the user selects a travel destination, the information is sent to the server. The server retrieves the information group related to the destination selected by the user from the database and transfers the information including geographical information, visual data, audio guides, etc. constituting the information group to the information processing device.
[0287] The information processing device generates an augmented reality environment using the information group received from the server. By using software such as ARKit (iOS) or ARCore (Android), the user can visually experience the digital information superimposed on the real world. For example, when the user selects "Virtual Paris Travel", detailed 3D models of the Eiffel Tower and the Louvre Museum are distributed and displayed in the augmented reality environment.
[0288] Furthermore, the information processing device uses speech recognition technology (such as Google Speech-to-Text) to interpret the user's voice commands. This allows the user to interact with and manipulate the environment through voice. For example, if the user asks, "What is the height of the Eiffel Tower?", the information will be provided immediately.
[0289] This system allows users to experience visiting tourist destinations around the world from the comfort of their homes, and to provide feedback based on their experiences, contributing to system improvements. By inputting prompts such as, "Provide detailed information about the travel destination chosen by the user as a virtual experience with a visually represented 3D model and audio support," into the generating AI model, a more realistic and immersive experience can be achieved.
[0290] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0291] Step 1:
[0292] The user selects a travel destination through the application interface of the information processing device. The input is the name of the travel destination selected by the user. Based on the user's selection, the information is sent to the server.
[0293] Step 2:
[0294] The server receives the name of the travel destination sent by the user, and retrieves a set of information related to the destination, such as visual data, geographical information, and audio guides, by referring to the database. The input is the travel destination selected by the user, and the output is a set of various information related to that travel destination.
[0295] Step 3:
[0296] The server encodes the acquired information and streams or downloads it to the information processing device. The input is the information acquired by the server, and the output is the data transferred to the information processing device.
[0297] Step 4:
[0298] The information processing device uses ARKit or ARCore to generate an augmented reality environment based on the information received from the server. The input is the received information, and the output is the augmented reality environment presented to the user.
[0299] Step 5:
[0300] Users use voice commands to ask questions of or operate the information processing device. The input is the voice commands spoken by the user.
[0301] Step 6:
[0302] The information processing device uses speech recognition technology to convert user voice commands into text data and displays appropriate information in an augmented reality environment based on that input. The input is the user's voice command, and the output is the corresponding information or operation result.
[0303] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0304] This invention provides a personalized experience tailored to the user's emotional state by combining an emotion engine with a travel experience system using an augmented reality environment. This system operates by integrating the user's terminal, server, and emotion engine.
[0305] First, the user selects the destination they want to visit via the application and enters the relevant information. When this information is sent from the terminal to the server, the server retrieves the content related to the destination from the database and streams it to the user's terminal.
[0306] The user's terminal generates an augmented reality environment based on the received content and provides visual information to the user through smart glasses or a headset. Furthermore, the terminal uses the built-in emotion engine to infer the emotion from the user's expression, voice tone, and other biometric information.
[0307] The emotion engine analyzes the user's emotion in real time and adjusts the content and interactions within the augmented reality environment based on this. For example, when the user is relaxed, it provides a calm environment, and when the user is excited, it provides a more stimulating experience.
[0308] As a specific example, assume the user selects to experience "Virtual New York". The server retrieves the 3D model of Times Square and event information and streams it. The terminal constructs this as an augmented reality environment and provides it to the user's field of vision. The emotion engine analyzes the user's reaction, and if the user is feeling surprised or happy, it introduces more enjoyable attractions or displays information according to the interest.
[0309] After the experience, the emotion data and feedback collected from the user are sent to the server and used to improve future experiences. Thereby, a travel experience optimized for individual users is provided. This invention contributes to enabling dynamic analysis of emotions and further personalizing the user's experience.
[0310] The following explains the processing flow.
[0311] Step 1:
[0312] The user launches the application and selects a destination they wish to visit. They then enter information such as related activities and preferred language, and send this information to the server via their device.
[0313] Step 2:
[0314] The server analyzes the received request and retrieves high-resolution images, 3D models, and audio data related to the destination from the database. It then packages this data and prepares it for streaming to the user's device.
[0315] Step 3:
[0316] The device receives content sent from the server. It then decompresses and decodes the content, preparing to generate the augmented reality environment.
[0317] Step 4:
[0318] The device's built-in emotion engine collects user facial expression data and voice tone. Based on this, it analyzes the user's emotional state in real time.
[0319] Step 5:
[0320] The device generates an augmented reality environment and displays it in the user's field of view using smart glasses or a headset. The content and audio guidance presented are dynamically adjusted according to the user's emotional state.
[0321] Step 6:
[0322] Users enjoy an interactive experience based on analysis by an emotion engine. For example, if a user is relaxed, calming music and scenery are emphasized, while if they are excited, active activities are presented.
[0323] Step 7:
[0324] When a user finishes an experience, the device sends emotional data and user feedback to the server. This information is stored in a database for system improvement and further personalization.
[0325] (Example 2)
[0326] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0327] Traditional augmented reality systems have struggled to dynamically adjust content based on user emotions and individual actions, making it difficult to provide an experience optimized for each user. Therefore, there is a growing need to provide travel experiences that are more adaptable to users' emotions and interests.
[0328] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0329] In this invention, the server includes means for acquiring destination-related information based on user input information, means for transferring the acquired information to the user's device, and means for generating an augmented reality environment using the information received by the device. This enables dynamic adjustment of display content based on the user's biometric information, analysis of emotions, and an optimized travel experience.
[0330] "User input information" refers to information that includes destinations the user wants to visit, activities of interest, and other personal preferences.
[0331] "Destination-related information" refers to information including geographical information, historical background, and data on tourist attractions related to the destination selected by the user.
[0332] "User device" refers to a smartphone, tablet, or dedicated device that a user uses to receive information and content and visualize the augmented reality environment.
[0333] An "augmented reality environment" refers to a technological environment that provides a richer and more interactive visual experience by overlaying digital information onto the physical reality environment.
[0334] "Biometric information" refers to data that includes a user's physiological and behavioral characteristics, such as facial expressions, tone of voice, and other physical responses.
[0335] "Analyzing and optimizing emotions" refers to the process of analyzing a user's biometric information and real-time emotional state, and then individually adjusting the user's experience based on the results.
[0336] "Collecting feedback and making improvements" refers to the process of gathering evaluations and comments from users after their experience and using that feedback to improve the system's content and functionality.
[0337] This invention is an augmented reality system designed to provide a personalized travel experience based on the user's emotions. Specific embodiments are described below.
[0338] The user begins a virtual journey using an application on their mobile device. First, the user selects a destination they wish to visit via the application and enters the necessary details. This information is sent from the user's device to the server. The server searches a database to retrieve various information related to the destination. This process uses web server software (e.g., generic name) and a database management system (e.g., generic name).
[0339] The acquired information is streamed to the user's device in real time, and the device uses this information to generate an augmented reality environment. The generation of augmented reality utilizes a corresponding AR platform (e.g., general name) and is visually presented to the user through smart glasses or compatible devices. This environment includes 3D models, guide information, and real-time information on local events.
[0340] Furthermore, the device is equipped with a built-in emotion engine. It collects the user's facial expressions and voice via the device's camera and microphone, and uses machine learning libraries (e.g., generic names) to analyze the user's emotions in real time. The emotion-based analysis results are used to adjust the augmented reality environment, optimizing the user experience.
[0341] For example, if a user selects "Virtual Times Square," the server provides relevant 3D models and information. If the user's emotions indicate excitement or surprise, the system adds more detailed information and interesting attractions.
[0342] After the experience ends, emotional data and feedback collected from users are stored on the server and used to improve future services. This allows for the creation of optimal travel experiences tailored to each user.
[0343] As an example of input to the generative AI model, prompt sentences such as "Explain how this system optimizes the augmented reality experience based on the user's emotions" can be used.
[0344] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0345] Step 1:
[0346] The user launches an application on their device, selects a destination they wish to visit and activities of interest, and enters detailed information. Based on this, the device sends destination information to the server in text format. The input data includes place names, planned activities, and the desired style of experience. The server receives the selected place names and associated user parameters as output.
[0347] Step 2:
[0348] The server searches the database based on the received destination information to retrieve relevant data. Using specific queries, it retrieves 3D model data and related tourist attractions and event information from the database management system. Input data consists of user selections sent to the server, while output data is content sent to the terminal in list format. Specifically, the server executes a backend program to extract information according to the user's preferences.
[0349] Step 3:
[0350] The terminal generates an augmented reality environment based on information received from the server. Using an AR-enabled device, 3D models and related information are integrated into the user's field of view. Using an AR development platform such as Unity, virtual objects are overlaid onto the real environment. This operation is performed by processing information packets sent from the server as input and providing the visualized AR content to the user as output.
[0351] Step 4:
[0352] The device activates a built-in emotion engine to collect the user's biometric information in real time. It analyzes facial expressions and voice tone captured from the camera and microphone, and uses machine learning algorithms to classify emotions. The input data consists of video and audio, and the output is the user's emotional state (e.g., surprise, joy). This process provides the capability to detect the user's reactions in real time.
[0353] Step 5:
[0354] The server dynamically adjusts the augmented reality content based on the output from the emotion engine. For example, if the user is surprised, it will present additional information or attractions that are more exciting. The input data is emotion information sent from the device, and the output data is the updated AR content. In this step, the server uses a generative AI model to create prompts and generate new information as content.
[0355] Step 6:
[0356] After the user experience ends, the device sends emotional data and feedback to the server. This information includes a log of the user's emotional changes and responses to a feedback form. As output, this data is stored in the server's records and used to improve future experiences. The data serves as crucial foundational information for improving the quality of the user experience.
[0357] (Application Example 2)
[0358] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0359] Traditional virtual shopping systems have faced challenges in providing an optimal shopping experience tailored to individual users, as they often fail to adequately personalize the user experience based on their emotional state. Furthermore, the lack of real-time suggestions for relevant products based on user interests and responses has limited the user experience.
[0360] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0361] In this invention, the server includes means for acquiring digital information related to the destination based on user input information, means for transferring the acquired digital information to the user's information processing device, and means for analyzing the user's emotional state and adjusting the components within the augmented reality environment accordingly. This enables the provision of a personalized shopping experience that responds to the user's emotions and the real-time presentation of relevant product information based on the user's emotional analysis results.
[0362] "User input information" refers to information about destinations and interests that users provide to the system.
[0363] "Digital information" refers to data and content expressed in a format that can be processed by a computer.
[0364] An "information processing device" refers to an electronic device that has the function of receiving, processing, and displaying digital information.
[0365] An "augmented reality environment" refers to a virtual environment where digital information is overlaid onto images and information from the real world and displayed to the user.
[0366] "User behavior" refers to the physical actions of the user, including body movements, gestures, and gaze.
[0367] "Interaction" refers to the exchange of operations and responses that takes place between a system and a user.
[0368] "Emotional state" refers to the mental reactions and facial expressions a user displays in response to a particular stimulus.
[0369] "Analysis" refers to the process of understanding the content of collected data and information in detail.
[0370] "Components" refer to the individual elements that make up the functions, devices, and information parts of a system.
[0371] To realize this application, it is necessary to build a system involving three main elements: user, server, and terminal. First, the user wears an information processing device such as smart glasses or a headset to access the system. When the user inputs information about the destination they wish to visit or their interests, that information is sent to the server. Based on the user's input, the server retrieves digital information related to the destination from a database and transfers that digital information to the user's terminal.
[0372] The device utilizes received digital information to generate an augmented reality environment and displays this environment to the user through the information processing device to which the device is attached. Here, high-precision microphones and cameras are used to analyze the user's actions and emotions. For example, based on video data of the user's face acquired by the camera and audio data collected by the microphone, an emotion recognition API (e.g., Microsoft Azure Face API) is used to analyze the user's emotional state in real time.
[0373] Based on these analysis results, the server dynamically adjusts the components within the augmented reality environment. If a user shows interest in a particular product, it can suggest related information and similar items to enhance the user's shopping experience. The software used includes a framework for augmented reality development (e.g., Unity).
[0374] For example, if a user becomes interested in a raincoat in a virtual shopping mall, the device that detects this interest will then present the user with related new collections and recommended items.
[0375] Furthermore, this system utilizes a generative AI model to improve services based on user responses. For example, input to the AI model can be provided through prompts like the following.
[0376] "Develop methods to predict how users will react to a particular product and provide them with the most relevant additional information."
[0377] "Analyze emotional data and propose the optimal algorithm for providing personalized experiences to users."
[0378] In this way, the present invention can provide a real-time, adaptive shopping experience based on the user's emotional state.
[0379] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0380] Step 1:
[0381] The user enters information about the destinations they wish to visit and their interests. This generates user input information, which is then sent from the user's device to the server.
[0382] Step 2:
[0383] The server retrieves digital information related to the destination from the database based on the received input information. The database stores information about pre-registered destinations. The retrieved digital information becomes the server's output, and this information is transferred to the user's terminal.
[0384] Step 3:
[0385] The terminal generates an augmented reality environment using digital information received from the server. This digital information includes visuals and content related to the destination. The generated augmented reality environment is the terminal's output, and this environment is presented to the user through an information processing device.
[0386] Step 4:
[0387] The user experiences an augmented reality environment through an information processing device. The user's movements, facial expressions, and voice are monitored by the device. As a result, user movement data is input into the terminal.
[0388] Step 5:
[0389] The device collects user behavior data using a high-precision microphone and camera, and analyzes this data using an emotion recognition API. The analyzed data generates an output indicating the user's emotional state. This output is used by the system to personalize the user experience.
[0390] Step 6:
[0391] The server dynamically adjusts the components of the augmented reality environment based on the user's emotional state. For example, if the server determines that the user is excited, it will present new event information or related products. The adjusted environment is the server's output, leading to an optimized user experience.
[0392] Step 7:
[0393] Using a generative AI model, information is generated based on user feedback and sentiment data to help further improve the service. This prompt message includes content such as, "Based on user responses, please consider what additional information should be provided." This output information will be used for future system improvements.
[0394] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0395] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include those described above. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions shown by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0396] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0397] [Third Embodiment]
[0398] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0399] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0400] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0401] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0402] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0403] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0404] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0405] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0406] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0407] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0408] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0409] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0410] This invention provides a system that enables users to experience travel from home through their own devices. In this system, the user selects a travel destination using an application. Information based on the selection is sent to a server, which retrieves the corresponding content from a database. The content includes high-resolution images, 3D models, and audio data of the destination. The server prepares this content for streaming to the user's device and begins the transfer.
[0411] The user's device processes the received data and generates an augmented reality environment through smart glasses or a headset. This environment is built based on the received data and provides realistic images in the user's field of view. When the user wears the device, they can enjoy a 360-degree virtual journey that responds to head movements and settings.
[0412] For example, if a user chooses to experience "Virtual Kyoto," the server prepares detailed 3D models and audio guides of Kyoto's famous landmarks. The user's device receives this information and uses AR technology to provide an experience of visiting various shrines and temples from the comfort of their home. During this process, the user can use hand movements and voice commands to ask for more detailed information or switch viewpoints.
[0413] Furthermore, users who have completed the experience can provide feedback on their satisfaction level and areas for improvement. This feedback is collected by the server and used to improve the system. In this way, the present invention enables diverse travel experiences that transcend physical limitations and provides new value to users.
[0414] The following describes the processing flow.
[0415] Step 1:
[0416] The user opens the application, selects their desired travel destination, and specifies related activities and preferred language. Once this information is entered, the device sends the request to the server.
[0417] Step 2:
[0418] The server parses the received request and retrieves content related to the destination from the database, such as high-resolution images, 3D models, and audio data. It then organizes this information appropriately and prepares it for data transfer.
[0419] Step 3:
[0420] The server begins streaming the organized content to the user's device. The transfer takes place in real time, and compression techniques are used as needed to optimize the process for smooth data reception.
[0421] Step 4:
[0422] The terminal receives data from the server and decompresses and decodes the received content. This prepares it for building the augmented reality environment. The terminal's processor performs the necessary processing to ensure the system operates smoothly.
[0423] Step 5:
[0424] The device generates an augmented reality environment through smart glasses or a headset. When the user wears the device, 3D objects are overlaid on their field of vision and displayed superimposed on the real world. This allows the user to feel as if they are right in front of their destination.
[0425] Step 6:
[0426] Users can interact within the augmented reality environment using voice commands and hand gestures. For example, they can select specific objects to display detailed information or converse with virtual guides.
[0427] Step 7:
[0428] Once a user finishes their experience, the device collects feedback from the user regarding their satisfaction level and areas for improvement, and sends it to the server. This allows information about system improvements to be accumulated on the server, enriching future experiences.
[0429] (Example 1)
[0430] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0431] There is a growing demand for virtual environments that allow users to experience the real world with greater realism without physical travel. However, current technology fails to deliver a satisfactory experience due to insufficient collection and real-time display of location-related data, as well as inadequate user interaction. Furthermore, it lacks the functionality to effectively utilize user feedback and continuously improve the virtual environment.
[0432] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0433] In this invention, the server includes means for searching for location-related data based on user selection information, means for transmitting the collected data to the user's information processing device, and means for constructing a virtual reality environment using the data received by the information processing device. This enables the user to have a virtual experience with a sense of presence.
[0434] A "user" refers to an individual or group that operates the system and experiences the virtual environment.
[0435] "Selection information" refers to information provided to users to determine the virtual location or environment they wish to visit.
[0436] "Location-related data" refers to a collection of information such as high-resolution images, 3D models, and audio guides about the location selected by the user.
[0437] An "information processing device" refers to an electronic device, such as a computer or smart device owned by a user, that has the function of receiving and processing data.
[0438] "Transmission" refers to the act or means of sending data from a server to an information processing device.
[0439] A "virtual reality environment" refers to a digitally constructed space that does not exist in the real world and can be experienced through user interaction.
[0440] "Interaction" refers to the two-way exchanges, such as operations and responses, that take place between the user and the virtual reality environment.
[0441] "Opinions" refer to feedback from users regarding their experiences and suggestions for improvement.
[0442] A "storage device" is an electronic storage medium used to store data so that it can be retrieved later.
[0443] "Eye movement" refers to actions related to changes in a user's visual focus or eye movements.
[0444] "Instant adjustment" refers to changing the displayed content of the virtual reality environment in real time in response to the user's actions and requests.
[0445] This invention relates to a system that realizes a virtual reality environment using a user's computer or smart device (information processing device). The user installs an application on their device and launches it to use the system. Based on the user's selection information, the user determines the virtual location they wish to visit. For example, the user can choose to experience "Virtual Kyoto."
[0446] The server retrieves data related to the selected location from its internal data store based on the customer's selection. This data includes high-resolution images, 3D models, and audio guides. The server then transmits this data to the user's device. This transmission requires efficient data transfer and typically uses real-time communication protocols over the internet.
[0447] The device receives streamed data and uses AR technologies such as Unity and Unreal Engine to construct a virtual reality environment. This constructed environment is displayed and experienced by the user through smart glasses or a headset. The device instantly adjusts the display within the virtual environment in response to the user's gaze and movements. This dynamic interaction enhances immersion and provides a more realistic experience for the user.
[0448] Users can manipulate the environment using hand movements and voice commands to acquire information and switch viewpoints. User feedback is collected by computer after the experience ends, sent to a server, and stored in memory. This feedback data is used to improve the virtual environment.
[0449] As an example of a generated AI model prompt, "I want to visualize a 3D model of Kyoto and play an audio guide about the temples" provides parameters useful for data retrieval and display. Users can have diverse and immersive travel experiences that go beyond physical limitations.
[0450] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0451] Step 1:
[0452] The user launches an application on their device and selects a virtual location they wish to visit. As input, the user provides destination selection information. This selection information serves as a basis for the system to identify the location data needed in the next step. As output, the selection information is sent to the server.
[0453] Step 2:
[0454] The server searches the data store for data related to the corresponding location based on the selection information it receives. It receives user selection information as input. Data processing involves executing database queries to extract datasets containing high-resolution images, 3D models, audio guides, etc. The collected data is then prepared on the server as output.
[0455] Step 3:
[0456] The server prepares the collected data for transmission to the user's terminal. The retrieved data is used as input. Data processing involves compression and encoding, converting the data into an optimized format for efficient data transfer. As output, the data is ready and transmission begins.
[0457] Step 4:
[0458] The terminal receives data from the server and constructs the virtual reality environment. It receives transmitted images, 3D models, and audio data as input. It uses Unity or Unreal Engine for data processing to generate the virtual reality environment. The generated AR environment is displayed on the display device as output.
[0459] Step 5:
[0460] Users experience the AR environment through smart glasses or a headset. They manipulate the environment through their gaze and body movements, and interact with it through voice and motion. The device receives user gaze and motion information as input, and adjusts the virtual environment display in real time based on this information. The output is an interactive experience that responds to movement.
[0461] Step 6:
[0462] After a user completes an experience, feedback is provided within the system. As input, user opinions and satisfaction data are entered into the device. This data is sent to a server and stored in memory. As output, the feedback data is used for improvement.
[0463] (Application Example 1)
[0464] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0465] Conventional travel experience systems make it difficult to realistically experience diverse travel destinations from the comfort of one's home. Furthermore, they lack features that allow for interaction that reflects user instructions and effectively incorporate post-experience feedback into the system. This makes it difficult to provide truly immersive remote experiences, and also hinders the continuous improvement of the system based on user input.
[0466] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0467] In this invention, the server includes means for acquiring a set of information related to a destination based on user input information, means for transferring the acquired set of information to the user's information processing device, means for generating an augmented reality environment using the information received by the information processing device, and means for operating the functions of the information processing device in response to the user's voice instructions using voice recognition technology. This makes it possible for the user to realistically experience a variety of travel destinations from the comfort of their home, and also enables flexible operation of the augmented reality environment based on the user's voice input and improvement of the system based on the feedback received.
[0468] "User input information" refers to instructions and information related to the travel destination or experience selected by the user through the information processing device.
[0469] "Destination-related information" refers to a collection of digital information, including geographical information, visual data, and audio guides for the selected travel destination.
[0470] An "information processing device" is a terminal used by users to experience virtual reality, and refers to devices such as smartphones and headsets.
[0471] "Transferring means" refers to the processes and technologies used to send data from a server to an information processing device.
[0472] An "augmented reality environment" is a computer-generated image environment that overlays digital information onto real-world environmental information for display.
[0473] "Speech recognition technology" is a technology that interprets the voice spoken by a user and allows a computer to respond appropriately based on that interpretation.
[0474] "Voice commands" are commands or requests that a user makes to a system or device using their voice.
[0475] A system implementing this invention begins by installing an application equipped with a user interface for the user to select a travel destination on an information processing device. When the user selects a travel destination, that information is sent to a server. The server retrieves a set of information related to the destination selected by the user from a database and transfers the information, including geographic information, visual data, and audio guides, that constitutes that set of information to the information processing device.
[0476] The information processing device generates an augmented reality environment using the information received from the server. Using software such as ARKit (iOS) and ARCore (Android), users can visually experience digital information superimposed on the real world. For example, if a user selects "Virtual Paris Trip," detailed 3D models of the Eiffel Tower and the Louvre Museum are delivered and displayed in the augmented reality environment.
[0477] Furthermore, the information processing device uses speech recognition technology (such as Google Speech-to-Text) to interpret the user's voice commands. This allows the user to interact with and manipulate the environment through voice. For example, if the user asks, "What is the height of the Eiffel Tower?", the information will be provided immediately.
[0478] This system allows users to experience visiting tourist destinations around the world from the comfort of their homes, and to provide feedback based on their experiences, contributing to system improvements. By inputting prompts such as, "Provide detailed information about the travel destination chosen by the user as a virtual experience with a visually represented 3D model and audio support," into the generating AI model, a more realistic and immersive experience can be achieved.
[0479] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0480] Step 1:
[0481] The user selects a travel destination through the application interface of the information processing device. The input is the name of the travel destination selected by the user. Based on the user's selection, the information is sent to the server.
[0482] Step 2:
[0483] The server receives the name of the travel destination sent by the user, and retrieves a set of information related to the destination, such as visual data, geographical information, and audio guides, by referring to the database. The input is the travel destination selected by the user, and the output is a set of various information related to that travel destination.
[0484] Step 3:
[0485] The server encodes the acquired information and streams or downloads it to the information processing device. The input is the information acquired by the server, and the output is the data transferred to the information processing device.
[0486] Step 4:
[0487] The information processing device uses ARKit or ARCore to generate an augmented reality environment based on the information received from the server. The input is the received information, and the output is the augmented reality environment presented to the user.
[0488] Step 5:
[0489] Users use voice commands to ask questions of or operate the information processing device. The input is the voice commands spoken by the user.
[0490] Step 6:
[0491] The information processing device uses speech recognition technology to convert user voice commands into text data and displays appropriate information in an augmented reality environment based on that input. The input is the user's voice command, and the output is the corresponding information or operation result.
[0492] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0493] This invention provides a personalized experience tailored to the user's emotional state by combining an emotion engine with a travel experience system using an augmented reality environment. This system operates by integrating the user's terminal, server, and emotion engine.
[0494] First, the user selects a destination they wish to visit through the application and enters relevant information. Once this information is sent to the server via the device, the server retrieves content related to the destination from its database and streams it to the user's device.
[0495] The user's device generates an augmented reality environment based on the received content and provides visual information to the user through smart glasses or a headset. Furthermore, the device uses a built-in emotion engine to infer emotions from the user's facial expressions, tone of voice, and other biometric information.
[0496] The emotion engine analyzes the user's emotions in real time and adjusts the content and interactions within the augmented reality environment based on this analysis. For example, if the user is relaxed, it provides a calm environment; if they are excited, it provides a more stimulating experience.
[0497] As a concrete example, suppose a user chooses to experience "Virtual New York." The server retrieves a 3D model of Times Square and event information and streams it. The device then constructs this as an augmented reality environment and presents it to the user's field of view. The emotion engine analyzes the user's reactions, and if the user is surprised or delighted, it suggests more fun attractions or displays information tailored to their interests.
[0498] After the experience ends, emotional data and feedback collected from users are sent to a server and used to improve future experiences. This allows for the provision of travel experiences optimized for each individual user. This invention enables dynamic analysis of emotions and contributes to further personalizing the user experience.
[0499] The following describes the processing flow.
[0500] Step 1:
[0501] The user launches the application and selects a destination they wish to visit. They then enter information such as related activities and preferred language, and send this information to the server via their device.
[0502] Step 2:
[0503] The server analyzes the received request and retrieves high-resolution images, 3D models, and audio data related to the destination from the database. It then packages this data and prepares it for streaming to the user's device.
[0504] Step 3:
[0505] The device receives content sent from the server. It then decompresses and decodes the content, preparing to generate the augmented reality environment.
[0506] Step 4:
[0507] The device's built-in emotion engine collects user facial expression data and voice tone. Based on this, it analyzes the user's emotional state in real time.
[0508] Step 5:
[0509] The device generates an augmented reality environment and displays it in the user's field of view using smart glasses or a headset. The content and audio guidance presented are dynamically adjusted according to the user's emotional state.
[0510] Step 6:
[0511] Users enjoy an interactive experience based on analysis by an emotion engine. For example, if a user is relaxed, calming music and scenery are emphasized, while if they are excited, active activities are presented.
[0512] Step 7:
[0513] When a user finishes an experience, the device sends emotional data and user feedback to the server. This information is stored in a database for system improvement and further personalization.
[0514] (Example 2)
[0515] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0516] Traditional augmented reality systems have struggled to dynamically adjust content based on user emotions and individual actions, making it difficult to provide an experience optimized for each user. Therefore, there is a growing need to provide travel experiences that are more adaptable to users' emotions and interests.
[0517] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0518] In this invention, the server includes means for acquiring destination-related information based on user input information, means for transferring the acquired information to the user's device, and means for generating an augmented reality environment using the information received by the device. This enables dynamic adjustment of display content based on the user's biometric information, analysis of emotions, and an optimized travel experience.
[0519] "User input information" refers to information that includes destinations the user wants to visit, activities of interest, and other personal preferences.
[0520] "Destination-related information" refers to information including geographical information, historical background, and data on tourist attractions related to the destination selected by the user.
[0521] "User device" refers to a smartphone, tablet, or dedicated device that a user uses to receive information and content and visualize the augmented reality environment.
[0522] An "augmented reality environment" refers to a technological environment that provides a richer and more interactive visual experience by overlaying digital information onto the physical reality environment.
[0523] "Biometric information" refers to data that includes a user's physiological and behavioral characteristics, such as facial expressions, tone of voice, and other physical responses.
[0524] "Analyzing and optimizing emotions" refers to the process of analyzing a user's biometric information and real-time emotional state, and then individually adjusting the user's experience based on the results.
[0525] "Collecting feedback and making improvements" refers to the process of gathering evaluations and comments from users after their experience and using that feedback to improve the system's content and functionality.
[0526] This invention is an augmented reality system designed to provide a personalized travel experience based on the user's emotions. Specific embodiments are described below.
[0527] The user begins a virtual journey using an application on their mobile device. First, the user selects a destination they wish to visit via the application and enters the necessary details. This information is sent from the user's device to the server. The server searches a database to retrieve various information related to the destination. This process uses web server software (e.g., generic name) and a database management system (e.g., generic name).
[0528] The acquired information is streamed to the user's device in real time, and the device uses this information to generate an augmented reality environment. The generation of augmented reality utilizes a corresponding AR platform (e.g., general name) and is visually presented to the user through smart glasses or compatible devices. This environment includes 3D models, guide information, and real-time information on local events.
[0529] Furthermore, the device is equipped with a built-in emotion engine. It collects the user's facial expressions and voice via the device's camera and microphone, and uses machine learning libraries (e.g., generic names) to analyze the user's emotions in real time. The emotion-based analysis results are used to adjust the augmented reality environment, optimizing the user experience.
[0530] For example, if a user selects "Virtual Times Square," the server provides relevant 3D models and information. If the user's emotions indicate excitement or surprise, the system adds more detailed information and interesting attractions.
[0531] After the experience ends, emotional data and feedback collected from users are stored on the server and used to improve future services. This allows for the creation of optimal travel experiences tailored to each user.
[0532] As an example of input to the generative AI model, prompt sentences such as "Explain how this system optimizes the augmented reality experience based on the user's emotions" can be used.
[0533] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0534] Step 1:
[0535] The user launches an application on their device, selects a destination they wish to visit and activities of interest, and enters detailed information. Based on this, the device sends destination information to the server in text format. The input data includes place names, planned activities, and the desired style of experience. The server receives the selected place names and associated user parameters as output.
[0536] Step 2:
[0537] The server searches the database based on the received destination information to retrieve relevant data. Using specific queries, it retrieves 3D model data and related tourist attractions and event information from the database management system. Input data consists of user selections sent to the server, while output data is content sent to the terminal in list format. Specifically, the server executes a backend program to extract information according to the user's preferences.
[0538] Step 3:
[0539] The terminal generates an augmented reality environment based on information received from the server. Using an AR-enabled device, 3D models and related information are integrated into the user's field of view. Using an AR development platform such as Unity, virtual objects are overlaid onto the real environment. This operation is performed by processing information packets sent from the server as input and providing the visualized AR content to the user as output.
[0540] Step 4:
[0541] The device activates a built-in emotion engine to collect the user's biometric information in real time. It analyzes facial expressions and voice tone captured from the camera and microphone, and uses machine learning algorithms to classify emotions. The input data consists of video and audio, and the output is the user's emotional state (e.g., surprise, joy). This process provides the capability to detect the user's reactions in real time.
[0542] Step 5:
[0543] The server dynamically adjusts the augmented reality content based on the output from the emotion engine. For example, if the user is surprised, it will present additional information or attractions that are more exciting. The input data is emotion information sent from the device, and the output data is the updated AR content. In this step, the server uses a generative AI model to create prompts and generate new information as content.
[0544] Step 6:
[0545] After the user experience ends, the device sends emotional data and feedback to the server. This information includes a log of the user's emotional changes and responses to a feedback form. As output, this data is stored in the server's records and used to improve future experiences. The data serves as crucial foundational information for improving the quality of the user experience.
[0546] (Application Example 2)
[0547] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0548] Traditional virtual shopping systems have faced challenges in providing an optimal shopping experience tailored to individual users, as they often fail to adequately personalize the user experience based on their emotional state. Furthermore, the lack of real-time suggestions for relevant products based on user interests and responses has limited the user experience.
[0549] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0550] In this invention, the server includes means for acquiring digital information related to the destination based on user input information, means for transferring the acquired digital information to the user's information processing device, and means for analyzing the user's emotional state and adjusting the components within the augmented reality environment accordingly. This enables the provision of a personalized shopping experience that responds to the user's emotions and the real-time presentation of relevant product information based on the user's emotional analysis results.
[0551] "User input information" refers to information about destinations and interests that users provide to the system.
[0552] "Digital information" refers to data and content expressed in a format that can be processed by a computer.
[0553] An "information processing device" refers to an electronic device that has the function of receiving, processing, and displaying digital information.
[0554] An "augmented reality environment" refers to a virtual environment where digital information is overlaid onto images and information from the real world and displayed to the user.
[0555] "User behavior" refers to the physical actions of the user, including body movements, gestures, and gaze.
[0556] "Interaction" refers to the exchange of operations and responses that takes place between a system and a user.
[0557] "Emotional state" refers to the mental reactions and facial expressions a user displays in response to a particular stimulus.
[0558] "Analysis" refers to the process of understanding the content of collected data and information in detail.
[0559] "Components" refer to the individual elements that make up the functions, devices, and information parts of a system.
[0560] To realize this application, it is necessary to build a system involving three main elements: user, server, and terminal. First, the user wears an information processing device such as smart glasses or a headset to access the system. When the user inputs information about the destination they wish to visit or their interests, that information is sent to the server. Based on the user's input, the server retrieves digital information related to the destination from a database and transfers that digital information to the user's terminal.
[0561] The device utilizes received digital information to generate an augmented reality environment and displays this environment to the user through the information processing device to which the device is attached. Here, high-precision microphones and cameras are used to analyze the user's actions and emotions. For example, based on video data of the user's face acquired by the camera and audio data collected by the microphone, an emotion recognition API (e.g., Microsoft Azure Face API) is used to analyze the user's emotional state in real time.
[0562] Based on these analysis results, the server dynamically adjusts the components within the augmented reality environment. If a user shows interest in a particular product, it can suggest related information and similar items to enhance the user's shopping experience. The software used includes a framework for augmented reality development (e.g., Unity).
[0563] For example, if a user becomes interested in a raincoat in a virtual shopping mall, the device that detects this interest will then present the user with related new collections and recommended items.
[0564] Furthermore, this system utilizes a generative AI model to improve services based on user responses. For example, input to the AI model can be provided through prompts like the following.
[0565] "Develop methods to predict how users will react to a particular product and provide them with the most relevant additional information."
[0566] "Analyze emotional data and propose the optimal algorithm for providing personalized experiences to users."
[0567] In this way, the present invention can provide a real-time, adaptive shopping experience based on the user's emotional state.
[0568] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0569] Step 1:
[0570] The user enters information about the destinations they wish to visit and their interests. This generates user input information, which is then sent from the user's device to the server.
[0571] Step 2:
[0572] The server retrieves digital information related to the destination from the database based on the received input information. The database stores information about pre-registered destinations. The retrieved digital information becomes the server's output, and this information is transferred to the user's terminal.
[0573] Step 3:
[0574] The terminal generates an augmented reality environment using digital information received from the server. This digital information includes visuals and content related to the destination. The generated augmented reality environment is the terminal's output, and this environment is presented to the user through an information processing device.
[0575] Step 4:
[0576] The user experiences an augmented reality environment through an information processing device. The user's movements, facial expressions, and voice are monitored by the device. As a result, user movement data is input into the terminal.
[0577] Step 5:
[0578] The device collects user behavior data using a high-precision microphone and camera, and analyzes this data using an emotion recognition API. The analyzed data generates an output indicating the user's emotional state. This output is used by the system to personalize the user experience.
[0579] Step 6:
[0580] The server dynamically adjusts the components of the augmented reality environment based on the user's emotional state. For example, if the server determines that the user is excited, it will present new event information or related products. The adjusted environment is the server's output, leading to an optimized user experience.
[0581] Step 7:
[0582] Using a generative AI model, information is generated based on user feedback and sentiment data to help further improve the service. This prompt message includes content such as, "Based on user responses, please consider what additional information should be provided." This output information will be used for future system improvements.
[0583] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0584] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include those described above. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions shown by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0585] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[0586] [Fourth Embodiment]
[0587] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0588] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0589] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0590] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0591] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0592] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0593] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0594] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0595] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0596] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0597] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0598] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0599] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0600] This invention provides a system that enables users to experience travel from home through their own devices. In this system, the user selects a travel destination using an application. Information based on the selection is sent to a server, which retrieves the corresponding content from a database. The content includes high-resolution images, 3D models, and audio data of the destination. The server prepares this content for streaming to the user's device and begins the transfer.
[0601] The user's device processes the received data and generates an augmented reality environment through smart glasses or a headset. This environment is built based on the received data and provides realistic images in the user's field of view. When the user wears the device, they can enjoy a 360-degree virtual journey that responds to head movements and settings.
[0602] For example, if a user chooses to experience "Virtual Kyoto," the server prepares detailed 3D models and audio guides of Kyoto's famous landmarks. The user's device receives this information and uses AR technology to provide an experience of visiting various shrines and temples from the comfort of their home. During this process, the user can use hand movements and voice commands to ask for more detailed information or switch viewpoints.
[0603] Furthermore, users who have completed the experience can provide feedback on their satisfaction level and areas for improvement. This feedback is collected by the server and used to improve the system. In this way, the present invention enables diverse travel experiences that transcend physical limitations and provides new value to users.
[0604] The following describes the processing flow.
[0605] Step 1:
[0606] The user opens the application, selects their desired travel destination, and specifies related activities and preferred language. Once this information is entered, the device sends the request to the server.
[0607] Step 2:
[0608] The server parses the received request and retrieves content related to the destination from the database, such as high-resolution images, 3D models, and audio data. It then organizes this information appropriately and prepares it for data transfer.
[0609] Step 3:
[0610] The server begins streaming the organized content to the user's device. The transfer takes place in real time, and compression techniques are used as needed to optimize the process for smooth data reception.
[0611] Step 4:
[0612] The terminal receives data from the server and decompresses and decodes the received content. This prepares it for building the augmented reality environment. The terminal's processor performs the necessary processing to ensure the system operates smoothly.
[0613] Step 5:
[0614] The device generates an augmented reality environment through smart glasses or a headset. When the user wears the device, 3D objects are overlaid on their field of vision and displayed superimposed on the real world. This allows the user to feel as if they are right in front of their destination.
[0615] Step 6:
[0616] Users can interact within the augmented reality environment using voice commands and hand gestures. For example, they can select specific objects to display detailed information or converse with virtual guides.
[0617] Step 7:
[0618] Once a user finishes their experience, the device collects feedback from the user regarding their satisfaction level and areas for improvement, and sends it to the server. This allows information about system improvements to be accumulated on the server, enriching future experiences.
[0619] (Example 1)
[0620] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0621] There is a growing demand for virtual environments that allow users to experience the real world with greater realism without physical travel. However, current technology fails to deliver a satisfactory experience due to insufficient collection and real-time display of location-related data, as well as inadequate user interaction. Furthermore, it lacks the functionality to effectively utilize user feedback and continuously improve the virtual environment.
[0622] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0623] In this invention, the server includes means for searching for location-related data based on user selection information, means for transmitting the collected data to the user's information processing device, and means for constructing a virtual reality environment using the data received by the information processing device. This enables the user to have a virtual experience with a sense of presence.
[0624] A "user" refers to an individual or group that operates the system and experiences the virtual environment.
[0625] "Selection information" refers to information provided to users to determine the virtual location or environment they wish to visit.
[0626] "Location-related data" refers to a collection of information such as high-resolution images, 3D models, and audio guides about the location selected by the user.
[0627] An "information processing device" refers to an electronic device, such as a computer or smart device owned by a user, that has the function of receiving and processing data.
[0628] "Transmission" refers to the act or means of sending data from a server to an information processing device.
[0629] A "virtual reality environment" refers to a digitally constructed space that does not exist in the real world and can be experienced through user interaction.
[0630] "Interaction" refers to the two-way exchanges, such as operations and responses, that take place between the user and the virtual reality environment.
[0631] "Opinions" refer to feedback from users regarding their experiences and suggestions for improvement.
[0632] A "storage device" is an electronic storage medium used to store data so that it can be retrieved later.
[0633] "Eye movement" refers to actions related to changes in a user's visual focus or eye movements.
[0634] "Instant adjustment" refers to changing the displayed content of the virtual reality environment in real time in response to the user's actions and requests.
[0635] This invention relates to a system that realizes a virtual reality environment using a user's computer or smart device (information processing device). The user installs an application on their device and launches it to use the system. Based on the user's selection information, the user determines the virtual location they wish to visit. For example, the user can choose to experience "Virtual Kyoto."
[0636] The server retrieves data related to the selected location from its internal data store based on the customer's selection. This data includes high-resolution images, 3D models, and audio guides. The server then transmits this data to the user's device. This transmission requires efficient data transfer and typically uses real-time communication protocols over the internet.
[0637] The device receives streamed data and uses AR technologies such as Unity and Unreal Engine to construct a virtual reality environment. This constructed environment is displayed and experienced by the user through smart glasses or a headset. The device instantly adjusts the display within the virtual environment in response to the user's gaze and movements. This dynamic interaction enhances immersion and provides a more realistic experience for the user.
[0638] Users can manipulate the environment using hand movements and voice commands to acquire information and switch viewpoints. User feedback is collected by computer after the experience ends, sent to a server, and stored in memory. This feedback data is used to improve the virtual environment.
[0639] As an example of a generated AI model prompt, "I want to visualize a 3D model of Kyoto and play an audio guide about the temples" provides parameters useful for data retrieval and display. Users can have diverse and immersive travel experiences that go beyond physical limitations.
[0640] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0641] Step 1:
[0642] The user launches an application on their device and selects a virtual location they wish to visit. As input, the user provides destination selection information. This selection information serves as a basis for the system to identify the location data needed in the next step. As output, the selection information is sent to the server.
[0643] Step 2:
[0644] The server searches the data store for data related to the corresponding location based on the selection information it receives. It receives user selection information as input. Data processing involves executing database queries to extract datasets containing high-resolution images, 3D models, audio guides, etc. The collected data is then prepared on the server as output.
[0645] Step 3:
[0646] The server prepares the collected data for transmission to the user's terminal. The retrieved data is used as input. Data processing involves compression and encoding, converting the data into an optimized format for efficient data transfer. As output, the data is ready and transmission begins.
[0647] Step 4:
[0648] The terminal receives data from the server and constructs the virtual reality environment. It receives transmitted images, 3D models, and audio data as input. It uses Unity or Unreal Engine for data processing to generate the virtual reality environment. The generated AR environment is displayed on the display device as output.
[0649] Step 5:
[0650] Users experience the AR environment through smart glasses or a headset. They manipulate the environment through their gaze and body movements, and interact with it through voice and motion. The device receives user gaze and motion information as input, and adjusts the virtual environment display in real time based on this information. The output is an interactive experience that responds to movement.
[0651] Step 6:
[0652] After a user completes an experience, feedback is provided within the system. As input, user opinions and satisfaction data are entered into the device. This data is sent to a server and stored in memory. As output, the feedback data is used for improvement.
[0653] (Application Example 1)
[0654] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0655] Conventional travel experience systems make it difficult to realistically experience diverse travel destinations from the comfort of one's home. Furthermore, they lack features that allow for interaction that reflects user instructions and effectively incorporate post-experience feedback into the system. This makes it difficult to provide truly immersive remote experiences, and also hinders the continuous improvement of the system based on user input.
[0656] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0657] In this invention, the server includes means for acquiring a set of information related to a destination based on user input information, means for transferring the acquired set of information to the user's information processing device, means for generating an augmented reality environment using the information received by the information processing device, and means for operating the functions of the information processing device in response to the user's voice instructions using voice recognition technology. This makes it possible for the user to realistically experience a variety of travel destinations from the comfort of their home, and also enables flexible operation of the augmented reality environment based on the user's voice input and improvement of the system based on the feedback received.
[0658] "User input information" refers to instructions and information related to the travel destination or experience selected by the user through the information processing device.
[0659] "Destination-related information" refers to a collection of digital information, including geographical information, visual data, and audio guides for the selected travel destination.
[0660] An "information processing device" is a terminal used by users to experience virtual reality, and refers to devices such as smartphones and headsets.
[0661] "Transferring means" refers to the processes and technologies used to send data from a server to an information processing device.
[0662] An "augmented reality environment" is a computer-generated image environment that overlays digital information onto real-world environmental information for display.
[0663] "Speech recognition technology" is a technology that interprets the voice spoken by a user and allows a computer to respond appropriately based on that interpretation.
[0664] "Voice commands" are commands or requests that a user makes to a system or device using their voice.
[0665] A system implementing this invention begins by installing an application equipped with a user interface for the user to select a travel destination on an information processing device. When the user selects a travel destination, that information is sent to a server. The server retrieves a set of information related to the destination selected by the user from a database and transfers the information, including geographic information, visual data, and audio guides, that constitutes that set of information to the information processing device.
[0666] The information processing device generates an augmented reality environment using the information received from the server. Using software such as ARKit (iOS) and ARCore (Android), users can visually experience digital information superimposed on the real world. For example, if a user selects "Virtual Paris Trip," detailed 3D models of the Eiffel Tower and the Louvre Museum are delivered and displayed in the augmented reality environment.
[0667] Furthermore, the information processing device uses speech recognition technology (such as Google Speech-to-Text) to interpret the user's voice commands. This allows the user to interact with and manipulate the environment through voice. For example, if the user asks, "What is the height of the Eiffel Tower?", the information will be provided immediately.
[0668] This system allows users to experience visiting tourist destinations around the world from the comfort of their homes, and to provide feedback based on their experiences, contributing to system improvements. By inputting prompts such as, "Provide detailed information about the travel destination chosen by the user as a virtual experience with a visually represented 3D model and audio support," into the generating AI model, a more realistic and immersive experience can be achieved.
[0669] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0670] Step 1:
[0671] The user selects a travel destination through the application interface of the information processing device. The input is the name of the travel destination selected by the user. Based on the user's selection, the information is sent to the server.
[0672] Step 2:
[0673] The server receives the name of the travel destination sent by the user, and retrieves a set of information related to the destination, such as visual data, geographical information, and audio guides, by referring to the database. The input is the travel destination selected by the user, and the output is a set of various information related to that travel destination.
[0674] Step 3:
[0675] The server encodes the acquired information and streams or downloads it to the information processing device. The input is the information acquired by the server, and the output is the data transferred to the information processing device.
[0676] Step 4:
[0677] The information processing device uses ARKit or ARCore to generate an augmented reality environment based on the information received from the server. The input is the received information, and the output is the augmented reality environment presented to the user.
[0678] Step 5:
[0679] Users use voice commands to ask questions of or operate the information processing device. The input is the voice commands spoken by the user.
[0680] Step 6:
[0681] The information processing device uses speech recognition technology to convert user voice commands into text data and displays appropriate information in an augmented reality environment based on that input. The input is the user's voice command, and the output is the corresponding information or operation result.
[0682] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0683] This invention provides a personalized experience tailored to the user's emotional state by combining an emotion engine with a travel experience system using an augmented reality environment. This system operates by integrating the user's terminal, server, and emotion engine.
[0684] First, the user selects a destination they wish to visit through the application and enters relevant information. Once this information is sent to the server via the device, the server retrieves content related to the destination from its database and streams it to the user's device.
[0685] The user's device generates an augmented reality environment based on the received content and provides visual information to the user through smart glasses or a headset. Furthermore, the device uses a built-in emotion engine to infer emotions from the user's facial expressions, tone of voice, and other biometric information.
[0686] The emotion engine analyzes the user's emotions in real time and adjusts the content and interactions within the augmented reality environment based on this analysis. For example, if the user is relaxed, it provides a calm environment; if they are excited, it provides a more stimulating experience.
[0687] As a concrete example, suppose a user chooses to experience "Virtual New York." The server retrieves a 3D model of Times Square and event information and streams it. The device then constructs this as an augmented reality environment and presents it to the user's field of view. The emotion engine analyzes the user's reactions, and if the user is surprised or delighted, it suggests more fun attractions or displays information tailored to their interests.
[0688] After the experience ends, emotional data and feedback collected from users are sent to a server and used to improve future experiences. This allows for the provision of travel experiences optimized for each individual user. This invention enables dynamic analysis of emotions and contributes to further personalizing the user experience.
[0689] The following describes the processing flow.
[0690] Step 1:
[0691] The user launches the application and selects a destination they wish to visit. They then enter information such as related activities and preferred language, and send this information to the server via their device.
[0692] Step 2:
[0693] The server analyzes the received request and retrieves high-resolution images, 3D models, and audio data related to the destination from the database. It then packages this data and prepares it for streaming to the user's device.
[0694] Step 3:
[0695] The device receives content sent from the server. It then decompresses and decodes the content, preparing to generate the augmented reality environment.
[0696] Step 4:
[0697] The device's built-in emotion engine collects user facial expression data and voice tone. Based on this, it analyzes the user's emotional state in real time.
[0698] Step 5:
[0699] The device generates an augmented reality environment and displays it in the user's field of view using smart glasses or a headset. The content and audio guidance presented are dynamically adjusted according to the user's emotional state.
[0700] Step 6:
[0701] Users enjoy an interactive experience based on analysis by an emotion engine. For example, if a user is relaxed, calming music and scenery are emphasized, while if they are excited, active activities are presented.
[0702] Step 7:
[0703] When a user finishes an experience, the device sends emotional data and user feedback to the server. This information is stored in a database for system improvement and further personalization.
[0704] (Example 2)
[0705] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0706] Traditional augmented reality systems have struggled to dynamically adjust content based on user emotions and individual actions, making it difficult to provide an experience optimized for each user. Therefore, there is a growing need to provide travel experiences that are more adaptable to users' emotions and interests.
[0707] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0708] In this invention, the server includes means for acquiring destination-related information based on user input information, means for transferring the acquired information to the user's device, and means for generating an augmented reality environment using the information received by the device. This enables dynamic adjustment of display content based on the user's biometric information, analysis of emotions, and an optimized travel experience.
[0709] "User input information" refers to information that includes destinations the user wants to visit, activities of interest, and other personal preferences.
[0710] "Destination-related information" refers to information including geographical information, historical background, and data on tourist attractions related to the destination selected by the user.
[0711] "User device" refers to a smartphone, tablet, or dedicated device that a user uses to receive information and content and visualize the augmented reality environment.
[0712] An "augmented reality environment" refers to a technological environment that provides a richer and more interactive visual experience by overlaying digital information onto the physical reality environment.
[0713] "Biometric information" refers to data that includes a user's physiological and behavioral characteristics, such as facial expressions, tone of voice, and other physical responses.
[0714] "Analyzing and optimizing emotions" refers to the process of analyzing a user's biometric information and real-time emotional state, and then individually adjusting the user's experience based on the results.
[0715] "Collecting feedback and making improvements" refers to the process of gathering evaluations and comments from users after their experience and using that feedback to improve the system's content and functionality.
[0716] This invention is an augmented reality system designed to provide a personalized travel experience based on the user's emotions. Specific embodiments are described below.
[0717] The user begins a virtual journey using an application on their mobile device. First, the user selects a destination they wish to visit via the application and enters the necessary details. This information is sent from the user's device to the server. The server searches a database to retrieve various information related to the destination. This process uses web server software (e.g., generic name) and a database management system (e.g., generic name).
[0718] The acquired information is streamed to the user's device in real time, and the device uses this information to generate an augmented reality environment. The generation of augmented reality utilizes a corresponding AR platform (e.g., general name) and is visually presented to the user through smart glasses or compatible devices. This environment includes 3D models, guide information, and real-time information on local events.
[0719] Furthermore, the device is equipped with a built-in emotion engine. It collects the user's facial expressions and voice via the device's camera and microphone, and uses machine learning libraries (e.g., generic names) to analyze the user's emotions in real time. The emotion-based analysis results are used to adjust the augmented reality environment, optimizing the user experience.
[0720] For example, if a user selects "Virtual Times Square," the server provides relevant 3D models and information. If the user's emotions indicate excitement or surprise, the system adds more detailed information and interesting attractions.
[0721] After the experience ends, emotional data and feedback collected from users are stored on the server and used to improve future services. This allows for the creation of optimal travel experiences tailored to each user.
[0722] As an example of input to the generative AI model, prompt sentences such as "Explain how this system optimizes the augmented reality experience based on the user's emotions" can be used.
[0723] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0724] Step 1:
[0725] The user launches an application on their device, selects a destination they wish to visit and activities of interest, and enters detailed information. Based on this, the device sends destination information to the server in text format. The input data includes place names, planned activities, and the desired style of experience. The server receives the selected place names and associated user parameters as output.
[0726] Step 2:
[0727] The server searches the database based on the received destination information to retrieve relevant data. Using specific queries, it retrieves 3D model data and related tourist attractions and event information from the database management system. Input data consists of user selections sent to the server, while output data is content sent to the terminal in list format. Specifically, the server executes a backend program to extract information according to the user's preferences.
[0728] Step 3:
[0729] The terminal generates an augmented reality environment based on information received from the server. Using an AR-enabled device, 3D models and related information are integrated into the user's field of view. Using an AR development platform such as Unity, virtual objects are overlaid onto the real environment. This operation is performed by processing information packets sent from the server as input and providing the visualized AR content to the user as output.
[0730] Step 4:
[0731] The device activates a built-in emotion engine to collect the user's biometric information in real time. It analyzes facial expressions and voice tone captured from the camera and microphone, and uses machine learning algorithms to classify emotions. The input data consists of video and audio, and the output is the user's emotional state (e.g., surprise, joy). This process provides the capability to detect the user's reactions in real time.
[0732] Step 5:
[0733] The server dynamically adjusts the augmented reality content based on the output from the emotion engine. For example, if the user is surprised, it will present additional information or attractions that are more exciting. The input data is emotion information sent from the device, and the output data is the updated AR content. In this step, the server uses a generative AI model to create prompts and generate new information as content.
[0734] Step 6:
[0735] After the user experience ends, the device sends emotional data and feedback to the server. This information includes a log of the user's emotional changes and responses to a feedback form. As output, this data is stored in the server's records and used to improve future experiences. The data serves as crucial foundational information for improving the quality of the user experience.
[0736] (Application Example 2)
[0737] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0738] Traditional virtual shopping systems have faced challenges in providing an optimal shopping experience tailored to individual users, as they often fail to adequately personalize the user experience based on their emotional state. Furthermore, the lack of real-time suggestions for relevant products based on user interests and responses has limited the user experience.
[0739] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0740] In this invention, the server includes means for acquiring digital information related to the destination based on user input information, means for transferring the acquired digital information to the user's information processing device, and means for analyzing the user's emotional state and adjusting the components within the augmented reality environment accordingly. This enables the provision of a personalized shopping experience that responds to the user's emotions and the real-time presentation of relevant product information based on the user's emotional analysis results.
[0741] "User input information" refers to information about destinations and interests that users provide to the system.
[0742] "Digital information" refers to data and content expressed in a format that can be processed by a computer.
[0743] An "information processing device" refers to an electronic device that has the function of receiving, processing, and displaying digital information.
[0744] An "augmented reality environment" refers to a virtual environment where digital information is overlaid onto images and information from the real world and displayed to the user.
[0745] "User behavior" refers to the physical actions of the user, including body movements, gestures, and gaze.
[0746] "Interaction" refers to the exchange of operations and responses that takes place between a system and a user.
[0747] "Emotional state" refers to the mental reactions and facial expressions a user displays in response to a particular stimulus.
[0748] "Analysis" refers to the process of understanding the content of collected data and information in detail.
[0749] "Components" refer to the individual elements that make up the functions, devices, and information parts of a system.
[0750] To realize this application, it is necessary to build a system involving three main elements: user, server, and terminal. First, the user wears an information processing device such as smart glasses or a headset to access the system. When the user inputs information about the destination they wish to visit or their interests, that information is sent to the server. Based on the user's input, the server retrieves digital information related to the destination from a database and transfers that digital information to the user's terminal.
[0751] The device utilizes received digital information to generate an augmented reality environment and displays this environment to the user through the information processing device to which the device is attached. Here, high-precision microphones and cameras are used to analyze the user's actions and emotions. For example, based on video data of the user's face acquired by the camera and audio data collected by the microphone, an emotion recognition API (e.g., Microsoft Azure Face API) is used to analyze the user's emotional state in real time.
[0752] Based on these analysis results, the server dynamically adjusts the components within the augmented reality environment. If a user shows interest in a particular product, it can suggest related information and similar items to enhance the user's shopping experience. The software used includes a framework for augmented reality development (e.g., Unity).
[0753] For example, if a user becomes interested in a raincoat in a virtual shopping mall, the device that detects this interest will then present the user with related new collections and recommended items.
[0754] Furthermore, this system utilizes a generative AI model to improve services based on user responses. For example, input to the AI model can be provided through prompts like the following.
[0755] "Develop methods to predict how users will react to a particular product and provide them with the most relevant additional information."
[0756] "Analyze emotional data and propose the optimal algorithm for providing personalized experiences to users."
[0757] In this way, the present invention can provide a real-time, adaptive shopping experience based on the user's emotional state.
[0758] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0759] Step 1:
[0760] The user enters information about the destinations they wish to visit and their interests. This generates user input information, which is then sent from the user's device to the server.
[0761] Step 2:
[0762] The server retrieves digital information related to the destination from the database based on the received input information. The database stores information about pre-registered destinations. The retrieved digital information becomes the server's output, and this information is transferred to the user's terminal.
[0763] Step 3:
[0764] The terminal generates an augmented reality environment using digital information received from the server. This digital information includes visuals and content related to the destination. The generated augmented reality environment is the terminal's output, and this environment is presented to the user through an information processing device.
[0765] Step 4:
[0766] The user experiences an augmented reality environment through an information processing device. The user's movements, facial expressions, and voice are monitored by the device. As a result, user movement data is input into the terminal.
[0767] Step 5:
[0768] The device collects user behavior data using a high-precision microphone and camera, and analyzes this data using an emotion recognition API. The analyzed data generates an output indicating the user's emotional state. This output is used by the system to personalize the user experience.
[0769] Step 6:
[0770] The server dynamically adjusts the components of the augmented reality environment based on the user's emotional state. For example, if the server determines that the user is excited, it will present new event information or related products. The adjusted environment is the server's output, leading to an optimized user experience.
[0771] Step 7:
[0772] Using a generative AI model, information is generated based on user feedback and sentiment data to help further improve the service. This prompt message includes content such as, "Based on user responses, please consider what additional information should be provided." This output information will be used for future system improvements.
[0773] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0774] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include those described above. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions shown by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0775] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[0776] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0777] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[0778] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[0779] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[0780] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[0781] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[0782] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[0783] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[0784] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[0785] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0786] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0787] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[0788] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[0789] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[0790] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[0791] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0792] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0793] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted as being incorporated by reference.
[0794] The following is further disclosed regarding the embodiments described above.
[0795] (Claim 1)
[0796] A means of obtaining destination-related content based on user input information,
[0797] A means for transferring the acquired content to the user's terminal,
[0798] A means for generating an augmented reality environment using the content received by the terminal,
[0799] Means for controlling interactions within the augmented reality environment based on user actions,
[0800] A system that includes this.
[0801] (Claim 2)
[0802] The system according to claim 1, further comprising means for collecting user feedback and storing it in a database for improving the augmented reality environment.
[0803] (Claim 3)
[0804] The system according to claim 1, further comprising means for adjusting the display in the augmented reality environment in real time in response to changes in the user's viewpoint.
[0805] "Example 1"
[0806] (Claim 1)
[0807] A means of searching for location-related data based on user selection information,
[0808] A means for transmitting the collected data to the user's information processing device,
[0809] The information processing device provides means for constructing a virtual reality environment using the data it has received,
[0810] Means for managing use within the virtual reality environment based on the user's physical movements,
[0811] A system that includes this.
[0812] (Claim 2)
[0813] The system according to claim 1, further comprising means for collecting user feedback and storing it in a storage device for improving the virtual reality environment.
[0814] (Claim 3)
[0815] The system according to claim 1, further comprising means for instantly adjusting the display in the virtual reality environment in response to the user's eye movements.
[0816] "Application Example 1"
[0817] (Claim 1)
[0818] A means for obtaining a set of information related to the destination based on user input information,
[0819] Means for transferring the acquired information group to the user's information processing device,
[0820] The information processing device provides means for generating an augmented reality environment using the information set it receives,
[0821] Means for controlling the interaction within the augmented reality environment based on user actions,
[0822] A means for operating the functions of an information processing device in response to a user's voice command using speech recognition technology,
[0823] A system that includes this.
[0824] (Claim 2)
[0825] The system according to claim 1, further comprising means for obtaining user evaluations and storing them on a recording medium for improving the augmented reality environment.
[0826] (Claim 3)
[0827] The system according to claim 1, further comprising means for instantly adjusting the display in the augmented reality environment in response to changes in the user's viewpoint.
[0828] "Example 2 of combining an emotion engine"
[0829] (Claim 1)
[0830] A means of obtaining destination-related information based on user input information,
[0831] Means for transferring the acquired information to the user's device,
[0832] The device includes means for generating an augmented reality environment using the information received by the device,
[0833] Means for dynamically adjusting the display content within the augmented reality environment based on the user's biometric information,
[0834] A means for analyzing user emotions and optimizing information and interactions in the augmented reality environment,
[0835] A system that includes this.
[0836] (Claim 2)
[0837] The system according to claim 1, further comprising means for collecting user feedback and storing it in records for improving the augmented reality environment.
[0838] (Claim 3)
[0839] The system according to claim 1, further comprising means for adjusting the display in the augmented reality environment in real time based on the user's visual information.
[0840] "Application example 2 when combining with an emotional engine"
[0841] (Claim 1)
[0842] A means of obtaining digital information related to the destination based on user input information,
[0843] means for transferring the acquired digital information to the user's information processing device,
[0844] The information processing device provides means for generating an augmented reality environment using digital information received by the information processing device,
[0845] Means for controlling interactions within the augmented reality environment based on user actions,
[0846] A means for analyzing the user's emotional state and adjusting the components within the augmented reality environment accordingly,
[0847] A system that includes this.
[0848] (Claim 2)
[0849] The system according to claim 1, further comprising means for collecting user feedback and storing it in a storage device for improving the augmented reality environment.
[0850] (Claim 3)
[0851] Means for adjusting the display within the augmented reality environment in real time in response to changes in the user's viewpoint,
[0852] The system according to claim 1, further comprising means for dynamically presenting information on related products based on the results of user sentiment analysis. [Explanation of symbols]
[0853] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. A means of obtaining destination-related content based on user input information, A means for transferring the acquired content to the user's terminal, A means for generating an augmented reality environment using the content received by the terminal, Means for controlling interactions within the augmented reality environment based on user actions, A system that includes this.
2. The system according to claim 1, further comprising means for collecting user feedback and storing it in a database for improving the augmented reality environment.
3. The system according to claim 1, further comprising means for adjusting the display in the augmented reality environment in real time in response to changes in the user's viewpoint.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A