System
The virtual travel system using head-mounted displays and AI chatbots provides immersive virtual experiences of historical events and scientific phenomena, addressing the limitations of traditional static media in conveying these subjects.
Patent Information
- Application Number
- JP2024181290
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-10-23
- Filing Date
- 2024-10-16
- Publication Date
- 2025-05-08
AI Technical Summary
Traditional methods for understanding historical events and scientific phenomena rely on static media like books and videos, making it difficult for users to experience these events directly and visually.
A virtual travel system using head-mounted displays that acquires and converts video data related to historical events or scientific phenomena, allowing users to experience these events virtually while interacting with AI chatbots for detailed explanations.
Enables users to gain a deeper understanding of historical events and scientific phenomena through immersive virtual experiences, overcoming the limitations of traditional static media.
Smart Images

Figure 2025071785000001_ABST
Abstract
Description
[Technical field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including a description and related instruction sentence regarding the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] JP 2022-180282 A Summary of the Invention [Problem to be solved by the invention]
[0004] The present invention aims to enable users to take a virtual trip to the heart of historical events and scientific phenomena. Conventional methods require users to refer to books or video materials in order to understand historical events and scientific phenomena, which makes it difficult to have a direct experience or visual understanding. [Means for solving the problem]
[0005] The present invention provides means by which a user can virtually travel to the heart of historical events and scientific phenomena using applications for head-mounted displays, including the following means:
[0006] A means of retrieving relevant visual data based on a user-selected historical event or scientific phenomenon.
[0007] A means of converting acquired video data to match the user's viewpoint and displaying it on a head-mounted display.
[0008] A means to provide users with detailed explanations of selected historical events or scientific phenomena via an AI chatbot as a guide.
[0009] By using the above means, the user can understand historical events and scientific phenomena through direct experience. The present invention makes it possible to solve the problems of direct experience and visual understanding that have been a problem in the past.
[0010] A "virtual travel system" is a system that allows a user to virtually travel to the center of a historical event or scientific phenomenon using a head-mounted display.
[0011] "Video Data" means data representing visual information related to a selected historical event or scientific phenomenon, which is transformed to match the user's viewpoint and displayed on a head-mounted display.
[0012] An "AI chatbot as a guide" is an artificial intelligence chatbot that is used to provide users with detailed explanations of selected historical events or scientific phenomena, providing information through dialogue with the user and deepening their understanding. [Brief description of the drawings]
[0013] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Diagram 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. FIG. [Diagram 3] FIG. 11 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Diagram 5] FIG. 13 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 13 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 13 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] 4 is a sequence diagram showing a process flow of the data processing system according to the first embodiment. FIG. [Figure 12] 11 is a sequence diagram showing a process flow of the data processing system in application example 1. FIG. [Figure 13] FIG. 11 is a sequence diagram showing the flow of processing of the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 11 is a sequence diagram showing the flow of processing in the data processing system in application example 2 when combined with an emotion engine. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0014] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.
[0015] First, the terms used in the following description will be explained.
[0016] In the following embodiments, a signed processor (hereinafter simply referred to as a "processor") may be one arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be one type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), etc.
[0017] In the following embodiments, a signed RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0018] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0019] In the following embodiments, a communication I / F (Interface) with a code is an interface including a communication processor and an antenna. The communication I / F controls communication between multiple computers. An example of a communication standard applied to the communication I / F is a wireless communication standard including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0020] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. In addition, in this specification, the same idea as "A and / or B" is also applied when three or more things are expressed by connecting them with "and / or."
[0021] [First embodiment]
[0022] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0023] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0024] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a wide area network (WAN) and / or a local area network (LAN).
[0025] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0026] The reception device 38 includes a touch panel 38A and a microphone 38B, and receives user input. The touch panel 38A detects contact with a pointer (e.g., a pen or a finger) to receive user input by the touch of the pointer. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0027] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (e.g., voice and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs voice according to instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, an aperture, and a shutter, and an imaging element such as a Complementary Metal-Oxide-Semiconductor (CMOS) image sensor or a Charge Coupled Device (CCD) image sensor.
[0028] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for transmitting and receiving various types of information between the processor 46 and the processor 28 via the network 54.
[0029] FIG. 2 shows an example of main functions of the data processing device 12 and the smart device 14.
[0030] As shown in Fig. 2, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32. The specific process program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific process program 56 from the storage 32, and executes the read specific process program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific process program 56 executed on the RAM 30.
[0031] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0032] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores a reception output program 60. The reception output program 60 is used together with the specific processing program 56 by the data processing system 10. The processor 46 reads out the reception output program 60 from the storage 50, and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0033] Next, a description will be given of the specific processing by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal".
[0034] An embodiment for implementing the present invention includes the following elements.
[0035] 1. Server
[0036] Build a database to capture video data related to historical events and scientific phenomena. Historical events and scientific phenomena are examples of "events."
[0037] Based on the user's selection, relevant video data is extracted and transformed to suit the user's perspective.
[0038] It provides an AI chatbot that acts as a guide and provides detailed explanations to users.
[0039] 2. Terminal
[0040] Video data is received from the server based on a user-selected historical event or scientific phenomenon.
[0041] The received video data is displayed on a head-mounted display.
[0042] It accepts user operations and sends the selected information to the server.
[0043] 3. Users
[0044] By operating the terminal, users select the historical event or scientific phenomenon they would like to virtually travel through.
[0045] Wearing a head-mounted display, participants visually experience a virtual site and central area.
[0046] Interact with an AI chatbot that acts as your guide and receives detailed explanations.
[0047] The above is an example of an embodiment of the present invention. The server acquires and converts video data and provides an AI chatbot, the terminal receives and displays the video data, and the user selects and experiences information and interacts with the AI chatbot. This allows the user to take a virtual trip and gain a deeper understanding of historical events and scientific phenomena.
[0048] The process flow will be explained below.
[0049] Step 1: The server acquires and converts the video data.
[0050] The server retrieves video data related to historical events or scientific phenomena selected by the user.
[0051] The acquired video data is extracted from a database in the server.
[0052] The server converts the captured video data to match the user's viewpoint, allowing the user to visually experience a virtual site or center.
[0053] Step 2: The device receives and displays the video data
[0054] The terminal receives the video data transmitted from the server.
[0055] The received video data is displayed on a head-mounted display inside the terminal.
[0056] Users can wear a head-mounted display and visually experience a virtual site or central area.
[0057] Step 3: User makes information selections and experiences
[0058] Users operate the terminal to select the historical event or scientific phenomenon about which they would like to take a virtual journey.
[0059] The selected information is transmitted from the terminal to the server.
[0060] Based on the user's selection, the server retrieves the relevant video data, converts it and sends it to the terminal.
[0061] Step 4: User interacts with AI chatbot and receives detailed instructions
[0062] Users interact with the AI chatbot through a head-mounted display.
[0063] The AI chatbot provides users with detailed explanations of selected historical events or scientific phenomena.
[0064] Users can gain deeper understanding through dialogue with AI chatbots.
[0065] This is the flow of the program's processing. The server acquires and converts the video data, and the terminal receives and displays the video data. The user can select and experience information, and receive detailed explanations through dialogue with the AI chatbot.
[0066] Example 1
[0067] Next, a description will be given of Example 1. In the following description, the data processing device 12 is referred to as a "server" and the smart device 14 is referred to as a "terminal."
[0068] Currently, there are only a limited number of systems that allow users to visually experience historical events or scientific phenomena, making it difficult for users to gain a deep understanding. In addition, existing systems are unable to convert in real time to match the user's perspective, which results in a lack of immersion. Furthermore, there are insufficient guides to provide detailed information, and information supplementation is not adequate.
[0069] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0070] In this invention, the server includes a means for constructing a database for acquiring video data, a means for extracting video data based on a specific event or phenomenon selected by the user, a means for converting the extracted video data to match the user's viewpoint and displaying it, and a means for interacting with an AI chatbot as a guide that provides detailed explanations to the user based on a request provided by the user, thereby enabling the user to gain a deeper understanding of historical events and scientific phenomena through a highly immersive virtual journey.
[0071] A "database" is a collection of data in which information is systematically collected, stored, and made available for rapid retrieval.
[0072] "Video data" is digital or analog data that contains visual information, and is content that includes moving images and still images.
[0073] "Transformation to fit the viewpoint" refers to the process of adjusting the original video data according to the user's field of view and visual needs, and displaying it appropriately.
[0074] An "AI chatbot as a guide" is an automated dialogue system that uses artificial intelligence to provide appropriate information in response to user questions.
[0075] A "visual display device" is a device that allows a user to visually experience video data, and primarily refers to a head-mounted display or a monitor.
[0076] "Additional information" is supplemental information about the event or phenomenon the user is experiencing, including detailed explanations and background knowledge.
[0077] A "request" refers to the act of a user sending a specific request or instruction to the system.
[0078] The embodiment for implementing the present invention includes the following elements.
[0079] 1. Server configuration and roles:
[0080] The server builds a database to acquire and store video data related to historical events and scientific phenomena. Specifically, the server manages a huge amount of video content and extracts relevant video data based on a specific event or phenomenon selected by the user. For example, the database can include "video of pyramid construction in ancient Egypt" and "video on the development of scientific theories." Furthermore, the server converts the extracted video data to match the user's viewpoint. This conversion is intended to enable the user to view the video from a 360-degree viewpoint. The server also provides an AI chatbot as a guide, providing detailed explanations for requests provided by the user.
[0081] 2. Terminal configuration and roles:
[0082] The terminal receives video data from the server based on historical events or scientific phenomena selected by the user. The received video data is displayed on a visual display device (e.g., a head-mounted display). A specific example is a device such as a VR headset. The terminal accepts user operations and transmits the operation contents to the server. For example, if the user requests to zoom in on a particular scene or to know more about a particular object, the request is transmitted to the server.
[0083] 3. User interaction and experience:
[0084] Users can operate the device to experience a virtual journey. For example, a user can select a specific historical event or scientific phenomenon, such as "the discovery of Machu Picchu" or "scientific experiment in a laboratory." After that, the user puts on the head-mounted display and visually experiences the selected scene from a 360-degree perspective. Furthermore, if the user wants to know more information, they can interact with an AI chatbot that acts as a guide. For example, if the user asks, "What were the most commonly sold products in this ancient market?", the AI chatbot will answer, "Grains and olive oil were commonly traded."
[0085] Examples:
[0086] As a concrete example, consider the case where the user selects "Darwin's Voyage of the Beagle." In this case, the user receives the video data from the server through the terminal and it is displayed on the head-mounted display. When the user asks a question about the plants and animals discovered during Darwin's voyage, the AI chatbot will explain the details.
[0087] Examples of prompts for a generative AI model include:
[0088] "What were the most commonly sold items in ancient Roman markets?"
[0089] "I would like to know more about what Darwin discovered on his Beagle voyage."
[0090] "I want to see an episode about the discovery of the theory of relativity."
[0091] This allows users to take a virtual journey and gain a deeper understanding of historical events or scientific phenomena.
[0092] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0093] Step 1:
[0094] Database construction (server)
[0095] The server creates a database that systematically collects and stores information.
[0096] Input: Video content related to historical events or scientific phenomena.
[0097] Data processing: Classify video content by category and create an index.
[0098] Output: The indexed database.
[0099] Specific operation: The server collects, for example, footage of "scientific experiments" or "ancient civilizations" from the Internet, adds metadata to them, and stores them in a database.
[0100] Step 2:
[0101] Receiving user request (terminal)
[0102] The terminal provides an interface for the user to select a particular event or phenomenon.
[0103] Input: A request for a user-selected event or phenomenon.
[0104] Data processing: Analyze the request content and send it to the server to extract the necessary data.
[0105] Output: The parsed request data.
[0106] Specific operation: When a user selects "Darwin's Voyage of the Beagle," the request is sent from the terminal to the server.
[0107] Step 3:
[0108] Extraction and conversion of video data (server)
[0109] The server extracts relevant video data from a database based on the request.
[0110] Input: The user's request data.
[0111] Data processing: Based on the request, relevant video data is searched for in the database and the video is transformed to suit the user's perspective.
[0112] Output: Video data transformed to match the user's viewpoint.
[0113] Specific operation: The server retrieves footage, for example, of "Darwin's Voyage of the Beagle" from a database and converts the footage into a 360-degree view.
[0114] Step 4:
[0115] Video data transmission (server)
[0116] The server transmits the converted video data to the terminal.
[0117] Input: Video data transformed to match the user's viewpoint.
[0118] Data processing: Encode video data into a transmittable format.
[0119] Output: The encoded video data is sent to the device.
[0120] Specific operation: The server transmits video data encoded in, for example, MP4 format to the terminal in real time.
[0121] Step 5:
[0122] Receiving video data (terminal)
[0123] The terminal receives the video data transmitted from the server.
[0124] Input: Encoded video data sent from the server.
[0125] Data processing: Decoding the encoded video data and converting it into a displayable format.
[0126] Output: Decoded video data.
[0127] Specific operation: The terminal decodes the received video data, for example in MP4 format, and stores it in its internal memory.
[0128] Step 6:
[0129] Display of video data (terminal)
[0130] The terminal displays the received video data on a visual display device.
[0131] Input: Decoded video data.
[0132] Data Processing: Converting the decoded video data into a resolution and format suitable for a visual display device.
[0133] Output: Video data displayed on a visual display device.
[0134] Specific operation: The terminal displays a 360-degree image on, for example, a head-mounted display (HMD), allowing the user to visually experience the scene.
[0135] Step 7:
[0136] Receiving user operations (terminal)
[0137] The terminal receives user operations and requests.
[0138] Input: Actions taken by the user or additional requests.
[0139] Data processing: Analyzes operation and request data and sends it to the server.
[0140] Output: Parsed operation and request data.
[0141] Specific operation: When a user makes a request such as "I would like to take a closer look at Darwin's notes," the request is sent to the server via the terminal.
[0142] Step 8:
[0143] Interaction with AI chatbot (server, user)
[0144] Based on the user's request, the server launches an AI chatbot to act as a guide and engages in a conversation.
[0145] Input: The request made by the user.
[0146] Data processing: The AI chatbot retrieves information corresponding to the request and provides the user with an appropriate answer.
[0147] Output: A detailed description provided to the user.
[0148] What it does: When a user asks, "What's in Darwin's notebooks?" the AI chatbot responds, "These contain Darwin's early thoughts on evolution."
[0149] (Application example 1)
[0150] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0151] Traditional education of historical events and scientific phenomena has often relied on static media such as textbooks and videos, and it has been difficult to actually experience them. This has made it difficult to maintain users' understanding and interest. In addition, guidebooks and human guides have limitations in their explanations, and the information is not always consistently detailed and individualized. This has created a need for an effective way for users to gain a deeper understanding of specific events and phenomena.
[0152] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0153] In this invention, the server includes a means for acquiring video data related to a historical event or a scientific phenomenon selected by a user based on the historical event or the scientific phenomenon, a means for converting the acquired video data to match the user's viewpoint and displaying it, a means for interacting with an AI chatbot that acts as a guide and provides the user with a detailed explanation of the selected historical event or the scientific phenomenon, a means for the user to experience the video data by wearing a head-mounted display in a physical store, and a means for the AI chatbot to provide a detailed explanation in real time as a guide. This allows the user to experience the historical event or scientific phenomenon in an immersive way and receive a detailed and personalized explanation from the AI chatbot.
[0154] A "user" is someone who uses the system to experience historical events or scientific phenomena.
[0155] A "historic event" refers to a specific occurrence or phenomenon that had a significant impact on human history.
[0156] A "scientific phenomenon" refers to an event that allows for the scientific understanding and analysis of laws and phenomena in the natural world.
[0157] "Video data" refers to video and image information relating to a specific event or phenomenon.
[0158] "Means of acquiring" refers to the method or mechanism for getting the required information or data into a system or device.
[0159] "Means for converting to suit the viewpoint" refers to technology that adjusts video data in an appropriate manner according to the user's viewpoint and position.
[0160] The term "means for displaying" refers to a mechanism or method for presenting video data in a form that can be visually recognized by a user.
[0161] An "AI chatbot as a guide" refers to a program that uses artificial intelligence technology to provide information to users and engage in dialogue.
[0162] "Means of dialogue" refers to the technology that allows users and AI chatbots to communicate with each other through questions and answers.
[0163] "System" refers to a set of mechanisms in which multiple means and devices work together to provide specific services and functions to users.
[0164] A "head-mounted display" refers to a display device that is worn on the head and allows the user to visually experience immersive images.
[0165] "Real-time" refers to data and information being processed and displayed instantly, without delay.
[0166] A "database" refers to a system for systematically organizing and storing specific information or data.
[0167] "Means for providing detailed explanations" refers to a method for providing specific and detailed information necessary for users to deepen their understanding.
[0168] A "physical store" refers to a sales location that has a physical presence to offer goods or services, and is a permanent commercial space.
[0169] "Means of experiencing" refers to the technology or methods that allow a user to actually experience a particular event through simulation or virtual reality.
[0170] In the embodiment of the present invention, the entire system is configured as follows: The system used by a user mainly comprises a server, a terminal, and a head mounted display.
[0171] 1. Server
[0172] The server is responsible for building a database to retrieve video data related to historical events or scientific phenomena selected by the user. The retrieved video data is then transformed to fit the user's perspective. The server also provides an AI chatbot to provide relevant detailed explanations to the user in real time.
[0173] The software used includes database access via REST API and a natural language processing engine for the AI chatbot. Specifically, the "requests" library is used to acquire video data, and a GPT series of generative AI models are used for natural language processing.
[0174] 2. Terminal
[0175] The terminal operated by the user receives the video data from the server and displays the data on the head mounted display. The terminal accepts the user's operations and transmits the information to the server.
[0176] The terminals are general-purpose smartphones or tablet devices equipped with a browser application that uses HTML5 or WebGL to receive and display video data.
[0177] 3. Head-mounted displays
[0178] The user wears a head-mounted display and visually experiences the video data transmitted from the server. The video data is updated in real time based on the user's viewpoint, providing an immersive experience.
[0179] The hardware used includes VR headsets, which track the movements of the user's head and adjust the image based on that point of view.
[0180] Examples
[0181] In the education section of the brick-and-mortar store, a visitor selects to experience "Egyptian pyramid construction." The visitor puts on a head-mounted display and is immersed in the pyramid construction site that unfolds before their eyes, witnessing workers stacking stones. In real time, an AI chatbot explains, "This site was built around 2500 BC," and provides additional details, "The stone used was mainly limestone."
[0182] Examples of prompt statements
[0183] "Generate guides based on video data about the construction of the Egyptian pyramids to allow visitors to virtually experience the sites."
[0184] As described above, the system of the present invention allows users to experience a specific historical event or scientific phenomenon in real time while receiving detailed explanations from an AI chatbot, which leads to deeper understanding and longer-lasting interest compared to traditional static educational methods.
[0185] The flow of the specific process in the application example 1 will be described with reference to FIG.
[0186] Step 1:
[0187] The user operates the terminal to select the historical event or scientific phenomenon they wish to experience.
[0188] Input: Event ID or theme selected by the user
[0189] Output: The selected event ID is sent to the server.
[0190] Specific operation: On the device interface, the user taps or clicks to select the item they want to experience. The device then sends this information to the server.
[0191] Step 2:
[0192] Based on the event ID received by the server, the server retrieves related video data from the database.
[0193] Input: Event ID sent by the user
[0194] Output: Captured video data
[0195] Specific operation: The server queries the database to retrieve video data associated with the specified event ID.
[0196] Step 3:
[0197] The server converts the acquired video data to match the user's viewpoint.
[0198] Input: Acquired video data, user's viewpoint information
[0199] Output: Converted video data
[0200] Specific operation: The server uses the user's head tracking data to convert the video data to match the appropriate viewpoint in real time.
[0201] Step 4:
[0202] The server transmits the converted video data to the terminal.
[0203] Input: Converted video data
[0204] Output: Video data sent to the device
[0205] Specific operation: The server streams the converted video data to the terminal via the network.
[0206] Step 5:
[0207] The video data received by the terminal is displayed on a head-mounted display.
[0208] Input: Video data sent from the server
[0209] Output: Image displayed on the head-mounted display
[0210] Specific operation: The terminal outputs video data to a head-mounted display, which the user experiences visually.
[0211] Step 6:
[0212] Users interact with an AI chatbot while visually experiencing the video using a head-mounted display.
[0213] Input: User questions and input
[0214] Output: Answers and explanations from the AI chatbot
[0215] How it works: During the experience, users input questions via voice or text, and the AI chatbot provides information in response to those questions in real time.
[0216] Step 7:
[0217] The AI chatbot provides detailed explanations to the user based on the database.
[0218] Input: User questions, database information
[0219] Output: Detailed explanation from the AI chatbot
[0220] Specific operation: The AI chatbot analyzes the user's question, retrieves relevant information from the database, and provides the user with appropriate explanations.
[0221] Through these steps, users can experience historical events and scientific phenomena in an immersive way within a physical store and receive detailed explanations.
[0222] Furthermore, an emotion engine that estimates the emotion of the user may be combined. That is, the identification processing unit 290 may estimate the emotion of the user using the emotion identification model 59, and perform identification processing using the emotion of the user.
[0223] An embodiment for implementing the present invention includes the following elements.
[0224] 1. Server
[0225] Build an emotion engine to recognize user emotions.
[0226] Based on the user's selection, relevant video data is extracted and transformed to suit the user's perspective.
[0227] It provides an AI chatbot that acts as a guide and provides detailed explanations to users.
[0228] The system will incorporate a means to recognize the user's emotions and customize the images and explanations accordingly.
[0229] It is equipped with sensors and cameras to recognize the user's emotions.
[0230] The user's selection is transmitted to the server and related video data is received.
[0231] The received video data is displayed on a head-mounted display.
[0232] Incorporate a means to recognise user emotions and tailor the travel experience accordingly.
[0233] 3. Users
[0234] By operating the terminal, users select the historical event or scientific phenomenon they would like to virtually travel through.
[0235] Wearing a head-mounted display, participants visually experience a virtual site and central area.
[0236] Interact with an AI chatbot that acts as your guide and receives detailed explanations.
[0237] The system recognizes the user's emotions through sensors and cameras and customizes the images and explanations accordingly.
[0238] The above is an example of an embodiment of the present invention. The server provides the emotion engine construction and customization function, the terminal recognizes emotions and displays images, and the user selects and experiences information and interacts with the AI chatbot. This makes it possible to provide a travel experience customized to the user's emotions.
[0239] The process flow will be explained below.
[0240] Step 1: The server customizes the video data incorporating the emotion engine.
[0241] The server builds an emotion engine to recognize the user's emotion.
[0242] Based on the user's selection, relevant video data is extracted and transformed to suit the user's perspective.
[0243] The server uses an emotion engine to recognize the user's emotions and customize the video and explanations accordingly.
[0244] The customized video data is transmitted to the terminal.
[0245] Step 2: The device receives and displays the customized video data
[0246] The terminal receives the customized video data transmitted from the server.
[0247] The received video data is displayed on a head-mounted display.
[0248] The user wears a head-mounted display and visually experiences customized images.
[0249] Step 3: User emotions are recognized and the travel experience is tailored
[0250] The device uses sensors and cameras to recognize the user's emotions.
[0251] The user's emotion data is sent from the terminal to a server and analyzed.
[0252] The server tailors the travel experience based on the user's emotions, providing customized footage and commentary.
[0253] Users will enjoy a travel experience that is tailored to their emotions and receive detailed instructions through interactions with AI chatbots.
[0254] This is the flow of the program's processing. The server incorporates an emotion engine to customize the video data. The terminal receives and displays the customized video data, recognizes the user's emotions, and adjusts the travel experience. The user can enjoy a customized experience and receive detailed explanations through dialogue with the AI chatbot.
[0255] Example 2
[0256] Next, a description will be given of Example 2. In the following description, the data processing device 12 is referred to as a "server" and the smart device 14 is referred to as a "terminal."
[0257] Current virtual reality systems can provide video data on historical events or scientific phenomena selected by the user, but they lack the ability to dynamically adjust the video and explanation according to the user's emotions, making the experience uniform for each individual user and making true customization difficult. In addition, even if an AI chatbot provides an interactive explanation, the content is fixed and cannot immediately respond to the user's interests and emotions. For this reason, it is necessary to provide a highly personalized virtual travel experience that corresponds to the individual emotions and interests of each user.
[0258] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0259] In this invention, the server includes a means for acquiring video data related to a historical event or a scientific phenomenon selected by a user, a means for converting the acquired video data to match the user's viewpoint and displaying it, a means for customizing the video data using an emotion engine for recognizing the user's emotions, a means for interacting with an AI chatbot that acts as a guide and provides the user with a detailed explanation of the selected historical event or scientific phenomenon, and a means for dynamically adjusting the video data and the explanation content according to changes in the user's emotions. This enables a highly customized virtual travel experience according to the user's emotions and interests.
[0260] "Video Data" is digital data containing visual information related to a historical event or scientific phenomenon selected by a user.
[0261] An "emotion engine" is a software or hardware function for recognizing a user's emotional state by analyzing the user's facial expressions and behavior.
[0262] "Transforming to fit the user's perspective" means adjusting visual information in real time to accommodate the user's perspective and movements as they experience the site or center.
[0263] An "AI chatbot" is a program that uses artificial intelligence to converse with users in natural language and act as a virtual guide, providing specific information and detailed explanations.
[0264] "Dynamic adjustment" refers to the process of instantly changing or adapting video data and narrative content based on the user's real-time emotions and reactions.
[0265] A "head-mounted display" is a device worn by a user to visually experience virtual reality environments and images.
[0266] "Database" means a collection of information that effectively stores and manages additional information related to a historical event or scientific phenomenon selected by a User.
[0267] "Customization" means tailoring and modifying the experience and information provided to suit a user's specific needs and circumstances, based on their individual feelings and interests.
[0268] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0269] The present invention relates to a system for providing a virtual travel experience customized according to a user's emotions. A specific embodiment of the system will be described below.
[0270] 1. Hardware and Software Used
[0271] In this system, a server, a terminal, and a user each play a role. The server is a computer system that includes an emotion engine, a database, and an AI chatbot. The terminal is a device equipped with sensors, a camera, and a head-mounted display to recognize the user's emotions. The emotion engine uses a general emotion analysis API (e.g., a natural language processing engine), and the AI chatbot uses a general generative AI model (e.g., a natural language generation model).
[0272] Specific hardware examples include a "depth camera" for emotion recognition and "VR goggles" for the head-mounted display. The server is operated on a "cloud computing platform" and data storage uses a "cloud storage service." Specific software examples include a "natural language processing API" for the emotion engine and a "natural language generation API" for the generative AI model.
[0273] 2. Program Processing
[0274] The server retrieves video data related to historical events or scientific phenomena selected by the user from the cloud storage. It uses a 3D engine (e.g., a graphics engine) to transform the video data to fit the user's viewpoint. The server then uses an emotion engine to analyze the user's emotions and customizes the video data based on the analysis results.
[0275] For example, if a user selects "Age of Dinosaurs," and the server determines the user's emotion as "joy" using the emotion engine, the video will be centered around playful scenes of dinosaurs.
[0276] Meanwhile, the device is equipped with sensors and cameras to recognize the user's emotions in real time. The device decodes the video data received from the server and displays it on the user's head-mounted display. While the user is experiencing the virtual space, the device continuously transmits emotional data to the server, and the server dynamically adjusts the video and explanations based on this data.
[0277] Moreover, the AI chatbot acts as a guide, providing detailed descriptions: when the user asks, "What is this?", the chatbot responds, "This dinosaur was a Tyrannosaurus, and it mainly ate meat."
[0278] Examples
[0279] The user selects "Dinosaur Era." The server uses an emotion engine to analyze the user's emotions and determines that the user is "excited." Based on this, the server retrieves a dynamic dinosaur scene and places it in 3D space. The device displays this video data on a head-mounted display, and the user is immersed in the world of dinosaurs in real time. Furthermore, depending on the user's excitement, further interactive scenes are pushed from the server.
[0280] Examples of prompt statements
[0281] "Provide action-packed footage from the age of dinosaurs for your excited users."
[0282] The system enables a highly customized virtual travel experience based on the user's emotions and interests.
[0283] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0284] Step 1:
[0285] The user selects a virtual travel theme on the terminal.
[0286] As a specific operation, the user operates the terminal interface and selects "Dinosaur Age."
[0287] Input: User's theme selection (e.g. "Age of Dinosaurs")
[0288] Output: The selected theme ID.
[0289] Step 2:
[0290] The terminal sends the user's request to the server.
[0291] A request data with the selected theme ID is created and sent to the server via Wi-Fi.
[0292] Input: Selected Theme ID (e.g. "Age of Dinosaurs")
[0293] Output: A request containing theme selection information.
[0294] Step 3:
[0295] The server acquires emotion data for recognizing the user's emotion.
[0296] Emotion data sent from the device is received and analyzed by the emotion engine.
[0297] Input: Emotion data sent from the device
[0298] Output: Analyzed emotion information (e.g. "joy")
[0299] Step 4:
[0300] The server extracts the relevant video data and transforms it to suit the user's perspective.
[0301] Video data related to the "Dinosaur Age" is retrieved from cloud storage and converted to fit the user's perspective using the Unity (registered trademark) 3D engine.
[0302] Input: Theme selection information, analyzed emotion information
[0303] Output: Converted video data
[0304] Step 5:
[0305] The server transmits the converted video data to the terminal.
[0306] Using streaming technology, the converted video data is transferred to the terminal in real time.
[0307] Input: Converted video data
[0308] Output: Video data sent to the device
[0309] Step 6:
[0310] The terminal displays the video data on a head-mounted display.
[0311] The received video data is decoded and displayed on a head-mounted display, allowing the user to experience a virtual journey.
[0312] Input: Received video data
[0313] Output: Image displayed on a head-mounted display
[0314] Step 7:
[0315] The terminal continuously transmits the user's emotion data to the server.
[0316] The data obtained from the emotion sensor is continuously sent to the server in real time.
[0317] Input: Real-time emotion data
[0318] Output: Continuous emotion data sent to the server.
[0319] Step 8:
[0320] The server provides detailed instructions to the user through an AI chatbot.
[0321] Depending on the user's movements and questions, the AI chatbot will provide detailed explanations, for example, "This dinosaur is a Tyrannosaurus."
[0322] Input: User questions, real-time sentiment data
[0323] Output: Detailed explanation by AI chatbot
[0324] Step 9:
[0325] The server customizes images and explanations in real time according to the user's emotions.
[0326] Based on the analyzed emotion information, the video and explanation are dynamically adjusted. For example, if the user feels "surprise", an interesting additional scene is provided.
[0327] Input: Continuous emotion data, current video and description information
[0328] Output: Real-time customized video and explanations
[0329] This provides a highly customized virtual travel experience that responds to the user's emotions and interests.
[0330] (Application example 2)
[0331] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0332] In modern factories, the emotions and stress levels of workers often have a significant impact on work efficiency and quality. However, conventional factory robots and management systems operate at a uniform work speed and in a uniform manner without taking into account the emotional state of workers, which causes problems of increased stress and fatigue among workers. Furthermore, since they do not respond to the individual conditions of workers, it is difficult to maximize the overall work efficiency and quality. Similarly, there has been a lack of technology that recognizes emotions and stress levels in real time and automatically adjusts based on them. To solve these issues, there is a demand for a system that recognizes the user's emotions in real time and automatically adjusts work based on them.
[0333] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0334] In this invention, the server includes a means for recognizing the user's emotions, a means for acquiring related video data based on information selected by the user, and a means for converting and displaying the acquired video data in accordance with the user's viewpoint, thereby making it possible to provide a customized experience and adjust the working environment according to the user's emotions.
[0335] A "user" is a person who interacts with the system and provides affective states and selections.
[0336] The "emotion engine" is a mechanism that analyzes data obtained from external sensors and cameras and recognizes the user's emotional state.
[0337] "Video data" is visual information about historical events or scientific phenomena, and is digital content provided to users.
[0338] An "AI chatbot" is an artificial intelligence-based dialogue system that interacts with users and provides detailed explanations and follow-ups.
[0339] A "head-mounted display" is a display device that is worn by a user to visually experience real-world events or scientific phenomena.
[0340] "Robot control means" refers to a control mechanism for automatically adjusting the speed and method of work based on the user's emotional state.
[0341] A "database" is an information management system that accumulates information about historical events and scientific phenomena and provides it to users.
[0342] A "sensor" is a hardware device for detecting a user's emotional state.
[0343] A "camera" is a photographing device for capturing the user's facial expressions and movements.
[0344] "Work speed" is the speed at which the robot progresses through a task, and is adjusted based on the user's emotional state.
[0345] In order to implement this invention, the server, the terminal, and the user each play their respective roles, thereby realizing a customized experience and adjustment of the working environment according to the user's emotions.
[0346] The server builds an emotion engine and provides a function to recognize the user's emotions. Specifically, the server analyzes data acquired from external sensors and cameras, and determines the user's emotional state based on the user's facial expressions and heart rate. Based on the acquired emotion data, the server customizes the user's experience and tasks in real time.
[0347] The device is equipped with sensors and cameras to recognize the user's emotions and transmits the acquired emotional data to a server. It also receives related video data based on the information selected by the user and displays it on the head-mounted display. This allows the user to visually experience historical events and scientific phenomena. In addition, the device interacts with an AI chatbot to provide detailed explanations.
[0348] As a concrete example, consider a case where a worker on a packaging line in a factory suddenly becomes stressed. In this case, the server recognizes the worker's emotions based on data acquired from cameras and vital sensors, and automatically adjusts the robot's work speed. In addition, an AI chatbot will speak to the worker at appropriate times, encouraging them to "slow down your work pace" and refresh themselves. Conversely, if the worker is relaxed, they will continue working at their normal speed.
[0349] The hardware used is a webcam and vital sensors (to monitor heart rate and sweat gland activity), while the software uses an image processing library, emotion recognition API, and interpretive programming language, etc. This makes it possible to analyze the user's emotions in real time and provide a customized experience and work environment based on that.
[0350] An example of a prompt would be:
[0351] Please explain a system that recognizes the emotional state (stress, relaxation, etc.) of workers on a packaging line in a factory in real time and automatically adjusts the speed and method of a robot's work based on that information. Also, please describe in detail the role of an AI chatbot that interacts with the workers.
[0352] This allows for a customized experience and adjustment of the working environment according to the user's emotional state.
[0353] The flow of the specific process in the application example 2 will be described with reference to FIG.
[0354] Step 1:
[0355] Terminal (user operation)
[0356] The terminal activates the camera and vital sensors when the worker starts working. The camera captures the worker's face, and the vital sensors monitor the heart rate and sweat gland activity. This allows data related to the user's (worker's) emotional state to be obtained. The input is the captured image data and vital data, and the output is processed data to be passed to the emotion recognition model.
[0357] Step 2:
[0358] Terminal (data transmission)
[0359] The device transmits the acquired emotion data to the server in real time. The transmitted data includes captured image data and vital data. The input is the data collected from the user, and the output is a data stream transmitted to the server. The data is converted into an appropriate format for analysis on the server side.
[0360] Step 3:
[0361] Server (emotion recognition)
[0362] The server analyzes the received data and recognizes the user's emotional state using an emotion engine. Specifically, it extracts emotions from image data using image processing libraries and emotion recognition APIs, and integrates vital data. The input is the transmitted data stream, and the output is the judgment result of the emotional state. The server does this in real time.
[0363] Step 4:
[0364] Server (Emotion-based customization)
[0365] The server determines the necessary work adjustments based on the emotion recognition results. For example, if it is determined that the worker is stressed, it processes the work robot to slow down. The input is the emotion recognition result, and the output is a control command for the work robot. This control command is sent to the robot control system.
[0366] Step 5:
[0367] Terminal (Interaction with AI chatbot)
[0368] The AI chatbot on the device provides feedback according to the user's emotional state. For example, if the user is feeling stressed, the AI chatbot will say, "It's okay to slow down your work pace." The input is the emotion recognition result, and the output is the content of the chatbot's dialogue. This allows the user to receive psychological support.
[0369] Step 6:
[0370] User (feedback on experience)
[0371] The user (worker) takes actions such as adjusting the work pace or taking a break based on the feedback from the AI chatbot. Based on this feedback, the user selects an action that will improve their emotional state. The input is the feedback from the chatbot, and the output is the user's behavior change.
[0372] Through these steps, a customized experience and adjustments to the working environment can be made in real time based on the user's emotional state.
[0373] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires a voice indicating a user input for the result of the specific processing. The control unit 46A transmits the voice data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[0374] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (registered trademark) (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by making a neural network perform deep learning. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating a voice, text data indicating a text, and image data indicating an image is input. The data generation model 58 performs inference on the input inference data according to the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0375] In the above embodiment, an example was given in which the specific process was performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0376] [Second embodiment]
[0377] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0378] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0379] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a wide area network (WAN) and / or a local area network (LAN).
[0380] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0381] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs the voice according to instructions from the processor 46.
[0382] Camera 42 is a small digital camera equipped with an optical system including a lens, an aperture, and a shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (e.g., an imaging range defined by an angle of view equivalent to the width of the field of vision of an average healthy person).
[0383] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for transmitting and receiving various types of information between the processor 46 and the processor 28 via the network 54. The transmission and reception of various types of information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is performed in a secure state.
[0384] Fig. 4 shows an example of main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0385] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32, and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0386] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0387] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50, and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0388] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal".
[0389] An embodiment for implementing the present invention includes the following elements.
[0390] 1. Server
[0391] Build a database to capture video data related to historical events and scientific phenomena.
[0392] Based on the user's selection, relevant video data is extracted and transformed to suit the user's perspective.
[0393] It provides an AI chatbot that acts as a guide and provides detailed explanations to users.
[0394] 2. Terminal
[0395] Video data is received from the server based on a user-selected historical event or scientific phenomenon.
[0396] The received video data is displayed on a head-mounted display.
[0397] It accepts user operations and sends the selected information to the server.
[0398] 3. Users
[0399] By operating the terminal, users select the historical event or scientific phenomenon they would like to virtually travel through.
[0400] Wearing a head-mounted display, participants visually experience a virtual site and central area.
[0401] Interact with an AI chatbot that acts as your guide and receives detailed explanations.
[0402] The above is an example of an embodiment of the present invention. The server acquires and converts video data and provides an AI chatbot, the terminal receives and displays the video data, and the user selects and experiences information and interacts with the AI chatbot. This allows the user to take a virtual trip and gain a deeper understanding of historical events and scientific phenomena.
[0403] The process flow will be explained below.
[0404] Step 1: The server acquires and converts the video data.
[0405] The server retrieves video data related to historical events or scientific phenomena selected by the user.
[0406] The acquired video data is extracted from a database in the server.
[0407] The server converts the captured video data to match the user's viewpoint, allowing the user to visually experience a virtual site or center.
[0408] Step 2: The device receives and displays the video data
[0409] The terminal receives the video data transmitted from the server.
[0410] The received video data is displayed on a head-mounted display inside the terminal.
[0411] Users can wear a head-mounted display and visually experience a virtual site or central area.
[0412] Step 3: User makes information selections and experiences
[0413] The user operates the terminal to select the historical event or scientific phenomenon about which they would like to virtually travel.
[0414] The selected information is transmitted from the terminal to the server.
[0415] Based on the user's selection, the server retrieves the relevant video data, converts it and sends it to the terminal.
[0416] Step 4: User interacts with AI chatbot and receives detailed instructions
[0417] Users interact with the AI chatbot through a head-mounted display.
[0418] The AI chatbot provides users with detailed explanations of selected historical events or scientific phenomena.
[0419] Users can gain deeper understanding through dialogue with AI chatbots.
[0420] This is the flow of the program's processing. The server acquires and converts the video data, and the terminal receives and displays the video data. The user can select and experience information, and receive detailed explanations through dialogue with the AI chatbot.
[0421] Example 1
[0422] Next, a description will be given of Example 1. In the following description, the data processing device 12 is referred to as a "server" and the smart glasses 214 are referred to as a "terminal".
[0423] Currently, there are only a limited number of systems that allow users to visually experience historical events or scientific phenomena, making it difficult for users to gain a deep understanding. In addition, existing systems are unable to convert in real time to match the user's perspective, which results in a lack of immersion. Furthermore, there are insufficient guides to provide detailed information, and information supplementation is not adequate.
[0424] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0425] In this invention, the server includes a means for constructing a database for acquiring video data, a means for extracting video data based on a specific event or phenomenon selected by the user, a means for converting the extracted video data to match the user's viewpoint and displaying it, and a means for interacting with an AI chatbot as a guide that provides detailed explanations to the user based on a request provided by the user, thereby enabling the user to gain a deeper understanding of historical events and scientific phenomena through a highly immersive virtual journey.
[0426] A "database" is a collection of data in which information is systematically collected, stored, and made available for rapid retrieval.
[0427] "Video data" is digital or analog data that contains visual information, and is content that includes moving images and still images.
[0428] "Transformation to fit the viewpoint" refers to the process of adjusting the original video data according to the user's field of view and visual needs, and displaying it appropriately.
[0429] An "AI chatbot as a guide" is an automated dialogue system that uses artificial intelligence to provide appropriate information in response to user questions.
[0430] A "visual display device" is a device that allows a user to visually experience video data, and primarily refers to a head-mounted display or a monitor.
[0431] "Additional information" is supplemental information about the event or phenomenon the user is experiencing, including detailed explanations and background knowledge.
[0432] A "request" refers to the act of a user sending a specific request or instruction to the system.
[0433] The embodiment for implementing the present invention includes the following elements.
[0434] 1. Server configuration and roles:
[0435] The server builds a database to acquire and store video data related to historical events and scientific phenomena. Specifically, the server manages a huge amount of video content and extracts relevant video data based on a specific event or phenomenon selected by the user. For example, the database can include "video of pyramid construction in ancient Egypt" and "video on the development of scientific theories." Furthermore, the server converts the extracted video data to match the user's viewpoint. This conversion is intended to enable the user to view the video from a 360-degree viewpoint. The server also provides an AI chatbot as a guide, providing detailed explanations for requests provided by the user.
[0436] 2. Terminal configuration and roles:
[0437] The terminal receives video data from the server based on historical events or scientific phenomena selected by the user. The received video data is displayed on a visual display device (e.g., a head-mounted display). A specific example is a device such as a VR headset. The terminal accepts user operations and transmits the operation contents to the server. For example, if the user requests to zoom in on a particular scene or to know more about a particular object, the request is transmitted to the server.
[0438] 3. User interaction and experience:
[0439] Users can operate the device to experience a virtual journey. For example, a user can select a specific historical event or scientific phenomenon, such as "the discovery of Machu Picchu" or "scientific experiment in a laboratory." After that, the user puts on the head-mounted display and visually experiences the selected scene from a 360-degree perspective. Furthermore, if the user wants to know more information, they can interact with an AI chatbot that acts as a guide. For example, if the user asks, "What were the most commonly sold products in this ancient market?", the AI chatbot will answer, "Grains and olive oil were commonly traded."
[0440] Examples:
[0441] As a concrete example, consider the case where the user selects "Darwin's Voyage of the Beagle." In this case, the user receives the video data from the server through the terminal and it is displayed on the head-mounted display. When the user asks a question about the plants and animals discovered during Darwin's voyage, the AI chatbot will explain the details.
[0442] Examples of prompts for a generative AI model include:
[0443] "What were the most commonly sold items in ancient Roman markets?"
[0444] "I would like to know more about what Darwin discovered on his Beagle voyage."
[0445] "I want to see an episode about the discovery of the theory of relativity."
[0446] This allows users to take a virtual journey and gain a deeper understanding of historical events or scientific phenomena.
[0447] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0448] Step 1:
[0449] Database construction (server)
[0450] The server creates a database that systematically collects and stores information.
[0451] Input: Video content related to historical events or scientific phenomena.
[0452] Data processing: Classify video content by category and create an index.
[0453] Output: The indexed database.
[0454] Specific operation: The server collects, for example, footage of "scientific experiments" or "ancient civilizations" from the Internet, adds metadata to them, and stores them in a database.
[0455] Step 2:
[0456] Receiving user request (terminal)
[0457] The terminal provides an interface for the user to select a particular event or phenomenon.
[0458] Input: A request for a user-selected event or phenomenon.
[0459] Data processing: Analyze the request content and send it to the server to extract the necessary data.
[0460] Output: The parsed request data.
[0461] Specific operation: When a user selects "Darwin's Voyage of the Beagle," the request is sent from the terminal to the server.
[0462] Step 3:
[0463] Extraction and conversion of video data (server)
[0464] The server extracts relevant video data from a database based on the request.
[0465] Input: The user's request data.
[0466] Data processing: Based on the request, relevant video data is searched for in the database and the video is transformed to suit the user's perspective.
[0467] Output: Video data transformed to match the user's viewpoint.
[0468] Specific operation: The server retrieves footage, for example, of "Darwin's Voyage of the Beagle" from a database and converts the footage into a 360-degree view.
[0469] Step 4:
[0470] Video data transmission (server)
[0471] The server transmits the converted video data to the terminal.
[0472] Input: Video data transformed to match the user's viewpoint.
[0473] Data processing: Encode video data into a transmittable format.
[0474] Output: The encoded video data is sent to the device.
[0475] Specific operation: The server transmits video data encoded in, for example, MP4 format to the terminal in real time.
[0476] Step 5:
[0477] Receiving video data (terminal)
[0478] The terminal receives the video data transmitted from the server.
[0479] Input: Encoded video data sent from the server.
[0480] Data processing: Decoding the encoded video data and converting it into a displayable format.
[0481] Output: Decoded video data.
[0482] Specific operation: The terminal decodes the received video data, for example in MP4 format, and stores it in its internal memory.
[0483] Step 6:
[0484] Display of video data (terminal)
[0485] The terminal displays the received video data on a visual display device.
[0486] Input: Decoded video data.
[0487] Data Processing: Converting the decoded video data into a resolution and format suitable for a visual display device.
[0488] Output: Video data displayed on a visual display device.
[0489] Specific operation: The terminal displays a 360-degree image on, for example, a head-mounted display (HMD), allowing the user to visually experience the scene.
[0490] Step 7:
[0491] Receiving user operations (terminal)
[0492] The terminal receives user operations and requests.
[0493] Input: Actions taken by the user or additional requests.
[0494] Data processing: Analyzes operation and request data and sends it to the server.
[0495] Output: Parsed operation and request data.
[0496] Specific operation: When a user makes a request such as "I would like to take a closer look at Darwin's notes," the request is sent to the server via the terminal.
[0497] Step 8:
[0498] Interaction with AI chatbot (server, user)
[0499] Based on the user's request, the server launches an AI chatbot to act as a guide and engages in a conversation.
[0500] Input: The request made by the user.
[0501] Data processing: The AI chatbot retrieves information corresponding to the request and provides the user with an appropriate answer.
[0502] Output: A detailed description provided to the user.
[0503] What it does: When a user asks, "What's in Darwin's notebooks?" the AI chatbot responds, "These contain Darwin's early thoughts on evolution."
[0504] (Application example 1)
[0505] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal".
[0506] Traditional education of historical events and scientific phenomena has often relied on static media such as textbooks and videos, and it has been difficult to actually experience them. This has made it difficult to maintain users' understanding and interest. In addition, guidebooks and human guides have limitations in their explanations, and the information is not always consistently detailed and individualized. This has created a need for an effective way for users to gain a deeper understanding of specific events and phenomena.
[0507] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0508] In this invention, the server includes a means for acquiring video data related to a historical event or a scientific phenomenon selected by a user based on the historical event or the scientific phenomenon, a means for converting the acquired video data to match the user's viewpoint and displaying it, a means for interacting with an AI chatbot that acts as a guide and provides the user with a detailed explanation of the selected historical event or the scientific phenomenon, a means for the user to experience the video data by wearing a head-mounted display in a physical store, and a means for the AI chatbot to provide a detailed explanation in real time as a guide. This allows the user to experience the historical event or scientific phenomenon in an immersive way and receive a detailed and personalized explanation from the AI chatbot.
[0509] A "user" is someone who uses the system to experience historical events or scientific phenomena.
[0510] A "historic event" refers to a specific occurrence or phenomenon that had a significant impact on human history.
[0511] A "scientific phenomenon" refers to an event that allows for the scientific understanding and analysis of laws and phenomena in the natural world.
[0512] "Video data" refers to video and image information relating to a specific event or phenomenon.
[0513] "Means of acquiring" refers to the method or mechanism for getting the required information or data into a system or device.
[0514] "Means for converting to suit the viewpoint" refers to technology that adjusts video data in an appropriate manner according to the user's viewpoint and position.
[0515] The term "means for displaying" refers to a mechanism or method for presenting video data in a form that can be visually recognized by a user.
[0516] An "AI chatbot as a guide" refers to a program that uses artificial intelligence technology to provide information to users and engage in dialogue.
[0517] "Means of dialogue" refers to the technology that allows users and AI chatbots to communicate with each other through questions and answers.
[0518] "System" refers to a set of mechanisms in which multiple means and devices work together to provide specific services and functions to users.
[0519] A "head-mounted display" refers to a display device that is worn on the head and allows the user to visually experience immersive images.
[0520] "Real-time" refers to data and information being processed and displayed instantly, without delay.
[0521] A "database" refers to a system for systematically organizing and storing specific information or data.
[0522] "Means for providing detailed explanations" refers to a method for providing specific and detailed information necessary for users to deepen their understanding.
[0523] A "physical store" refers to a sales location that has a physical presence to offer goods or services, and is a permanent commercial space.
[0524] "Means of experiencing" refers to the technology or methods that allow a user to actually experience a particular event through simulation or virtual reality.
[0525] In the embodiment of the present invention, the entire system is configured as follows: The system used by a user mainly comprises a server, a terminal, and a head mounted display.
[0526] 1. Server
[0527] The server is responsible for building a database to retrieve video data related to historical events or scientific phenomena selected by the user. The retrieved video data is then transformed to fit the user's perspective. The server also provides an AI chatbot to provide relevant detailed explanations to the user in real time.
[0528] The software used includes database access via REST API and a natural language processing engine for the AI chatbot. Specifically, the "requests" library is used to acquire video data, and a GPT series of generative AI models are used for natural language processing.
[0529] 2. Terminal
[0530] The terminal operated by the user receives the video data from the server and displays the data on the head mounted display. The terminal accepts the user's operations and transmits the information to the server.
[0531] The terminals are general-purpose smartphones or tablet devices equipped with a browser application that uses HTML5 or WebGL to receive and display video data.
[0532] 3. Head-mounted displays
[0533] The user wears a head-mounted display and visually experiences the video data transmitted from the server. The video data is updated in real time based on the user's viewpoint, providing an immersive experience.
[0534] The hardware used includes VR headsets, which track the movements of the user's head and adjust the image based on that point of view.
[0535] Examples
[0536] In the education section of the brick-and-mortar store, a visitor selects to experience "Egyptian pyramid construction." The visitor puts on a head-mounted display and is immersed in the pyramid construction site that unfolds before their eyes, witnessing workers stacking stones. In real time, an AI chatbot explains, "This site was built around 2500 BC," and provides additional details, "The stone used was mainly limestone."
[0537] Examples of prompt statements
[0538] "Generate guides based on video data about the construction of the Egyptian pyramids to allow visitors to virtually experience the sites."
[0539] As described above, the system of the present invention allows users to experience a specific historical event or scientific phenomenon in real time while receiving detailed explanations from an AI chatbot, which leads to deeper understanding and longer-lasting interest compared to traditional static educational methods.
[0540] The flow of the specific process in the application example 1 will be described with reference to FIG.
[0541] Step 1:
[0542] The user operates the terminal to select the historical event or scientific phenomenon they wish to experience.
[0543] Input: Event ID or theme selected by the user
[0544] Output: The selected event ID is sent to the server.
[0545] Specific operation: On the device interface, the user taps or clicks to select the item they want to experience. The device then sends this information to the server.
[0546] Step 2:
[0547] Based on the event ID received by the server, the server retrieves related video data from the database.
[0548] Input: Event ID sent by the user
[0549] Output: Captured video data
[0550] Specific operation: The server queries the database to retrieve video data associated with the specified event ID.
[0551] Step 3:
[0552] The server converts the acquired video data to match the user's viewpoint.
[0553] Input: Acquired video data, user's viewpoint information
[0554] Output: Converted video data
[0555] Specific operation: The server uses the user's head tracking data to convert the video data to match the appropriate viewpoint in real time.
[0556] Step 4:
[0557] The server transmits the converted video data to the terminal.
[0558] Input: Converted video data
[0559] Output: Video data sent to the device
[0560] Specific operation: The server streams the converted video data to the terminal via the network.
[0561] Step 5:
[0562] The video data received by the terminal is displayed on a head-mounted display.
[0563] Input: Video data sent from the server
[0564] Output: Image displayed on the head-mounted display
[0565] Specific operation: The terminal outputs video data to a head-mounted display, which the user experiences visually.
[0566] Step 6:
[0567] Users interact with an AI chatbot while visually experiencing the video using a head-mounted display.
[0568] Input: User questions and input
[0569] Output: Answers and explanations from the AI chatbot
[0570] How it works: During the experience, users input questions via voice or text, and the AI chatbot provides information in response to those questions in real time.
[0571] Step 7:
[0572] The AI chatbot provides detailed explanations to the user based on the database.
[0573] Input: User questions, database information
[0574] Output: Detailed explanation from the AI chatbot
[0575] Specific operation: The AI chatbot analyzes the user's question, retrieves relevant information from the database, and provides the user with appropriate explanations.
[0576] Through these steps, users can experience historical events and scientific phenomena in an immersive way within a physical store and receive detailed explanations.
[0577] In addition, an emotion engine that estimates the emotion of the user may be further combined. That is, the identification processing unit 290 may estimate the emotion of the user using the emotion identification model 59, and perform identification processing using the emotion of the user.
[0578] An embodiment for implementing the present invention includes the following elements.
[0579] 1. Server
[0580] Build an emotion engine to recognize user emotions.
[0581] Based on the user's selection, relevant video data is extracted and transformed to suit the user's perspective.
[0582] It provides an AI chatbot that acts as a guide and provides detailed explanations to users.
[0583] The system will incorporate a means to recognize the user's emotions and customize the images and explanations accordingly.
[0584] It is equipped with sensors and cameras to recognize the user's emotions.
[0585] The user's selection is transmitted to the server and related video data is received.
[0586] The received video data is displayed on a head-mounted display.
[0587] Incorporate a means to recognise user emotions and tailor the travel experience accordingly.
[0588] 3. Users
[0589] By operating the terminal, users select the historical event or scientific phenomenon they would like to virtually travel through.
[0590] Wearing a head-mounted display, participants visually experience a virtual site and central area.
[0591] Interact with an AI chatbot that acts as your guide and receives detailed explanations.
[0592] The system recognizes the user's emotions through sensors and cameras and customizes the images and explanations accordingly.
[0593] The above is an example of an embodiment of the present invention. The server provides the emotion engine construction and customization function, the terminal recognizes emotions and displays images, and the user selects and experiences information and interacts with the AI chatbot. This makes it possible to provide a travel experience customized to the user's emotions.
[0594] The process flow will be explained below.
[0595] Step 1: The server customizes the video data incorporating the emotion engine.
[0596] The server builds an emotion engine to recognize the user's emotion.
[0597] Based on the user's selection, relevant video data is extracted and transformed to suit the user's perspective.
[0598] The server uses an emotion engine to recognize the user's emotions and customize the video and explanations accordingly.
[0599] The customized video data is transmitted to the terminal.
[0600] Step 2: The device receives and displays the customized video data
[0601] The terminal receives the customized video data transmitted from the server.
[0602] The received video data is displayed on a head-mounted display.
[0603] The user wears a head-mounted display and visually experiences customized images.
[0604] Step 3: User emotions are recognized and the travel experience is tailored
[0605] The device uses sensors and cameras to recognize the user's emotions.
[0606] The user's emotion data is sent from the terminal to a server and analyzed.
[0607] The server tailors the travel experience based on the user's emotions, providing customized footage and commentary.
[0608] Users will enjoy a travel experience that is tailored to their emotions and receive detailed instructions through interactions with AI chatbots.
[0609] This is the flow of the program. The server incorporates the emotion engine and customizes the video data. The terminal receives the customized video data, displays it, and
[0610] It recognizes users' emotions and tailors their travel experience. Users can enjoy a customized experience and get detailed explanations through interactions with AI chatbots.
[0611] Example 2
[0612] Next, a description will be given of Example 2. In the following description, the data processing device 12 is referred to as a "server" and the smart glasses 214 are referred to as a "terminal".
[0613] Current virtual reality systems can provide video data on historical events or scientific phenomena selected by the user, but they lack the ability to dynamically adjust the video and explanation according to the user's emotions, making the experience uniform for each individual user and making true customization difficult. In addition, even if an AI chatbot provides an interactive explanation, the content is fixed and cannot immediately respond to the user's interests and emotions. For this reason, it is necessary to provide a highly personalized virtual travel experience that corresponds to the individual emotions and interests of each user.
[0614] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0615] In this invention, the server includes a means for acquiring video data related to a historical event or a scientific phenomenon selected by a user, a means for converting the acquired video data to match the user's viewpoint and displaying it, a means for customizing the video data using an emotion engine for recognizing the user's emotions, a means for interacting with an AI chatbot that acts as a guide and provides the user with a detailed explanation of the selected historical event or scientific phenomenon, and a means for dynamically adjusting the video data and the explanation content according to changes in the user's emotions. This enables a highly customized virtual travel experience according to the user's emotions and interests.
[0616] "Video Data" is digital data containing visual information related to a historical event or scientific phenomenon selected by a user.
[0617] An "emotion engine" is a software or hardware function for recognizing a user's emotional state by analyzing the user's facial expressions and behavior.
[0618] "Transforming to fit the user's perspective" means adjusting visual information in real time to accommodate the user's perspective and movements as they experience the site or center.
[0619] An "AI chatbot" is a program that uses artificial intelligence to converse with users in natural language and act as a virtual guide, providing specific information and detailed explanations.
[0620] "Dynamic adjustment" refers to the process of instantly changing or adapting video data and narrative content based on the user's real-time emotions and reactions.
[0621] A "head-mounted display" is a device worn by a user to visually experience virtual reality environments and images.
[0622] "Database" means a collection of information that effectively stores and manages additional information related to a historical event or scientific phenomenon selected by a User.
[0623] "Customization" means tailoring and modifying the experience and information provided to suit a user's specific needs and circumstances, based on their individual feelings and interests.
[0624] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0625] The present invention relates to a system for providing a virtual travel experience customized according to a user's emotions. A specific embodiment of the system will be described below.
[0626] 1. Hardware and Software Used
[0627] In this system, a server, a terminal, and a user each play a role. The server is a computer system that includes an emotion engine, a database, and an AI chatbot. The terminal is a device equipped with sensors, a camera, and a head-mounted display to recognize the user's emotions. The emotion engine uses a general emotion analysis API (e.g., a natural language processing engine), and the AI chatbot uses a general generative AI model (e.g., a natural language generation model).
[0628] Specific hardware examples include a "depth camera" for emotion recognition and "VR goggles" for the head-mounted display. The server is operated on a "cloud computing platform" and data storage uses a "cloud storage service." Specific software examples include a "natural language processing API" for the emotion engine and a "natural language generation API" for the generative AI model.
[0629] 2. Program Processing
[0630] The server retrieves video data related to historical events or scientific phenomena selected by the user from the cloud storage. It uses a 3D engine (e.g., a graphics engine) to transform the video data to fit the user's viewpoint. The server then uses an emotion engine to analyze the user's emotions and customizes the video data based on the analysis results.
[0631] For example, if a user selects "Age of Dinosaurs," and the server determines the user's emotion as "joy" using the emotion engine, the video will be centered around playful scenes of dinosaurs.
[0632] Meanwhile, the device is equipped with sensors and cameras to recognize the user's emotions in real time. The device decodes the video data received from the server and displays it on the user's head-mounted display. While the user is experiencing the virtual space, the device continuously transmits emotional data to the server, and the server dynamically adjusts the video and explanations based on this data.
[0633] Moreover, the AI chatbot acts as a guide, providing detailed descriptions: when the user asks, "What is this?", the chatbot responds, "This dinosaur was a Tyrannosaurus, and it mainly ate meat."
[0634] Examples
[0635] The user selects "Dinosaur Era." The server uses an emotion engine to analyze the user's emotions and determines that the user is "excited." Based on this, the server retrieves a dynamic dinosaur scene and places it in 3D space. The device displays this video data on a head-mounted display, and the user is immersed in the world of dinosaurs in real time. Furthermore, depending on the user's excitement, further interactive scenes are pushed from the server.
[0636] Examples of prompt statements
[0637] "Provide action-packed footage from the age of dinosaurs for your excited users."
[0638] The system enables a highly customized virtual travel experience based on the user's emotions and interests.
[0639] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0640] Step 1:
[0641] The user selects a virtual travel theme on the terminal.
[0642] As a specific operation, the user operates the terminal interface and selects "Dinosaur Age."
[0643] Input: User's theme selection (e.g. "Age of Dinosaurs")
[0644] Output: The selected theme ID.
[0645] Step 2:
[0646] The terminal sends the user's request to the server.
[0647] A request data with the selected theme ID is created and sent to the server via Wi-Fi.
[0648] Input: Selected Theme ID (e.g. "Age of Dinosaurs")
[0649] Output: A request containing theme selection information.
[0650] Step 3:
[0651] The server acquires emotion data for recognizing the user's emotion.
[0652] Emotion data sent from the device is received and analyzed by the emotion engine.
[0653] Input: Emotion data sent from the device
[0654] Output: Analyzed emotion information (e.g. "joy")
[0655] Step 4:
[0656] The server extracts the relevant video data and transforms it to suit the user's perspective.
[0657] Video data related to the "Dinosaur Age" is retrieved from cloud storage and converted to fit the user's perspective using the Unity3D engine.
[0658] Input: Theme selection information, analyzed emotion information
[0659] Output: Converted video data
[0660] Step 5:
[0661] The server transmits the converted video data to the terminal.
[0662] Using streaming technology, the converted video data is transferred to the terminal in real time.
[0663] Input: Converted video data
[0664] Output: Video data sent to the device
[0665] Step 6:
[0666] The terminal displays the video data on a head-mounted display.
[0667] The received video data is decoded and displayed on a head-mounted display, allowing the user to experience a virtual journey.
[0668] Input: Received video data
[0669] Output: Image displayed on a head-mounted display
[0670] Step 7:
[0671] The terminal continuously transmits the user's emotion data to the server.
[0672] The data obtained from the emotion sensor is continuously sent to the server in real time.
[0673] Input: Real-time emotion data
[0674] Output: Continuous emotion data sent to the server.
[0675] Step 8:
[0676] The server provides detailed instructions to the user through an AI chatbot.
[0677] Depending on the user's movements and questions, the AI chatbot will provide detailed explanations, for example, "This dinosaur is a Tyrannosaurus."
[0678] Input: User questions, real-time sentiment data
[0679] Output: Detailed explanation by AI chatbot
[0680] Step 9:
[0681] The server customizes images and explanations in real time according to the user's emotions.
[0682] Based on the analyzed emotion information, the video and explanation are dynamically adjusted. For example, if the user feels "surprise", an interesting additional scene is provided.
[0683] Input: Continuous emotion data, current video and description information
[0684] Output: Real-time customized video and explanations
[0685] This provides a highly customized virtual travel experience that responds to the user's emotions and interests.
[0686] (Application example 2)
[0687] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal".
[0688] In modern factories, the emotions and stress levels of workers often have a significant impact on work efficiency and quality. However, conventional factory robots and management systems operate at a uniform work speed and in a uniform manner without taking into account the emotional state of workers, which causes problems of increased stress and fatigue among workers. Furthermore, since they do not respond to the individual conditions of workers, it is difficult to maximize the overall work efficiency and quality. Similarly, there has been a lack of technology that recognizes emotions and stress levels in real time and automatically adjusts based on them. To solve these issues, there is a demand for a system that recognizes the user's emotions in real time and automatically adjusts work based on them.
[0689] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0690] In this invention, the server includes a means for recognizing the user's emotions, a means for acquiring related video data based on information selected by the user, and a means for converting and displaying the acquired video data in accordance with the user's viewpoint, thereby making it possible to provide a customized experience and adjust the working environment according to the user's emotions.
[0691] A "user" is a person who interacts with the system and provides affective states and selections.
[0692] The "emotion engine" is a mechanism that analyzes data obtained from external sensors and cameras and recognizes the user's emotional state.
[0693] "Video data" is visual information about historical events or scientific phenomena, and is digital content provided to users.
[0694] An "AI chatbot" is an artificial intelligence-based dialogue system that interacts with users and provides detailed explanations and follow-ups.
[0695] A "head-mounted display" is a display device that is worn by a user to visually experience real-world events or scientific phenomena.
[0696] "Robot control means" refers to a control mechanism for automatically adjusting the speed and method of work based on the user's emotional state.
[0697] A "database" is an information management system that accumulates information about historical events and scientific phenomena and provides it to users.
[0698] A "sensor" is a hardware device for detecting a user's emotional state.
[0699] A "camera" is a photographing device for capturing the user's facial expressions and movements.
[0700] "Work speed" is the speed at which the robot progresses through a task, and is adjusted based on the user's emotional state.
[0701] In order to implement this invention, the server, the terminal, and the user each play their respective roles, thereby realizing a customized experience and adjustment of the working environment according to the user's emotions.
[0702] The server builds an emotion engine and provides a function to recognize the user's emotions. Specifically, the server analyzes data acquired from external sensors and cameras, and determines the user's emotional state based on the user's facial expressions and heart rate. Based on the acquired emotion data, the server customizes the user's experience and tasks in real time.
[0703] The device is equipped with sensors and cameras to recognize the user's emotions and transmits the acquired emotional data to a server. It also receives related video data based on the information selected by the user and displays it on the head-mounted display. This allows the user to visually experience historical events and scientific phenomena. In addition, the device interacts with an AI chatbot to provide detailed explanations.
[0704] As a concrete example, consider a case where a worker on a packaging line in a factory suddenly becomes stressed. In this case, the server recognizes the worker's emotions based on data acquired from cameras and vital sensors, and automatically adjusts the robot's work speed. In addition, an AI chatbot will speak to the worker at appropriate times, encouraging them to "slow down your work pace" and refresh themselves. Conversely, if the worker is relaxed, they will continue working at their normal speed.
[0705] The hardware used is a webcam and vital sensors (to monitor heart rate and sweat gland activity), while the software uses an image processing library, emotion recognition API, and interpretive programming language, etc. This makes it possible to analyze the user's emotions in real time and provide a customized experience and work environment based on that.
[0706] An example of a prompt would be:
[0707] Please explain a system that recognizes the emotional state (stress, relaxation, etc.) of workers on a packaging line in a factory in real time and automatically adjusts the speed and method of a robot's work based on that information. Also, please describe in detail the role of an AI chatbot that interacts with the workers.
[0708] This allows for a customized experience and adjustment of the working environment according to the user's emotional state.
[0709] The flow of the specific process in the application example 2 will be described with reference to FIG.
[0710] Step 1:
[0711] Terminal (user operation)
[0712] The terminal activates the camera and vital sensors when the worker starts working. The camera captures the worker's face, and the vital sensors monitor the heart rate and sweat gland activity. This allows data related to the user's (worker's) emotional state to be obtained. The input is the captured image data and vital data, and the output is processed data to be passed to the emotion recognition model.
[0713] Step 2:
[0714] Terminal (data transmission)
[0715] The device transmits the acquired emotion data to the server in real time. The transmitted data includes captured image data and vital data. The input is the data collected from the user, and the output is a data stream transmitted to the server. The data is converted into an appropriate format for analysis on the server side.
[0716] Step 3:
[0717] Server (emotion recognition)
[0718] The server analyzes the received data and recognizes the user's emotional state using an emotion engine. Specifically, it extracts emotions from image data using image processing libraries and emotion recognition APIs, and integrates vital data. The input is the transmitted data stream, and the output is the judgment result of the emotional state. The server does this in real time.
[0719] Step 4:
[0720] Server (Emotion-based customization)
[0721] The server determines the necessary work adjustments based on the emotion recognition results. For example, if it is determined that the worker is stressed, it processes the work robot to slow down. The input is the emotion recognition result, and the output is a control command for the work robot. This control command is sent to the robot control system.
[0722] Step 5:
[0723] Terminal (Interaction with AI chatbot)
[0724] The AI chatbot on the device provides feedback according to the user's emotional state. For example, if the user is feeling stressed, the AI chatbot will say, "It's okay to slow down your work pace." The input is the emotion recognition result, and the output is the content of the chatbot's dialogue. This allows the user to receive psychological support.
[0725] Step 6:
[0726] User (feedback on experience)
[0727] The user (worker) takes actions such as adjusting the work pace or taking a break based on the feedback from the AI chatbot. Based on this feedback, the user selects an action that will improve their emotional state. The input is the feedback from the chatbot, and the output is the user's behavior change.
[0728] Through these steps, a customized experience and adjustments to the working environment can be made in real time based on the user's emotional state.
[0729] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires a voice indicating a user input for the result of the specific processing. The control unit 46A transmits the voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[0730] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by making a neural network perform deep learning. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating a voice, text data indicating a text, and image data indicating an image is input. The data generation model 58 performs inference on the input inference data according to the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0731] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0732] [Third embodiment]
[0733] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0734] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[0735] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a wide area network (WAN) and / or a local area network (LAN).
[0736] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0737] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs the voice according to instructions from the processor 46.
[0738] Camera 42 is a small digital camera equipped with an optical system including a lens, an aperture, and a shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (e.g., an imaging range defined by an angle of view equivalent to the width of the field of vision of an average healthy person).
[0739] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for transmitting and receiving various types of information between the processor 46 and the processor 28 via the network 54. The transmission and reception of various types of information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is performed in a secure state.
[0740] Fig. 6 shows an example of main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0741] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32, and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0742] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0743] In the headset type terminal 314, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50, and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0744] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server", and the headset type terminal 314 will be referred to as the "terminal".
[0745] An embodiment for implementing the present invention includes the following elements.
[0746] 1. Server
[0747] Build a database to capture video data related to historical events and scientific phenomena.
[0748] Based on the user's selection, relevant video data is extracted and transformed to suit the user's perspective.
[0749] It provides an AI chatbot that acts as a guide and provides detailed explanations to users.
[0750] 2. Terminal
[0751] Video data is received from the server based on a user-selected historical event or scientific phenomenon.
[0752] The received video data is displayed on a head-mounted display.
[0753] It accepts user operations and sends the selected information to the server.
[0754] 3. Users
[0755] By operating the terminal, users select the historical event or scientific phenomenon they would like to virtually travel through.
[0756] Wearing a head-mounted display, participants visually experience a virtual site and central area.
[0757] Interact with an AI chatbot that acts as your guide and receives detailed explanations.
[0758] The above is an example of an embodiment of the present invention. The server acquires and converts video data and provides an AI chatbot, the terminal receives and displays the video data, and the user selects and experiences information and interacts with the AI chatbot. This allows the user to take a virtual trip and gain a deeper understanding of historical events and scientific phenomena.
[0759] The process flow will be explained below.
[0760] Step 1: The server acquires and converts the video data.
[0761] The server retrieves video data related to historical events or scientific phenomena selected by the user.
[0762] The acquired video data is extracted from a database in the server.
[0763] The server converts the captured video data to match the user's viewpoint, allowing the user to visually experience a virtual site or center.
[0764] Step 2: The device receives and displays the video data
[0765] The terminal receives the video data transmitted from the server.
[0766] The received video data is displayed on a head-mounted display inside the terminal.
[0767] Users can wear a head-mounted display and visually experience a virtual site or central area.
[0768] Step 3: User makes information selections and experiences
[0769] The user operates the terminal to select the historical event or scientific phenomenon about which they would like to virtually travel.
[0770] The selected information is transmitted from the terminal to the server.
[0771] Based on the user's selection, the server retrieves the relevant video data, converts it and sends it to the terminal.
[0772] Step 4: User interacts with AI chatbot and receives detailed instructions
[0773] Users interact with the AI chatbot through a head-mounted display.
[0774] The AI chatbot provides users with detailed explanations of selected historical events or scientific phenomena.
[0775] Users can gain deeper understanding through dialogue with AI chatbots.
[0776] This is the flow of the program's processing. The server acquires and converts the video data, and the terminal receives and displays the video data. The user can select and experience information, and receive detailed explanations through dialogue with the AI chatbot.
[0777] Example 1
[0778] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal".
[0779] Currently, there are only a limited number of systems that allow users to visually experience historical events or scientific phenomena, making it difficult for users to gain a deep understanding. In addition, existing systems are unable to convert in real time to match the user's perspective, which results in a lack of immersion. Furthermore, there are insufficient guides to provide detailed information, and information supplementation is not adequate.
[0780] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0781] In this invention, the server includes a means for constructing a database for acquiring video data, a means for extracting video data based on a specific event or phenomenon selected by the user, a means for converting the extracted video data to match the user's viewpoint and displaying it, and a means for interacting with an AI chatbot as a guide that provides detailed explanations to the user based on a request provided by the user, thereby enabling the user to gain a deeper understanding of historical events and scientific phenomena through a highly immersive virtual journey.
[0782] A "database" is a collection of data in which information is systematically collected, stored, and made available for rapid retrieval.
[0783] "Video data" is digital or analog data that contains visual information, and is content that includes moving images and still images.
[0784] "Transformation to fit the viewpoint" refers to the process of adjusting the original video data according to the user's field of view and visual needs, and displaying it appropriately.
[0785] An "AI chatbot as a guide" is an automated dialogue system that uses artificial intelligence to provide appropriate information in response to user questions.
[0786] A "visual display device" is a device that allows a user to visually experience video data, and primarily refers to a head-mounted display or a monitor.
[0787] "Additional information" is supplemental information about the event or phenomenon the user is experiencing, including detailed explanations and background knowledge.
[0788] A "request" refers to the act of a user sending a specific request or instruction to the system.
[0789] The embodiment for implementing the present invention includes the following elements.
[0790] 1. Server configuration and roles:
[0791] The server builds a database to acquire and store video data related to historical events and scientific phenomena. Specifically, the server manages a huge amount of video content and extracts relevant video data based on a specific event or phenomenon selected by the user. For example, the database can include "video of pyramid construction in ancient Egypt" and "video on the development of scientific theories." Furthermore, the server converts the extracted video data to match the user's viewpoint. This conversion is intended to enable the user to view the video from a 360-degree viewpoint. The server also provides an AI chatbot as a guide, providing detailed explanations for requests provided by the user.
[0792] 2. Terminal configuration and roles:
[0793] The terminal receives video data from the server based on historical events or scientific phenomena selected by the user. The received video data is displayed on a visual display device (e.g., a head-mounted display). A specific example is a device such as a VR headset. The terminal accepts user operations and transmits the operation contents to the server. For example, if the user requests to zoom in on a particular scene or to know more about a particular object, the request is transmitted to the server.
[0794] 3. User interaction and experience:
[0795] Users can operate the device to experience a virtual journey. For example, a user can select a specific historical event or scientific phenomenon, such as "the discovery of Machu Picchu" or "scientific experiment in a laboratory." After that, the user puts on the head-mounted display and visually experiences the selected scene from a 360-degree perspective. Furthermore, if the user wants to know more information, they can interact with an AI chatbot that acts as a guide. For example, if the user asks, "What were the most commonly sold products in this ancient market?", the AI chatbot will answer, "Grains and olive oil were commonly traded."
[0796] Examples:
[0797] As a concrete example, consider the case where the user selects "Darwin's Voyage of the Beagle." In this case, the user receives the video data from the server through the terminal and it is displayed on the head-mounted display. When the user asks a question about the plants and animals discovered during Darwin's voyage, the AI chatbot will explain the details.
[0798] Examples of prompts for a generative AI model include:
[0799] "What were the most commonly sold items in ancient Roman markets?"
[0800] "I would like to know more about what Darwin discovered on his Beagle voyage."
[0801] "I want to see an episode about the discovery of the theory of relativity."
[0802] This allows users to take a virtual journey and gain a deeper understanding of historical events or scientific phenomena.
[0803] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0804] Step 1:
[0805] Database construction (server)
[0806] The server creates a database that systematically collects and stores information.
[0807] Input: Video content related to historical events or scientific phenomena.
[0808] Data processing: Classify video content by category and create an index.
[0809] Output: The indexed database.
[0810] Specific operation: The server collects, for example, footage of "scientific experiments" or "ancient civilizations" from the Internet, adds metadata to them, and stores them in a database.
[0811] Step 2:
[0812] Receiving user request (terminal)
[0813] The terminal provides an interface for the user to select a particular event or phenomenon.
[0814] Input: A request for a user-selected event or phenomenon.
[0815] Data processing: Analyze the request content and send it to the server to extract the necessary data.
[0816] Output: The parsed request data.
[0817] Specific operation: When a user selects "Darwin's Voyage of the Beagle," the request is sent from the terminal to the server.
[0818] Step 3:
[0819] Extraction and conversion of video data (server)
[0820] The server extracts relevant video data from a database based on the request.
[0821] Input: The user's request data.
[0822] Data processing: Based on the request, relevant video data is searched for in the database and the video is transformed to suit the user's perspective.
[0823] Output: Video data transformed to match the user's viewpoint.
[0824] Specific operation: The server retrieves footage, for example, of "Darwin's Voyage of the Beagle" from a database and converts the footage into a 360-degree view.
[0825] Step 4:
[0826] Video data transmission (server)
[0827] The server transmits the converted video data to the terminal.
[0828] Input: Video data transformed to match the user's viewpoint.
[0829] Data processing: Encode video data into a transmittable format.
[0830] Output: The encoded video data is sent to the device.
[0831] Specific operation: The server transmits video data encoded in, for example, MP4 format to the terminal in real time.
[0832] Step 5:
[0833] Receiving video data (terminal)
[0834] The terminal receives the video data transmitted from the server.
[0835] Input: Encoded video data sent from the server.
[0836] Data processing: Decoding the encoded video data and converting it into a displayable format.
[0837] Output: Decoded video data.
[0838] Specific operation: The terminal decodes the received video data, for example in MP4 format, and stores it in its internal memory.
[0839] Step 6:
[0840] Display of video data (terminal)
[0841] The terminal displays the received video data on a visual display device.
[0842] Input: Decoded video data.
[0843] Data Processing: Converting the decoded video data into a resolution and format suitable for a visual display device.
[0844] Output: Video data displayed on a visual display device.
[0845] Specific operation: The terminal displays a 360-degree image on, for example, a head-mounted display (HMD), allowing the user to visually experience the scene.
[0846] Step 7:
[0847] Receiving user operations (terminal)
[0848] The terminal receives user operations and requests.
[0849] Input: Actions taken by the user or additional requests.
[0850] Data processing: Analyzes operation and request data and sends it to the server.
[0851] Output: Parsed operation and request data.
[0852] Specific operation: When a user makes a request such as "I would like to take a closer look at Darwin's notes," the request is sent to the server via the terminal.
[0853] Step 8:
[0854] Interaction with AI chatbot (server, user)
[0855] Based on the user's request, the server launches an AI chatbot to act as a guide and engages in a conversation.
[0856] Input: The request made by the user.
[0857] Data processing: The AI chatbot retrieves information corresponding to the request and provides the user with an appropriate answer.
[0858] Output: A detailed description provided to the user.
[0859] What it does: When a user asks, "What's in Darwin's notebooks?" the AI chatbot responds, "These contain Darwin's early thoughts on evolution."
[0860] (Application example 1)
[0861] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0862] Traditional education of historical events and scientific phenomena has often relied on static media such as textbooks and videos, and it has been difficult to actually experience them. This has made it difficult to maintain users' understanding and interest. In addition, guidebooks and human guides have limitations in their explanations, and the information is not always consistently detailed and individualized. This has created a need for an effective way for users to gain a deeper understanding of specific events and phenomena.
[0863] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0864] In this invention, the server includes a means for acquiring video data related to a historical event or a scientific phenomenon selected by a user based on the historical event or the scientific phenomenon, a means for converting the acquired video data to match the user's viewpoint and displaying it, a means for interacting with an AI chatbot that acts as a guide and provides the user with a detailed explanation of the selected historical event or the scientific phenomenon, a means for the user to experience the video data by wearing a head-mounted display in a physical store, and a means for the AI chatbot to provide a detailed explanation in real time as a guide. This allows the user to experience the historical event or scientific phenomenon in an immersive way and receive a detailed and personalized explanation from the AI chatbot.
[0865] A "user" is someone who uses the system to experience historical events or scientific phenomena.
[0866] A "historic event" refers to a specific occurrence or phenomenon that had a significant impact on human history.
[0867] A "scientific phenomenon" refers to an event that allows for the scientific understanding and analysis of laws and phenomena in the natural world.
[0868] "Video data" refers to video and image information relating to a specific event or phenomenon.
[0869] "Means of acquiring" refers to the method or mechanism for getting the required information or data into a system or device.
[0870] "Means for converting to suit the viewpoint" refers to technology that adjusts video data in an appropriate manner according to the user's viewpoint and position.
[0871] The term "means for displaying" refers to a mechanism or method for presenting video data in a form that can be visually recognized by a user.
[0872] An "AI chatbot as a guide" refers to a program that uses artificial intelligence technology to provide information to users and engage in dialogue.
[0873] "Means of dialogue" refers to the technology that allows users and AI chatbots to communicate with each other through questions and answers.
[0874] "System" refers to a set of mechanisms in which multiple means and devices work together to provide specific services and functions to users.
[0875] A "head-mounted display" refers to a display device that is worn on the head and allows the user to visually experience immersive images.
[0876] "Real-time" refers to data and information being processed and displayed instantly, without delay.
[0877] A "database" refers to a system for systematically organizing and storing specific information or data.
[0878] "Means for providing detailed explanations" refers to a method for providing specific and detailed information necessary for users to deepen their understanding.
[0879] A "physical store" refers to a sales location that has a physical presence to offer goods or services, and is a permanent commercial space.
[0880] "Means of experiencing" refers to the technology or methods that allow a user to actually experience a particular event through simulation or virtual reality.
[0881] In the embodiment of the present invention, the entire system is configured as follows: The system used by a user mainly comprises a server, a terminal, and a head mounted display.
[0882] 1. Server
[0883] The server is responsible for building a database to retrieve video data related to historical events or scientific phenomena selected by the user. The retrieved video data is then transformed to fit the user's perspective. The server also provides an AI chatbot to provide relevant detailed explanations to the user in real time.
[0884] The software used includes database access via REST API and a natural language processing engine for the AI chatbot. Specifically, the "requests" library is used to acquire video data, and a GPT series of generative AI models are used for natural language processing.
[0885] 2. Terminal
[0886] The terminal operated by the user receives the video data from the server and displays the data on the head mounted display. The terminal accepts the user's operations and transmits the information to the server.
[0887] The terminals are general-purpose smartphones or tablet devices equipped with a browser application that uses HTML5 or WebGL to receive and display video data.
[0888] 3. Head-mounted displays
[0889] The user wears a head-mounted display and visually experiences the video data transmitted from the server. The video data is updated in real time based on the user's viewpoint, providing an immersive experience.
[0890] The hardware used includes VR headsets, which track the movements of the user's head and adjust the image based on that point of view.
[0891] Examples
[0892] In the education section of the brick-and-mortar store, a visitor selects to experience "Egyptian pyramid construction." The visitor puts on a head-mounted display and is immersed in the pyramid construction site that unfolds before their eyes, witnessing workers stacking stones. In real time, an AI chatbot explains, "This site was built around 2500 BC," and provides additional details, "The stone used was mainly limestone."
[0893] Examples of prompt statements
[0894] "Generate guides based on video data about the construction of the Egyptian pyramids to allow visitors to virtually experience the sites."
[0895] As described above, the system of the present invention allows users to experience a specific historical event or scientific phenomenon in real time while receiving detailed explanations from an AI chatbot, which leads to deeper understanding and longer-lasting interest compared to traditional static educational methods.
[0896] The flow of the specific process in the application example 1 will be described with reference to FIG.
[0897] Step 1:
[0898] The user operates the terminal to select the historical event or scientific phenomenon they wish to experience.
[0899] Input: Event ID or theme selected by the user
[0900] Output: The selected event ID is sent to the server.
[0901] Specific operation: On the device interface, the user taps or clicks to select the item they want to experience. The device then sends this information to the server.
[0902] Step 2:
[0903] Based on the event ID received by the server, the server retrieves related video data from the database.
[0904] Input: Event ID sent by the user
[0905] Output: Captured video data
[0906] Specific operation: The server queries the database to retrieve video data associated with the specified event ID.
[0907] Step 3:
[0908] The server converts the acquired video data to match the user's viewpoint.
[0909] Input: Acquired video data, user's viewpoint information
[0910] Output: Converted video data
[0911] Specific operation: The server uses the user's head tracking data to convert the video data to match the appropriate viewpoint in real time.
[0912] Step 4:
[0913] The server transmits the converted video data to the terminal.
[0914] Input: Converted video data
[0915] Output: Video data sent to the device
[0916] Specific operation: The server streams the converted video data to the terminal via the network.
[0917] Step 5:
[0918] The video data received by the terminal is displayed on a head-mounted display.
[0919] Input: Video data sent from the server
[0920] Output: Image displayed on the head-mounted display
[0921] Specific operation: The terminal outputs video data to a head-mounted display, which the user experiences visually.
[0922] Step 6:
[0923] Users interact with an AI chatbot while visually experiencing the video using a head-mounted display.
[0924] Input: User questions and input
[0925] Output: Answers and explanations from the AI chatbot
[0926] How it works: During the experience, users input questions via voice or text, and the AI chatbot provides information in response to those questions in real time.
[0927] Step 7:
[0928] The AI chatbot provides detailed explanations to the user based on the database.
[0929] Input: User questions, database information
[0930] Output: Detailed explanation from the AI chatbot
[0931] Specific operation: The AI chatbot analyzes the user's question, retrieves relevant information from the database, and provides the user with appropriate explanations.
[0932] Through these steps, users can experience historical events and scientific phenomena in an immersive way within a physical store and receive detailed explanations.
[0933] In addition, an emotion engine that estimates the emotion of the user may be further combined. That is, the identification processing unit 290 may estimate the emotion of the user using the emotion identification model 59, and perform identification processing using the emotion of the user.
[0934] An embodiment for implementing the present invention includes the following elements.
[0935] 1. Server
[0936] Build an emotion engine to recognize user emotions.
[0937] Based on the user's selection, relevant video data is extracted and transformed to suit the user's perspective.
[0938] It provides an AI chatbot that acts as a guide and provides detailed explanations to users.
[0939] The system will incorporate a means to recognize the user's emotions and customize the images and explanations accordingly.
[0940] It is equipped with sensors and cameras to recognize the user's emotions.
[0941] The user's selection is transmitted to the server and related video data is received.
[0942] The received video data is displayed on a head-mounted display.
[0943] Incorporate a means to recognise user emotions and tailor the travel experience accordingly.
[0944] 3. Users
[0945] By operating the terminal, users select the historical event or scientific phenomenon they would like to virtually travel through.
[0946] Wearing a head-mounted display, participants visually experience a virtual site and central area.
[0947] Interact with an AI chatbot that acts as your guide and receives detailed explanations.
[0948] The system recognizes the user's emotions through sensors and cameras and customizes the images and explanations accordingly.
[0949] The above is an example of an embodiment of the present invention. The server provides the emotion engine construction and customization function, the terminal recognizes emotions and displays images, and the user selects and experiences information and interacts with the AI chatbot. This makes it possible to provide a travel experience customized to the user's emotions.
[0950] The process flow will be explained below.
[0951] Step 1: The server customizes the video data incorporating the emotion engine.
[0952] The server builds an emotion engine to recognize the user's emotion.
[0953] Based on the user's selection, relevant video data is extracted and transformed to suit the user's perspective.
[0954] The server uses an emotion engine to recognize the user's emotions and customize the video and explanations accordingly.
[0955] The customized video data is transmitted to the terminal.
[0956] Step 2: The device receives and displays the customized video data
[0957] The terminal receives the customized video data transmitted from the server.
[0958] The received video data is displayed on a head-mounted display.
[0959] The user wears a head-mounted display and visually experiences customized images.
[0960] Step 3: User emotions are recognized and the travel experience is tailored
[0961] The device uses sensors and cameras to recognize the user's emotions.
[0962] The user's emotion data is sent from the terminal to a server and analyzed.
[0963] The server tailors the travel experience based on the user's emotions, providing customized footage and commentary.
[0964] Users will enjoy a travel experience that is tailored to their emotions and receive detailed instructions through interactions with AI chatbots.
[0965] This is the flow of the program's processing. The server incorporates an emotion engine to customize the video data. The terminal receives and displays the customized video data, recognizes the user's emotions, and adjusts the travel experience. The user can enjoy a customized experience and receive detailed explanations through dialogue with the AI chatbot.
[0966] Example 2
[0967] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal".
[0968] Current virtual reality systems can provide video data on historical events or scientific phenomena selected by the user, but they lack the ability to dynamically adjust the video and explanation according to the user's emotions, making the experience uniform for each individual user and making true customization difficult. In addition, even if an AI chatbot provides an interactive explanation, the content is fixed and cannot immediately respond to the user's interests and emotions. For this reason, it is necessary to provide a highly personalized virtual travel experience that corresponds to the individual emotions and interests of each user.
[0969] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0970] In this invention, the server includes a means for acquiring video data related to a historical event or a scientific phenomenon selected by a user, a means for converting the acquired video data to match the user's viewpoint and displaying it, a means for customizing the video data using an emotion engine for recognizing the user's emotions, a means for interacting with an AI chatbot that acts as a guide and provides the user with a detailed explanation of the selected historical event or scientific phenomenon, and a means for dynamically adjusting the video data and the explanation content according to changes in the user's emotions. This enables a highly customized virtual travel experience according to the user's emotions and interests.
[0971] "Video Data" is digital data containing visual information related to a historical event or scientific phenomenon selected by a user.
[0972] An "emotion engine" is a software or hardware function for recognizing a user's emotional state by analyzing the user's facial expressions and behavior.
[0973] "Transforming to fit the user's perspective" means adjusting visual information in real time to accommodate the user's perspective and movements as they experience the site or center.
[0974] An "AI chatbot" is a program that uses artificial intelligence to converse with users in natural language and act as a virtual guide, providing specific information and detailed explanations.
[0975] "Dynamic adjustment" refers to the process of instantly changing or adapting video data and narrative content based on the user's real-time emotions and reactions.
[0976] A "head-mounted display" is a device worn by a user to visually experience virtual reality environments and images.
[0977] "Database" means a collection of information that effectively stores and manages additional information related to a historical event or scientific phenomenon selected by a User.
[0978] "Customization" means tailoring and modifying the experience and information provided to suit a user's specific needs and circumstances, based on their individual feelings and interests.
[0979] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0980] The present invention relates to a system for providing a virtual travel experience customized according to a user's emotions. A specific embodiment of the system will be described below.
[0981] 1. Hardware and Software Used
[0982] In this system, a server, a terminal, and a user each play a role. The server is a computer system that includes an emotion engine, a database, and an AI chatbot. The terminal is a device equipped with sensors, a camera, and a head-mounted display to recognize the user's emotions. The emotion engine uses a general emotion analysis API (e.g., a natural language processing engine), and the AI chatbot uses a general generative AI model (e.g., a natural language generation model).
[0983] Specific hardware examples include a "depth camera" for emotion recognition and "VR goggles" for the head-mounted display. The server is operated on a "cloud computing platform" and data storage uses a "cloud storage service." Specific software examples include a "natural language processing API" for the emotion engine and a "natural language generation API" for the generative AI model.
[0984] 2. Program Processing
[0985] The server retrieves video data related to historical events or scientific phenomena selected by the user from the cloud storage. It uses a 3D engine (e.g., a graphics engine) to transform the video data to fit the user's viewpoint. The server then uses an emotion engine to analyze the user's emotions and customizes the video data based on the analysis results.
[0986] For example, if a user selects "Age of Dinosaurs," and the server determines the user's emotion as "joy" using the emotion engine, the video will be centered around playful scenes of dinosaurs.
[0987] Meanwhile, the device is equipped with sensors and cameras to recognize the user's emotions in real time. The device decodes the video data received from the server and displays it on the user's head-mounted display. While the user is experiencing the virtual space, the device continuously transmits emotional data to the server, and the server dynamically adjusts the video and explanations based on this data.
[0988] Moreover, the AI chatbot acts as a guide, providing detailed descriptions: when the user asks, "What is this?", the chatbot responds, "This dinosaur was a Tyrannosaurus, and it mainly ate meat."
[0989] Examples
[0990] The user selects "Dinosaur Era." The server uses an emotion engine to analyze the user's emotions and determines that the user is "excited." Based on this, the server retrieves a dynamic dinosaur scene and places it in 3D space. The device displays this video data on a head-mounted display, and the user is immersed in the world of dinosaurs in real time. Furthermore, depending on the user's excitement, further interactive scenes are pushed from the server.
[0991] Examples of prompt statements
[0992] "Provide action-packed footage from the age of dinosaurs for your excited users."
[0993] The system enables a highly customized virtual travel experience based on the user's emotions and interests.
[0994] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0995] Step 1:
[0996] The user selects a virtual travel theme on the terminal.
[0997] As a specific operation, the user operates the terminal interface and selects "Dinosaur Age."
[0998] Input: User's theme selection (e.g. "Age of Dinosaurs")
[0999] Output: The selected theme ID.
[1000] Step 2:
[1001] The terminal sends the user's request to the server.
[1002] A request data with the selected theme ID is created and sent to the server via Wi-Fi.
[1003] Input: Selected Theme ID (e.g. "Age of Dinosaurs")
[1004] Output: A request containing theme selection information.
[1005] Step 3:
[1006] The server acquires emotion data for recognizing the user's emotion.
[1007] Emotion data sent from the device is received and analyzed by the emotion engine.
[1008] Input: Emotion data sent from the device
[1009] Output: Analyzed emotion information (e.g. "joy")
[1010] Step 4:
[1011] The server extracts the relevant video data and transforms it to suit the user's perspective.
[1012] Video data related to the "Dinosaur Age" is retrieved from cloud storage and converted to fit the user's perspective using the Unity3D engine.
[1013] Input: Theme selection information, analyzed emotion information
[1014] Output: Converted video data
[1015] Step 5:
[1016] The server transmits the converted video data to the terminal.
[1017] Using streaming technology, the converted video data is transferred to the terminal in real time.
[1018] Input: Converted video data
[1019] Output: Video data sent to the device
[1020] Step 6:
[1021] The terminal displays the video data on a head-mounted display.
[1022] The received video data is decoded and displayed on a head-mounted display, allowing the user to experience a virtual journey.
[1023] Input: Received video data
[1024] Output: Image displayed on a head-mounted display
[1025] Step 7:
[1026] The terminal continuously transmits the user's emotion data to the server.
[1027] The data obtained from the emotion sensor is continuously sent to the server in real time.
[1028] Input: Real-time emotion data
[1029] Output: Continuous emotion data sent to the server.
[1030] Step 8:
[1031] The server provides detailed instructions to the user through an AI chatbot.
[1032] Depending on the user's movements and questions, the AI chatbot will provide detailed explanations, for example, "This dinosaur is a Tyrannosaurus."
[1033] Input: User questions, real-time sentiment data
[1034] Output: Detailed explanation by AI chatbot
[1035] Step 9:
[1036] The server customizes images and explanations in real time according to the user's emotions.
[1037] Based on the analyzed emotion information, the video and explanation are dynamically adjusted. For example, if the user feels "surprise", an interesting additional scene is provided.
[1038] Input: Continuous emotion data, current video and description information
[1039] Output: Real-time customized video and explanations
[1040] This provides a highly customized virtual travel experience that responds to the user's emotions and interests.
[1041] (Application example 2)
[1042] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server", and the headset type terminal 314 will be referred to as a "terminal".
[1043] In modern factories, the emotions and stress levels of workers often have a significant impact on work efficiency and quality. However, conventional factory robots and management systems operate at a uniform work speed and in a uniform manner without taking into account the emotional state of workers, which causes problems of increased stress and fatigue among workers. Furthermore, since they do not respond to the individual conditions of workers, it is difficult to maximize the overall work efficiency and quality. Similarly, there has been a lack of technology that recognizes emotions and stress levels in real time and automatically adjusts based on them. To solve these issues, there is a demand for a system that recognizes the user's emotions in real time and automatically adjusts work based on them.
[1044] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1045] In this invention, the server includes a means for recognizing the user's emotions, a means for acquiring related video data based on information selected by the user, and a means for converting and displaying the acquired video data in accordance with the user's viewpoint, thereby making it possible to provide a customized experience and adjust the working environment according to the user's emotions.
[1046] A "user" is a person who interacts with the system and provides affective states and selections.
[1047] The "emotion engine" is a mechanism that analyzes data obtained from external sensors and cameras and recognizes the user's emotional state.
[1048] "Video data" is visual information about historical events or scientific phenomena, and is digital content provided to users.
[1049] An "AI chatbot" is an artificial intelligence-based dialogue system that interacts with users and provides detailed explanations and follow-ups.
[1050] A "head-mounted display" is a display device that is worn by a user to visually experience real-world events or scientific phenomena.
[1051] "Robot control means" refers to a control mechanism for automatically adjusting the speed and method of work based on the user's emotional state.
[1052] A "database" is an information management system that accumulates information about historical events and scientific phenomena and provides it to users.
[1053] A "sensor" is a hardware device for detecting a user's emotional state.
[1054] A "camera" is a photographing device for capturing the user's facial expressions and movements.
[1055] "Work speed" is the speed at which the robot progresses through a task, and is adjusted based on the user's emotional state.
[1056] In order to implement this invention, the server, the terminal, and the user each play their respective roles, thereby realizing a customized experience and adjustment of the working environment according to the user's emotions.
[1057] The server builds an emotion engine and provides a function to recognize the user's emotions. Specifically, the server analyzes data acquired from external sensors and cameras, and determines the user's emotional state based on the user's facial expressions and heart rate. Based on the acquired emotion data, the server customizes the user's experience and tasks in real time.
[1058] The device is equipped with sensors and cameras to recognize the user's emotions and transmits the acquired emotional data to a server. It also receives related video data based on the information selected by the user and displays it on the head-mounted display. This allows the user to visually experience historical events and scientific phenomena. In addition, the device interacts with an AI chatbot to provide detailed explanations.
[1059] As a concrete example, consider a case where a worker on a packaging line in a factory suddenly becomes stressed. In this case, the server recognizes the worker's emotions based on data acquired from cameras and vital sensors, and automatically adjusts the robot's work speed. In addition, an AI chatbot will speak to the worker at appropriate times, encouraging them to "slow down your work pace" and refresh themselves. Conversely, if the worker is relaxed, they will continue working at their normal speed.
[1060] The hardware used is a webcam and vital sensors (to monitor heart rate and sweat gland activity), while the software uses an image processing library, emotion recognition API, and interpretive programming language, etc. This makes it possible to analyze the user's emotions in real time and provide a customized experience and work environment based on that.
[1061] An example of a prompt would be:
[1062] Please explain a system that recognizes the emotional state (stress, relaxation, etc.) of workers on a packaging line in a factory in real time and automatically adjusts the speed and method of a robot's work based on that information. Also, please describe in detail the role of an AI chatbot that interacts with the workers.
[1063] This allows for a customized experience and adjustment of the working environment according to the user's emotional state.
[1064] The flow of the specific process in the application example 2 will be described with reference to FIG.
[1065] Step 1:
[1066] Terminal (user operation)
[1067] The terminal activates the camera and vital sensors when the worker starts working. The camera captures the worker's face, and the vital sensors monitor the heart rate and sweat gland activity. This allows data related to the user's (worker's) emotional state to be obtained. The input is the captured image data and vital data, and the output is processed data to be passed to the emotion recognition model.
[1068] Step 2:
[1069] Terminal (data transmission)
[1070] The device transmits the acquired emotion data to the server in real time. The transmitted data includes captured image data and vital data. The input is the data collected from the user, and the output is a data stream transmitted to the server. The data is converted into an appropriate format for analysis on the server side.
[1071] Step 3:
[1072] Server (emotion recognition)
[1073] The server analyzes the received data and recognizes the user's emotional state using an emotion engine. Specifically, it extracts emotions from image data using image processing libraries and emotion recognition APIs, and integrates vital data. The input is the transmitted data stream, and the output is the judgment result of the emotional state. The server does this in real time.
[1074] Step 4:
[1075] Server (Emotion-based customization)
[1076] The server determines the necessary work adjustments based on the emotion recognition results. For example, if it is determined that the worker is stressed, it processes the work robot to slow down. The input is the emotion recognition result, and the output is a control command for the work robot. This control command is sent to the robot control system.
[1077] Step 5:
[1078] Terminal (Interaction with AI chatbot)
[1079] The AI chatbot on the device provides feedback according to the user's emotional state. For example, if the user is feeling stressed, the AI chatbot will say, "It's okay to slow down your work pace." The input is the emotion recognition result, and the output is the content of the chatbot's dialogue. This allows the user to receive psychological support.
[1080] Step 6:
[1081] User (feedback on experience)
[1082] The user (worker) takes actions such as adjusting the work pace or taking a break based on the feedback from the AI chatbot. Based on this feedback, the user selects an action that will improve their emotional state. The input is the feedback from the chatbot, and the output is the user's behavior change.
[1083] Through these steps, a customized experience and adjustments to the working environment can be made in real time based on the user's emotional state.
[1084] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input for the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1085] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by making a neural network perform deep learning. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating a voice, text data indicating a text, and image data indicating an image is input. The data generation model 58 performs inference on the input inference data according to the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1086] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1087] [Fourth embodiment]
[1088] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1089] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1090] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a wide area network (WAN) and / or a local area network (LAN).
[1091] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. In addition, the microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1092] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs the voice according to instructions from the processor 46.
[1093] Camera 42 is a small digital camera equipped with an optical system including a lens, an aperture, and a shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (e.g., an imaging range defined by an angle of view equivalent to the width of the field of vision of an average healthy person).
[1094] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for transmitting and receiving various types of information between the processor 46 and the processor 28 via the network 54. The transmission and reception of various types of information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is performed in a secure state.
[1095] The control target 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, legs, etc. The posture and behavior of the robot 414 are controlled by controlling the motors of the arms, hands, legs, etc. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1096] Fig. 8 shows an example of main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1097] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32, and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1098] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1099] In the robot 414, the reception and output process is performed by the processor 46. A reception and output program 60 is stored in the storage 50. The processor 46 reads the reception and output program 60 from the storage 50, and executes the read reception and output program 60 on the RAM 48. The reception and output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception and output program 60 executed on the RAM 48.
[1100] Next, a description will be given of the specific processing by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal".
[1101] An embodiment for implementing the present invention includes the following elements.
[1102] 1. Server
[1103] Build a database to capture video data related to historical events and scientific phenomena.
[1104] Based on the user's selection, relevant video data is extracted and transformed to suit the user's perspective.
[1105] It provides an AI chatbot that acts as a guide and provides detailed explanations to users.
[1106] 2. Terminal
[1107] Video data is received from the server based on a user-selected historical event or scientific phenomenon.
[1108] The received video data is displayed on a head-mounted display.
[1109] It accepts user operations and sends the selected information to the server.
[1110] 3. Users
[1111] By operating the terminal, users select the historical event or scientific phenomenon they would like to virtually travel through.
[1112] Wearing a head-mounted display, participants visually experience a virtual site and central area.
[1113] Interact with an AI chatbot that acts as your guide and receives detailed explanations.
[1114] The above is an example of an embodiment of the present invention. The server acquires and converts video data and provides an AI chatbot, the terminal receives and displays the video data, and the user selects and experiences information and interacts with the AI chatbot. This allows the user to take a virtual trip and gain a deeper understanding of historical events and scientific phenomena.
[1115] The process flow will be explained below.
[1116] Step 1: The server acquires and converts the video data.
[1117] The server retrieves video data related to historical events or scientific phenomena selected by the user.
[1118] The acquired video data is extracted from a database in the server.
[1119] The server converts the captured video data to match the user's viewpoint, allowing the user to visually experience a virtual site or center.
[1120] Step 2: The device receives and displays the video data
[1121] The terminal receives the video data transmitted from the server.
[1122] The received video data is displayed on a head-mounted display inside the terminal.
[1123] Users can wear a head-mounted display and visually experience a virtual site or central area.
[1124] Step 3: User makes information selections and experiences
[1125] The user operates the terminal to select the historical event or scientific phenomenon about which they would like to virtually travel.
[1126] The selected information is transmitted from the terminal to the server.
[1127] Based on the user's selection, the server retrieves the relevant video data, converts it and sends it to the terminal.
[1128] Step 4: User interacts with AI chatbot and receives detailed instructions
[1129] Users interact with the AI chatbot through a head-mounted display.
[1130] The AI chatbot provides users with detailed explanations of selected historical events or scientific phenomena.
[1131] Users can gain deeper understanding through dialogue with AI chatbots.
[1132] This is the flow of the program's processing. The server acquires and converts the video data, and the terminal receives and displays the video data. The user can select and experience information, and receive detailed explanations through dialogue with the AI chatbot.
[1133] Example 1
[1134] Next, a description will be given of Example 1. In the following description, the data processing device 12 is referred to as a "server" and the robot 414 is referred to as a "terminal."
[1135] Currently, there are only a limited number of systems that allow users to visually experience historical events or scientific phenomena, making it difficult for users to gain a deep understanding. In addition, existing systems are unable to convert in real time to match the user's perspective, which results in a lack of immersion. Furthermore, there are insufficient guides to provide detailed information, and information supplementation is not adequate.
[1136] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1137] In this invention, the server includes a means for constructing a database for acquiring video data, a means for extracting video data based on a specific event or phenomenon selected by the user, a means for converting the extracted video data to match the user's viewpoint and displaying it, and a means for interacting with an AI chatbot as a guide that provides detailed explanations to the user based on a request provided by the user, thereby enabling the user to gain a deeper understanding of historical events and scientific phenomena through a highly immersive virtual journey.
[1138] A "database" is a collection of data in which information is systematically collected, stored, and made available for rapid retrieval.
[1139] "Video data" is digital or analog data that contains visual information, and is content that includes moving images and still images.
[1140] "Transformation to fit the viewpoint" refers to the process of adjusting the original video data according to the user's field of view and visual needs, and displaying it appropriately.
[1141] An "AI chatbot as a guide" is an automated dialogue system that uses artificial intelligence to provide appropriate information in response to user questions.
[1142] A "visual display device" is a device that allows a user to visually experience video data, and primarily refers to a head-mounted display or a monitor.
[1143] "Additional information" is supplemental information about the event or phenomenon the user is experiencing, including detailed explanations and background knowledge.
[1144] A "request" refers to the act of a user sending a specific request or instruction to the system.
[1145] The embodiment for implementing the present invention includes the following elements.
[1146] 1. Server configuration and roles:
[1147] The server builds a database to acquire and store video data related to historical events and scientific phenomena. Specifically, the server manages a huge amount of video content and extracts relevant video data based on a specific event or phenomenon selected by the user. For example, the database can include "video of pyramid construction in ancient Egypt" and "video on the development of scientific theories." Furthermore, the server converts the extracted video data to match the user's viewpoint. This conversion is intended to enable the user to view the video from a 360-degree viewpoint. The server also provides an AI chatbot as a guide, providing detailed explanations for requests provided by the user.
[1148] 2. Terminal configuration and roles:
[1149] The terminal receives video data from the server based on historical events or scientific phenomena selected by the user. The received video data is displayed on a visual display device (e.g., a head-mounted display). A specific example is a device such as a VR headset. The terminal accepts user operations and transmits the operation contents to the server. For example, if the user requests to zoom in on a particular scene or to know more about a particular object, the request is transmitted to the server.
[1150] 3. User interaction and experience:
[1151] Users can operate the device to experience a virtual journey. For example, a user can select a specific historical event or scientific phenomenon, such as "the discovery of Machu Picchu" or "scientific experiment in a laboratory." After that, the user puts on the head-mounted display and visually experiences the selected scene from a 360-degree perspective. Furthermore, if the user wants to know more information, they can interact with an AI chatbot that acts as a guide. For example, if the user asks, "What were the most commonly sold products in this ancient market?", the AI chatbot will answer, "Grains and olive oil were commonly traded."
[1152] Examples:
[1153] As a concrete example, consider the case where the user selects "Darwin's Voyage of the Beagle." In this case, the user receives the video data from the server through the terminal and it is displayed on the head-mounted display. When the user asks a question about the plants and animals discovered during Darwin's voyage, the AI chatbot will explain the details.
[1154] Examples of prompts for a generative AI model include:
[1155] "What were the most commonly sold items in ancient Roman markets?"
[1156] "I would like to know more about what Darwin discovered on his Beagle voyage."
[1157] "I want to see an episode about the discovery of the theory of relativity."
[1158] This allows users to take a virtual journey and gain a deeper understanding of historical events or scientific phenomena.
[1159] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1160] Step 1:
[1161] Database construction (server)
[1162] The server creates a database that systematically collects and stores information.
[1163] Input: Video content related to historical events or scientific phenomena.
[1164] Data processing: Classify video content by category and create an index.
[1165] Output: The indexed database.
[1166] Specific operation: The server collects, for example, footage of "scientific experiments" or "ancient civilizations" from the Internet, adds metadata to them, and stores them in a database.
[1167] Step 2:
[1168] Receiving user request (terminal)
[1169] The terminal provides an interface for the user to select a particular event or phenomenon.
[1170] Input: A request for a user-selected event or phenomenon.
[1171] Data processing: Analyze the request content and send it to the server to extract the necessary data.
[1172] Output: The parsed request data.
[1173] Specific operation: When a user selects "Darwin's Voyage of the Beagle," the request is sent from the terminal to the server.
[1174] Step 3:
[1175] Extraction and conversion of video data (server)
[1176] The server extracts relevant video data from a database based on the request.
[1177] Input: The user's request data.
[1178] Data processing: Based on the request, relevant video data is searched for in the database and the video is transformed to suit the user's perspective.
[1179] Output: Video data transformed to match the user's viewpoint.
[1180] Specific operation: The server retrieves footage, for example, of "Darwin's Voyage of the Beagle" from a database and converts the footage into a 360-degree view.
[1181] Step 4:
[1182] Video data transmission (server)
[1183] The server transmits the converted video data to the terminal.
[1184] Input: Video data transformed to match the user's viewpoint.
[1185] Data processing: Encode video data into a transmittable format.
[1186] Output: The encoded video data is sent to the device.
[1187] Specific operation: The server transmits video data encoded in, for example, MP4 format to the terminal in real time.
[1188] Step 5:
[1189] Receiving video data (terminal)
[1190] The terminal receives the video data transmitted from the server.
[1191] Input: Encoded video data sent from the server.
[1192] Data processing: Decoding the encoded video data and converting it into a displayable format.
[1193] Output: Decoded video data.
[1194] Specific operation: The terminal decodes the received video data, for example in MP4 format, and stores it in its internal memory.
[1195] Step 6:
[1196] Display of video data (terminal)
[1197] The terminal displays the received video data on a visual display device.
[1198] Input: Decoded video data.
[1199] Data Processing: Converting the decoded video data into a resolution and format suitable for a visual display device.
[1200] Output: Video data displayed on a visual display device.
[1201] Specific operation: The terminal displays a 360-degree image on, for example, a head-mounted display (HMD), allowing the user to visually experience the scene.
[1202] Step 7:
[1203] Receiving user operations (terminal)
[1204] The terminal receives user operations and requests.
[1205] Input: Actions taken by the user or additional requests.
[1206] Data processing: Analyzes operation and request data and sends it to the server.
[1207] Output: Parsed operation and request data.
[1208] Specific operation: When a user makes a request such as "I would like to take a closer look at Darwin's notes," the request is sent to the server via the terminal.
[1209] Step 8:
[1210] Interaction with AI chatbot (server, user)
[1211] Based on the user's request, the server launches an AI chatbot to act as a guide and engages in a conversation.
[1212] Input: The request made by the user.
[1213] Data processing: The AI chatbot retrieves information corresponding to the request and provides the user with an appropriate answer.
[1214] Output: A detailed description provided to the user.
[1215] What it does: When a user asks, "What's in Darwin's notebooks?" the AI chatbot responds, "These contain Darwin's early thoughts on evolution."
[1216] (Application example 1)
[1217] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1218] Traditional education of historical events and scientific phenomena has often relied on static media such as textbooks and videos, and it has been difficult to actually experience them. This has made it difficult to maintain users' understanding and interest. In addition, guidebooks and human guides have limitations in their explanations, and the information is not always consistently detailed and individualized. This has created a need for an effective way for users to gain a deeper understanding of specific events and phenomena.
[1219] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1220] In this invention, the server includes a means for acquiring video data related to a historical event or a scientific phenomenon selected by a user based on the historical event or the scientific phenomenon, a means for converting the acquired video data to match the user's viewpoint and displaying it, a means for interacting with an AI chatbot that acts as a guide and provides the user with a detailed explanation of the selected historical event or the scientific phenomenon, a means for the user to experience the video data by wearing a head-mounted display in a physical store, and a means for the AI chatbot to provide a detailed explanation in real time as a guide. This allows the user to experience the historical event or scientific phenomenon in an immersive way and receive a detailed and personalized explanation from the AI chatbot.
[1221] A "user" is someone who uses the system to experience historical events or scientific phenomena.
[1222] A "historic event" refers to a specific occurrence or phenomenon that had a significant impact on human history.
[1223] A "scientific phenomenon" refers to an event that allows for the scientific understanding and analysis of laws and phenomena in the natural world.
[1224] "Video data" refers to video and image information relating to a specific event or phenomenon.
[1225] "Means of acquiring" refers to the method or mechanism for getting the required information or data into a system or device.
[1226] "Means for converting to suit the viewpoint" refers to technology that adjusts video data in an appropriate manner according to the user's viewpoint and position.
[1227] The term "means for displaying" refers to a mechanism or method for presenting video data in a form that can be visually recognized by a user.
[1228] An "AI chatbot as a guide" refers to a program that uses artificial intelligence technology to provide information to users and engage in dialogue.
[1229] "Means of dialogue" refers to the technology that allows users and AI chatbots to communicate with each other through questions and answers.
[1230] "System" refers to a set of mechanisms in which multiple means and devices work together to provide specific services and functions to users.
[1231] A "head-mounted display" refers to a display device that is worn on the head and allows the user to visually experience immersive images.
[1232] "Real-time" refers to data and information being processed and displayed instantly, without delay.
[1233] A "database" refers to a system for systematically organizing and storing specific information or data.
[1234] "Means for providing detailed explanations" refers to a method for providing specific and detailed information necessary for users to deepen their understanding.
[1235] A "physical store" refers to a sales location that has a physical presence to offer goods or services, and is a permanent commercial space.
[1236] "Means of experiencing" refers to the technology or methods that allow a user to actually experience a particular event through simulation or virtual reality.
[1237] In the embodiment of the present invention, the entire system is configured as follows: The system used by a user mainly comprises a server, a terminal, and a head mounted display.
[1238] 1. Server
[1239] The server is responsible for building a database to retrieve video data related to historical events or scientific phenomena selected by the user. The retrieved video data is then transformed to fit the user's perspective. The server also provides an AI chatbot to provide relevant detailed explanations to the user in real time.
[1240] The software used includes database access via REST API and a natural language processing engine for the AI chatbot. Specifically, the "requests" library is used to acquire video data, and a GPT series of generative AI models are used for natural language processing.
[1241] 2. Terminal
[1242] The terminal operated by the user receives the video data from the server and displays the data on the head mounted display. The terminal accepts the user's operations and transmits the information to the server.
[1243] The terminals are general-purpose smartphones or tablet devices equipped with a browser application that uses HTML5 or WebGL to receive and display video data.
[1244] 3. Head-mounted displays
[1245] The user wears a head-mounted display and visually experiences the video data transmitted from the server. The video data is updated in real time based on the user's viewpoint, providing an immersive experience.
[1246] The hardware used includes VR headsets, which track the movements of the user's head and adjust the image based on that point of view.
[1247] Examples
[1248] In the education section of the brick-and-mortar store, a visitor selects to experience "Egyptian pyramid construction." The visitor puts on a head-mounted display and is immersed in the pyramid construction site that unfolds before their eyes, witnessing workers stacking stones. In real time, an AI chatbot explains, "This site was built around 2500 BC," and provides additional details, "The stone used was mainly limestone."
[1249] Examples of prompt statements
[1250] "Generate guides based on video data about the construction of the Egyptian pyramids to allow visitors to virtually experience the sites."
[1251] As described above, the system of the present invention allows users to experience a specific historical event or scientific phenomenon in real time while receiving detailed explanations from an AI chatbot, which leads to deeper understanding and longer-lasting interest compared to traditional static educational methods.
[1252] The flow of the specific process in the application example 1 will be described with reference to FIG.
[1253] Step 1:
[1254] The user operates the terminal to select the historical event or scientific phenomenon they wish to experience.
[1255] Input: Event ID or theme selected by the user
[1256] Output: The selected event ID is sent to the server.
[1257] Specific operation: On the device interface, the user taps or clicks to select the item they want to experience. The device then sends this information to the server.
[1258] Step 2:
[1259] Based on the event ID received by the server, the server retrieves related video data from the database.
[1260] Input: Event ID sent by the user
[1261] Output: Captured video data
[1262] Specific operation: The server queries the database to retrieve video data associated with the specified event ID.
[1263] Step 3:
[1264] The server converts the acquired video data to match the user's viewpoint.
[1265] Input: Acquired video data, user's viewpoint information
[1266] Output: Converted video data
[1267] Specific operation: The server uses the user's head tracking data to convert the video data to match the appropriate viewpoint in real time.
[1268] Step 4:
[1269] The server transmits the converted video data to the terminal.
[1270] Input: Converted video data
[1271] Output: Video data sent to the device
[1272] Specific operation: The server streams the converted video data to the terminal via the network.
[1273] Step 5:
[1274] The video data received by the terminal is displayed on a head-mounted display.
[1275] Input: Video data sent from the server
[1276] Output: Image displayed on the head-mounted display
[1277] Specific operation: The terminal outputs video data to a head-mounted display, which the user experiences visually.
[1278] Step 6:
[1279] Users interact with an AI chatbot while visually experiencing the video using a head-mounted display.
[1280] Input: User questions and input
[1281] Output: Answers and explanations from the AI chatbot
[1282] How it works: During the experience, users input questions via voice or text, and the AI chatbot provides information in response to those questions in real time.
[1283] Step 7:
[1284] The AI chatbot provides detailed explanations to the user based on the database.
[1285] Input: User questions, database information
[1286] Output: Detailed explanation from the AI chatbot
[1287] Specific operation: The AI chatbot analyzes the user's question, retrieves relevant information from the database, and provides the user with appropriate explanations.
[1288] Through these steps, users can experience historical events and scientific phenomena in an immersive way within a physical store and receive detailed explanations.
[1289] In addition, an emotion engine that estimates the emotion of the user may be further combined. That is, the identification processing unit 290 may estimate the emotion of the user using the emotion identification model 59, and perform identification processing using the emotion of the user.
[1290] An embodiment for implementing the present invention includes the following elements.
[1291] 1. Server
[1292] Build an emotion engine to recognize user emotions.
[1293] Based on the user's selection, relevant video data is extracted and transformed to suit the user's perspective.
[1294] It provides an AI chatbot that acts as a guide and provides detailed explanations to users.
[1295] The system will incorporate a means to recognize the user's emotions and customize the images and explanations accordingly.
[1296] It is equipped with sensors and cameras to recognize the user's emotions.
[1297] The user's selection is transmitted to the server and related video data is received.
[1298] The received video data is displayed on a head-mounted display.
[1299] Incorporate a means to recognise user emotions and tailor the travel experience accordingly.
[1300] 3. Users
[1301] By operating the terminal, users select the historical event or scientific phenomenon they would like to virtually travel through.
[1302] Wearing a head-mounted display, participants visually experience a virtual site and central area.
[1303] Interact with an AI chatbot that acts as your guide and receives detailed explanations.
[1304] The system recognizes the user's emotions through sensors and cameras and customizes the images and explanations accordingly.
[1305] The above is an example of an embodiment of the present invention. The server provides the emotion engine construction and customization function, the terminal recognizes emotions and displays images, and the user selects and experiences information and interacts with the AI chatbot. This makes it possible to provide a travel experience customized to the user's emotions.
[1306] The process flow will be explained below.
[1307] Step 1: The server customizes the video data incorporating the emotion engine.
[1308] The server builds an emotion engine to recognize the user's emotion.
[1309] Based on the user's selection, relevant video data is extracted and transformed to suit the user's perspective.
[1310] The server uses an emotion engine to recognize the user's emotions and customize the video and explanations accordingly.
[1311] The customized video data is transmitted to the terminal.
[1312] Step 2: The device receives and displays the customized video data
[1313] The terminal receives the customized video data transmitted from the server.
[1314] The received video data is displayed on a head-mounted display.
[1315] The user wears a head-mounted display and visually experiences customized images.
[1316] Step 3: User emotions are recognized and the travel experience is tailored
[1317] The device uses sensors and cameras to recognize the user's emotions.
[1318] The user's emotion data is sent from the terminal to a server and analyzed.
[1319] The server tailors the travel experience based on the user's emotions, providing customized footage and commentary.
[1320] Users will enjoy a travel experience that is tailored to their emotions and receive detailed instructions through interactions with AI chatbots.
[1321] This is the flow of the program's processing. The server incorporates an emotion engine to customize the video data. The terminal receives and displays the customized video data, recognizes the user's emotions, and adjusts the travel experience. The user can enjoy a customized experience and receive detailed explanations through dialogue with the AI chatbot.
[1322] Example 2
[1323] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal".
[1324] Current virtual reality systems can provide video data on historical events or scientific phenomena selected by the user, but they lack the ability to dynamically adjust the video and explanation according to the user's emotions, making the experience uniform for each individual user and making true customization difficult. In addition, even if an AI chatbot provides an interactive explanation, the content is fixed and cannot immediately respond to the user's interests and emotions. For this reason, it is necessary to provide a highly personalized virtual travel experience that corresponds to the individual emotions and interests of each user.
[1325] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1326] In this invention, the server includes a means for acquiring video data related to a historical event or a scientific phenomenon selected by a user, a means for converting the acquired video data to match the user's viewpoint and displaying it, a means for customizing the video data using an emotion engine for recognizing the user's emotions, a means for interacting with an AI chatbot that acts as a guide and provides the user with a detailed explanation of the selected historical event or scientific phenomenon, and a means for dynamically adjusting the video data and the explanation content according to changes in the user's emotions. This enables a highly customized virtual travel experience according to the user's emotions and interests.
[1327] "Video Data" is digital data containing visual information related to a historical event or scientific phenomenon selected by a user.
[1328] An "emotion engine" is a software or hardware function for recognizing a user's emotional state by analyzing the user's facial expressions and behavior.
[1329] "Transforming to fit the user's perspective" means adjusting visual information in real time to accommodate the user's perspective and movements as they experience the site or center.
[1330] An "AI chatbot" is a program that uses artificial intelligence to converse with users in natural language and act as a virtual guide, providing specific information and detailed explanations.
[1331] "Dynamic adjustment" refers to the process of instantly changing or adapting video data and narrative content based on the user's real-time emotions and reactions.
[1332] A "head-mounted display" is a device worn by a user to visually experience virtual reality environments and images.
[1333] "Database" means a collection of information that effectively stores and manages additional information related to a historical event or scientific phenomenon selected by a User.
[1334] "Customization" means tailoring and modifying the experience and information provided to suit a user's specific needs and circumstances, based on their individual feelings and interests.
[1335] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[1336] The present invention relates to a system for providing a virtual travel experience customized according to a user's emotions. A specific embodiment of the system will be described below.
[1337] 1. Hardware and Software Used
[1338] In this system, a server, a terminal, and a user each play a role. The server is a computer system that includes an emotion engine, a database, and an AI chatbot. The terminal is a device equipped with sensors, a camera, and a head-mounted display to recognize the user's emotions. The emotion engine uses a general emotion analysis API (e.g., a natural language processing engine), and the AI chatbot uses a general generative AI model (e.g., a natural language generation model).
[1339] Specific hardware examples include a "depth camera" for emotion recognition and "VR goggles" for the head-mounted display. The server is operated on a "cloud computing platform" and data storage uses a "cloud storage service." Specific software examples include a "natural language processing API" for the emotion engine and a "natural language generation API" for the generative AI model.
[1340] 2. Program Processing
[1341] The server retrieves video data related to historical events or scientific phenomena selected by the user from the cloud storage. It uses a 3D engine (e.g., a graphics engine) to transform the video data to fit the user's viewpoint. The server then uses an emotion engine to analyze the user's emotions and customizes the video data based on the analysis results.
[1342] For example, if a user selects "Age of Dinosaurs," and the server determines the user's emotion as "joy" using the emotion engine, the video will be centered around playful scenes of dinosaurs.
[1343] Meanwhile, the device is equipped with sensors and cameras to recognize the user's emotions in real time. The device decodes the video data received from the server and displays it on the user's head-mounted display. While the user is experiencing the virtual space, the device continuously transmits emotional data to the server, and the server dynamically adjusts the video and explanations based on this data.
[1344] Moreover, the AI chatbot acts as a guide, providing detailed descriptions: when the user asks, "What is this?", the chatbot responds, "This dinosaur was a Tyrannosaurus, and it mainly ate meat."
[1345] Examples
[1346] The user selects "Dinosaur Era." The server uses an emotion engine to analyze the user's emotions and determines that the user is "excited." Based on this, the server retrieves a dynamic dinosaur scene and places it in 3D space. The device displays this video data on a head-mounted display, and the user is immersed in the world of dinosaurs in real time. Furthermore, depending on the user's excitement, further interactive scenes are pushed from the server.
[1347] Examples of prompt statements
[1348] "Provide action-packed footage from the age of dinosaurs for your excited users."
[1349] The system enables a highly customized virtual travel experience based on the user's emotions and interests.
[1350] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1351] Step 1:
[1352] The user selects a virtual travel theme on the terminal.
[1353] As a specific operation, the user operates the terminal interface and selects "Dinosaur Age."
[1354] Input: User's theme selection (e.g. "Age of Dinosaurs")
[1355] Output: The selected theme ID.
[1356] Step 2:
[1357] The terminal sends the user's request to the server.
[1358] A request data with the selected theme ID is created and sent to the server via Wi-Fi.
[1359] Input: Selected Theme ID (e.g. "Age of Dinosaurs")
[1360] Output: A request containing theme selection information.
[1361] Step 3:
[1362] The server acquires emotion data for recognizing the user's emotion.
[1363] Emotion data sent from the device is received and analyzed by the emotion engine.
[1364] Input: Emotion data sent from the device
[1365] Output: Analyzed emotion information (e.g. "joy")
[1366] Step 4:
[1367] The server extracts the relevant video data and transforms it to suit the user's perspective.
[1368] Video data related to the "Dinosaur Age" is retrieved from cloud storage and converted to fit the user's perspective using the Unity3D engine.
[1369] Input: Theme selection information, analyzed emotion information
[1370] Output: Converted video data
[1371] Step 5:
[1372] The server transmits the converted video data to the terminal.
[1373] Using streaming technology, the converted video data is transferred to the terminal in real time.
[1374] Input: Converted video data
[1375] Output: Video data sent to the device
[1376] Step 6:
[1377] The terminal displays the video data on a head-mounted display.
[1378] The received video data is decoded and displayed on a head-mounted display, allowing the user to experience a virtual journey.
[1379] Input: Received video data
[1380] Output: Image displayed on a head-mounted display
[1381] Step 7:
[1382] The terminal continuously transmits the user's emotion data to the server.
[1383] The data obtained from the emotion sensor is continuously sent to the server in real time.
[1384] Input: Real-time emotion data
[1385] Output: Continuous emotion data sent to the server.
[1386] Step 8:
[1387] The server provides detailed instructions to the user through an AI chatbot.
[1388] Depending on the user's movements and questions, the AI chatbot will provide detailed explanations, for example, "This dinosaur is a Tyrannosaurus."
[1389] Input: User questions, real-time sentiment data
[1390] Output: Detailed explanation by AI chatbot
[1391] Step 9:
[1392] The server customizes images and explanations in real time according to the user's emotions.
[1393] Based on the analyzed emotion information, the video and explanation are dynamically adjusted. For example, if the user feels "surprise", an interesting additional scene is provided.
[1394] Input: Continuous emotion data, current video and description information
[1395] Output: Real-time customized video and explanations
[1396] This provides a highly customized virtual travel experience that responds to the user's emotions and interests.
[1397] (Application example 2)
[1398] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal".
[1399] In modern factories, the emotions and stress levels of workers often have a significant impact on work efficiency and quality. However, conventional factory robots and management systems operate at a uniform work speed and in a uniform manner without taking into account the emotional state of workers, which causes problems of increased stress and fatigue among workers. Furthermore, since they do not respond to the individual conditions of workers, it is difficult to maximize the overall work efficiency and quality. Similarly, there has been a lack of technology that recognizes emotions and stress levels in real time and automatically adjusts based on them. To solve these issues, there is a demand for a system that recognizes the user's emotions in real time and automatically adjusts work based on them.
[1400] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1401] In this invention, the server includes a means for recognizing the user's emotions, a means for acquiring related video data based on information selected by the user, and a means for converting and displaying the acquired video data in accordance with the user's viewpoint, thereby making it possible to provide a customized experience and adjust the working environment according to the user's emotions.
[1402] A "user" is a person who interacts with the system and provides affective states and selections.
[1403] The "emotion engine" is a mechanism that analyzes data obtained from external sensors and cameras and recognizes the user's emotional state.
[1404] "Video data" is visual information about historical events or scientific phenomena, and is digital content provided to users.
[1405] An "AI chatbot" is an artificial intelligence-based dialogue system that interacts with users and provides detailed explanations and follow-ups.
[1406] A "head-mounted display" is a display device that is worn by a user to visually experience real-world events or scientific phenomena.
[1407] "Robot control means" refers to a control mechanism for automatically adjusting the speed and method of work based on the user's emotional state.
[1408] A "database" is an information management system that accumulates information about historical events and scientific phenomena and provides it to users.
[1409] A "sensor" is a hardware device for detecting a user's emotional state.
[1410] A "camera" is a photographing device for capturing the user's facial expressions and movements.
[1411] "Work speed" is the speed at which the robot progresses through a task, and is adjusted based on the user's emotional state.
[1412] In order to implement this invention, the server, the terminal, and the user each play their respective roles, thereby realizing a customized experience and adjustment of the working environment according to the user's emotions.
[1413] The server builds an emotion engine and provides a function to recognize the user's emotions. Specifically, the server analyzes data acquired from external sensors and cameras, and determines the user's emotional state based on the user's facial expressions and heart rate. Based on the acquired emotion data, the server customizes the user's experience and tasks in real time.
[1414] The device is equipped with sensors and cameras to recognize the user's emotions and transmits the acquired emotional data to a server. It also receives related video data based on the information selected by the user and displays it on the head-mounted display. This allows the user to visually experience historical events and scientific phenomena. In addition, the device interacts with an AI chatbot to provide detailed explanations.
[1415] As a concrete example, consider a case where a worker on a packaging line in a factory suddenly becomes stressed. In this case, the server recognizes the worker's emotions based on data acquired from cameras and vital sensors, and automatically adjusts the robot's work speed. In addition, an AI chatbot will speak to the worker at appropriate times, encouraging them to "slow down your work pace" and refresh themselves. Conversely, if the worker is relaxed, they will continue working at their normal speed.
[1416] The hardware used is a webcam and vital sensors (to monitor heart rate and sweat gland activity), while the software uses an image processing library, emotion recognition API, and interpretive programming language, etc. This makes it possible to analyze the user's emotions in real time and provide a customized experience and work environment based on that.
[1417] An example of a prompt would be:
[1418] Please explain a system that recognizes the emotional state (stress, relaxation, etc.) of workers on a packaging line in a factory in real time and automatically adjusts the speed and method of a robot's work based on that information. Also, please describe in detail the role of an AI chatbot that interacts with the workers.
[1419] This allows for a customized experience and adjustment of the working environment according to the user's emotional state.
[1420] The flow of the specific process in the application example 2 will be described with reference to FIG.
[1421] Step 1:
[1422] Terminal (user operation)
[1423] The terminal activates the camera and vital sensors when the worker starts working. The camera captures the worker's face, and the vital sensors monitor the heart rate and sweat gland activity. This allows data related to the user's (worker's) emotional state to be obtained. The input is the captured image data and vital data, and the output is processed data to be passed to the emotion recognition model.
[1424] Step 2:
[1425] Terminal (data transmission)
[1426] The device transmits the acquired emotion data to the server in real time. The transmitted data includes captured image data and vital data. The input is the data collected from the user, and the output is a data stream transmitted to the server. The data is converted into an appropriate format for analysis on the server side.
[1427] Step 3:
[1428] Server (emotion recognition)
[1429] The server analyzes the received data and recognizes the user's emotional state using an emotion engine. Specifically, it extracts emotions from image data using image processing libraries and emotion recognition APIs, and integrates vital data. The input is the transmitted data stream, and the output is the judgment result of the emotional state. The server does this in real time.
[1430] Step 4:
[1431] Server (Emotion-based customization)
[1432] The server determines the necessary work adjustments based on the emotion recognition results. For example, if it is determined that the worker is stressed, it processes the work robot to slow down. The input is the emotion recognition result, and the output is a control command for the work robot. This control command is sent to the robot control system.
[1433] Step 5:
[1434] Terminal (Interaction with AI chatbot)
[1435] The AI chatbot on the device provides feedback according to the user's emotional state. For example, if the user is feeling stressed, the AI chatbot will say, "It's okay to slow down your work pace." The input is the emotion recognition result, and the output is the content of the chatbot's dialogue. This allows the user to receive psychological support.
[1436] Step 6:
[1437] User (feedback on experience)
[1438] The user (worker) takes actions such as adjusting the work pace or taking a break based on the feedback from the AI chatbot. Based on this feedback, the user selects an action that will improve their emotional state. The input is the feedback from the chatbot, and the output is the user's behavior change.
[1439] Through these steps, a customized experience and adjustments to the working environment can be made in real time based on the user's emotional state.
[1440] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires a voice indicating a user input for the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1441] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by making a neural network perform deep learning. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating a voice, text data indicating a text, and image data indicating an image is input. The data generation model 58 performs inference on the input inference data according to the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1442] In the above embodiment, an example was given in which the specific process was performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the robot 414.
[1443] The emotion identification model 59 as an emotion engine may determine the emotion of the user according to a specific mapping. Specifically, the emotion identification model 59 may determine the emotion of the user according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the emotion of the robot, and the identification processing unit 290 may perform identification processing using the emotion of the robot.
[1444] FIG. 9 is a diagram showing an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive emotions are arranged. The more outside the concentric circles, the more emotions that represent states and actions that arise from a state of mind are arranged. Emotions are a concept that includes emotions and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions that occur in the brain are arranged. On the right side of the concentric circles, emotions that are generally induced by situational judgment are arranged. On the upper and lower sides of the concentric circles, emotions that are generally generated from reactions that occur in the brain and are induced by situational judgment are arranged. In addition, on the upper side of the concentric circles, emotions of "pleasure" are arranged, and on the lower side, emotions of "discomfort" are arranged. In this way, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1445] These emotions are distributed in the 3 o'clock direction of emotion map 400 and usually fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1446] The inside of emotion map 400 represents what is going on inside one's mind, and the outside of emotion map 400 represents behavior, so the further out you go on emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1447] Here, human emotions are based on various balances such as posture and blood sugar level, and when these balances are far from the ideal, it indicates an unpleasant state, and when they are close to the ideal, it indicates a pleasant state. Emotions can also be created for robots, cars, motorcycles, etc., based on various balances such as posture and battery level, so that when these balances are far from the ideal, it indicates an unpleasant state, and when they are close to the ideal, it indicates a pleasant state. The emotion map may be generated, for example, based on the emotion map of Dr. Mitsuyoshi (Research on speech emotion recognition and emotion brain physiological signal analysis system, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). On the left half of the emotion map, emotions belonging to an area called "reaction" where sensation is dominant are lined up. On the right half of the emotion map, emotions belonging to an area called "situation" where situation recognition is dominant are lined up.
[1448] The emotion map defines two emotions that promote learning. The first is the negative emotion around the middle of "repentance" or "remorse" on the situation side. In other words, this is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the positive emotion around "desire" on the response side. In other words, this is when the robot has positive feelings such as "I want more" or "I want to know more."
[1449] The emotion identification model 59 inputs the user input to a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the emotion of the user. This neural network is pre-trained based on multiple learning data that are combinations of the user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in Fig. 10. Fig. 10 shows an example in which multiple emotions, "relief," "calm," and "encouraging," have similar emotion values.
[1450] Although the system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, the system according to the present disclosure is not necessarily implemented in a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program that runs on a personal computer, or an application that runs on a smartphone or the like. The method according to the present disclosure may be provided to a user in the form of SaaS (Software as a Service).
[1451] In the above embodiment, an example is given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to input data.
[1452] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable computer-readable non-transitory storage medium such as a Universal Serial Bus (USB) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1453] In addition, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 upon request from the data processing device 12.
[1454] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1455] As the hardware resource for executing the specific process, various processors as shown below can be used. An example of the processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing the specific process by executing software, i.e., a program. Another example of the processor is a dedicated electric circuit, which is a processor having a circuit configuration designed exclusively for executing the specific process, such as a Field-Programmable Gate Array (FPGA), a Programmable Logic Device (PLD), or an Application Specific Integrated Circuit (ASIC). Each processor has a built-in or connected memory, and each processor executes the specific process by using the memory.
[1456] The hardware resource that executes the specific process may be one of these various processors, or may be a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific process may be a single processor.
[1457] As an example of a configuration using one processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a configuration using a processor that realizes the functions of the entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1458] Furthermore, more specifically, the hardware structure of these various processors can be an electric circuit that combines circuit elements such as semiconductor elements. The specific processes described above are merely examples. It goes without saying that unnecessary steps may be deleted, new steps may be added, or the order of processes may be changed without departing from the spirit of the invention.
[1459] The above description and illustrations are detailed descriptions of the parts related to the technology of the present disclosure, and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, function, action, and effect is an example of the configuration, function, action, and effect of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above description and illustrations, within the scope of the gist of the technology of the present disclosure. In addition, in order to avoid confusion and to facilitate understanding of the parts related to the technology of the present disclosure, the above description and illustrations omit explanations of technical common sense that do not require explanation in order to enable the implementation of the technology of the present disclosure.
[1460] All publications, patent applications, and standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, and standard was specifically and individually indicated to be incorporated by reference.
[1461] The following is further disclosed regarding the above embodiment.
[1462] (Claim 1) means for acquiring video data related to a historical event or a scientific phenomenon based on a historical event or a scientific phenomenon selected by a user; a means for converting the acquired video data to match the user's viewpoint and displaying the converted video data; A means for interacting with an AI chatbot that serves as a guide and provides the user with a detailed explanation of the selected historical event or scientific phenomenon; A system including: (Claim 2) In the system of claim 1, A system including means for using a head mounted display for said user to visually experience the site of said historical event or the epicenter of said scientific phenomenon. (Claim 3) In the system according to claim 1 or 2, The system includes a means for constructing a database to provide additional information related to the historical event or scientific phenomenon selected by the user, and for the AI chatbot to act as a guide and provide a detailed explanation based on the database. (Claim 4) In the system of claim 1, The system includes an emotion engine that recognizes the user's emotion and a means for recognizing the user's emotion and customizing images and explanations according to the emotion. (Claim 5) In a system according to claim 1 or 4, The system includes an emotion engine for recognizing emotions of the user, the emotion engine including means for recognizing emotions of the user and adjusting a travel experience based on the emotions.
[1463] "Example 1" (Claim 1) A means for constructing a database for acquiring video data; means for extracting video data based on a particular event or phenomenon selected by a user; A means for converting the extracted video data to match a user's viewpoint and displaying the converted video data; A means for interacting with an AI chatbot that acts as a guide and provides detailed instructions to the user based on a request provided by the user; A system including: (Claim 2) 10. The system of claim 1, employing a visual display device for allowing a user to visually experience the site of an event or the heart of a phenomenon. (Claim 3) The system of claim 1 builds a database to provide additional information related to a specific event or phenomenon selected by the user, and an AI chatbot acts as a guide to provide a detailed explanation based on the database.
[1464] "Application example 1" (Claim 1) means for acquiring video data related to a historical event or a scientific phenomenon based on a historical event or a scientific phenomenon selected by a user; a means for converting the acquired video data to match the user's viewpoint and displaying the converted video data; A means for interacting with an AI chatbot that serves as a guide and provides the user with a detailed explanation of the selected historical event or scientific phenomenon; A means for allowing a user to experience the video data by wearing a head mounted display in a physical store; A means for the AI chatbot to provide detailed explanations in real time as a guide; A system including: (Claim 2) The system of claim 1 , wherein a user can experience historical events or scientific phenomena through the head-mounted display. (Claim 3) The system of claim 1, further comprising: a database for providing additional information related to the selected historical event or scientific phenomenon; and the AI chatbot acting as a guide provides a detailed explanation based on the database.
[1465] "Example 2 of combining emotion engines" (Claim 1) means for acquiring video data related to a historical event or a scientific phenomenon based on a historical event or a scientific phenomenon selected by a user; a means for converting the acquired video data to match the user's viewpoint and displaying the converted video data; means for customizing said video data using an emotion engine for recognizing an emotion of a user; A means for interacting with an AI chatbot that serves as a guide and provides the user with a detailed explanation of the selected historical event or scientific phenomenon; A means for dynamically adjusting the video data and the explanatory content in response to a change in the user's emotion; A system including: (Claim 2) 10. The system of claim 1, further comprising means for using a head mounted display for the user to visually experience the site of the historical event or the epicenter of the scientific phenomenon. (Claim 3) The system of claim 1, further comprising: means for constructing a database for providing additional information related to the historical event or scientific phenomenon selected by the user, and for the AI chatbot to provide a detailed explanation as a guide based on the database.
[1466] "Application example 2 when combining emotion engines" (Claim 1) means for acquiring video data related to a historical event or a scientific phenomenon based on a historical event or a scientific phenomenon selected by a user; a means for converting the acquired video data to match the user's viewpoint and displaying the converted video data; A means for interacting with an AI chatbot that serves as a guide and provides the user with a detailed explanation of the selected historical event or scientific phenomenon; means for recognizing the user's emotions and customizing video and descriptions based thereon; a robot control means for automatically adjusting tasks based on the emotional state of the user; A system including: (Claim 2) 2. The system of claim 1, further comprising means for using a head mounted display for the user to visually experience the site of the historical event or the epicenter of the scientific phenomenon. (Claim 3) The system of claim 1 further comprising a means for constructing a database for providing additional information related to the historical event or scientific phenomenon selected by the user, and for the AI chatbot to act as a guide and provide a detailed explanation based on the database. (Claim 4) The system of claim 1, further comprising means for using a camera and a sensor to recognize the user's emotions. (Claim 5) 2. The system of claim 1, wherein the robot control means includes means for adjusting a working speed based on the emotion recognition information of the user. [Explanation of symbols]
[1467] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. means for acquiring video data related to a user-selected event, the event including at least one of a historical event and a scientific phenomenon; a means for converting the acquired video data to match the user's viewpoint and displaying the converted video data; a means for recognizing an emotion of the user using an emotion engine for recognizing an emotion, and customizing the video data according to the recognized emotion of the user; A means for interacting with an AI chatbot that acts as a guide and provides the user with a detailed explanation of the selected event; means for dynamically adjusting the video data and the explanatory content in response to a change in the user's emotion; A system including:
2. means for using a visual display device for allowing the user to visually experience the scene of the historical event; The system of claim 1 .
3. means for using a visual display device to allow the user to visually experience the core of the scientific phenomenon; The system of claim 1 .
4. A means for constructing a database for providing additional information related to the event selected by the user, and for the AI chatbot as a guide to provide a detailed explanation based on the database; The system of claim 1 .
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A