system
The system addresses the challenge of VR-AR interaction by collecting and analyzing VR user data with a generative AI model, enabling real-time, bidirectional interaction and shared experiences between VR and AR users.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-11-13
- Publication Date
- 2026-05-25
AI Technical Summary
Conventional virtual reality (VR) technologies face limitations in enabling real-time interaction and experience sharing between VR and augmented reality (AR) users, with VR users' interactions being difficult to convey to non-VR users, and both parties often unable to share the same experience.
A system that collects user interaction data in a VR environment, analyzes it using a generative AI model, and visualizes it in an AR environment, allowing AR users to influence the VR environment in real-time, thereby facilitating bidirectional interaction and shared experiences.
Enables real-time, bidirectional interaction and shared experiences between VR and AR users, allowing them to collaborate and enhance each other's experiences by reflecting AR users' actions in the VR environment.
Smart Images

Figure 2026085789000001_ABST
Abstract
Description
Technical Field
[0001] The technology of this disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a method for controlling a persona chatbot performed by at least one processor, the method including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance as a response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In conventional virtual reality (VR) technology, communication between VR experience users and non-VR experience users is limited, and it is difficult to share and interfere with VR experiences in real time with others. In particular, the interaction of VR users cannot be effectively conveyed to non-VR users, and both parties cannot share the same experience. Therefore, there is a need for a system that combines VR and augmented reality (AR) so that both can interfere with each other bidirectionally in real time.
Means for Solving the Problems
[0005] This invention provides an analysis means for collecting user interaction data in a virtual reality environment and visualizing it in an augmented reality environment. Furthermore, it includes means for reflecting operation data from the augmented reality user into the acquired virtual reality environment. In this system, the operations of the augmented reality user can affect objects in the virtual reality environment, realizing bidirectional real-time interaction. As a result, it becomes possible for VR users and non-VR users to share the same experience and cooperate with each other.
[0006] A "virtual reality environment" is an artificial three-dimensional space created using computer technology, where users can experience immersion through sight and sound.
[0007] "Interaction data" refers to information about actions and operations performed by a user within a virtual reality or augmented reality environment, and is data acquired by sensors and controllers.
[0008] An "augmented reality environment" is a technological environment that overlays computer-generated digital information onto visual information from the real world and presents it to the user, creating an environment where reality and virtuality merge.
[0009] "Analysis means" refers to a technology or method for analyzing acquired interaction data for a specific purpose and converting it into meaningful information.
[0010] "Visualization" is the process of representing data and information visually, enabling them to be presented in a way that is easy for users to understand.
[0011] "Operation data" refers to data that represents information about the actions and commands a user performs within an augmented reality environment, and is used to influence the virtual reality environment.
[0012] "Reflecting" refers to the process of making changes or adaptations to the virtual reality or augmented reality environment based on acquired information and data. [Brief explanation of the drawing]
[0013] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] This is a sequence diagram showing the processing flow of the data processing system in Example 2, which incorporates an emotion engine. [Figure 14] This is a sequence diagram showing the processing flow of the data processing system in Application Example 2, which combines an emotion engine. [Modes for carrying out the invention]
[0014] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.
[0015] First, the terms used in the following description will be explained.
[0016] In the following embodiments, the numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.
[0017] In the following embodiments, the numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor
[0018] In the following embodiments, the numbered storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, and the like.
[0019] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0020] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0021] [First Embodiment]
[0022] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0023] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0024] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0025] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0026] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0027] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0028] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0029] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0030] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0031] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0032] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0033] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0034] This invention provides a system that enables real-time information sharing and bidirectional interaction between a virtual reality (VR) environment and an augmented reality (AR) environment. Specific embodiments of the system are described below.
[0035] The user (VR user) wears a VR headset and controllers and interacts in a virtual reality environment. The VR user's movements and location information are transmitted to the VR terminal via wireless and wired connections, where interaction data is collected. The terminal converts this data into a standardized data format and sends it to the server.
[0036] After receiving this data, the server analyzes it using a generative AI model. This analysis converts user interactions in the VR environment into natural language or visual information, processing it into a format that is easy for AR users to understand.
[0037] The user (AR user) operates the AR device to visualize and confirm information sent from the server. The AR user can provide feedback to the VR environment through their own interface, inputting information for example through gestures or voice input. The AR device then standardizes this input data and sends it back to the server.
[0038] The server receives operation data from the AR user and reflects it in the VR environment. As a result, objects and situations in the virtual reality environment are updated based on the AR user's actions. For example, if the AR user directs an object in the VR environment with a gesture, the object will move in the VR user's field of view accordingly.
[0039] Thus, the system of the present invention makes it possible for VR users and AR users to simultaneously experience the same virtual environment and realize two-way communication in real time. Users can leverage each other's interactions and collaborate to obtain an experience that transcends the boundaries between virtual and reality.
[0040] The following describes the processing flow.
[0041] Step 1:
[0042] The user (VR user) engages in activities within a virtual reality environment using a VR device. The terminal acquires interaction data such as the user's location and gestures from the VR device and temporarily stores this data.
[0043] Step 2:
[0044] The terminal processes the stored interaction data in real time and converts it into a standardized data format. This standardized data is then prepared for transmission to the server.
[0045] Step 3:
[0046] The terminal sends the converted interaction data to the server. The server receives this data and begins the analysis process.
[0047] Step 4:
[0048] The server uses a generated AI model based on the received data to analyze the VR user's interactions. Based on the analysis results, it converts the data into visual or linguistic data.
[0049] Step 5:
[0050] The server sends the converted visual or linguistic data to the device (AR user side).
[0051] Step 6:
[0052] The user (AR user) receives information transmitted from the server using an AR device and confirms the visualized VR interaction. Based on this information, the user provides feedback to the virtual reality environment through gestures and voice.
[0053] Step 7:
[0054] The device (AR user side) acquires user feedback information and standardizes that data for transmission to the server.
[0055] Step 8:
[0056] The server receives feedback data from AR users and performs analysis. Based on the analysis results, it generates data that instructs the server to make necessary changes to the VR environment.
[0057] Step 9:
[0058] The server sends the generated instruction data to the terminal (VR user side) and prepares to update objects and actions in the virtual reality environment.
[0059] Step 10:
[0060] The terminal (VR user side) receives instruction data from the server and updates the VR environment. As a result, the user (VR user) experiences virtual reality that reflects the latest environmental changes.
[0061] (Example 1)
[0062] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0063] There is a need for technologies that enable users to share information bidirectionally and engage in real-time, collaborative interactions in both virtual and augmented reality spaces. However, current technologies face challenges in smoothly visualizing and reflecting information due to delays in information analysis and compatibility issues.
[0064] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0065] In this invention, the server includes a device for collecting user activity information in a virtual reality space, a processing device for analyzing the collected activity information and converting it into various types of information that can be displayed in an augmented reality space, a device for acquiring operation information from augmented reality users and reflecting the acquired operation information in the virtual reality space, and a device for analyzing received data using a generative AI model. This enables users to share information with each other in real time and apply timely and appropriate feedback in their respective real spaces.
[0066] A "virtual reality space" is an artificial three-dimensional environment created by a computer that users can experience through their sight and hearing.
[0067] "User activity information" refers to data about the actions, location, and other interactions that users take in the virtual reality space.
[0068] A "collection device" is a hardware and software system for detecting, storing, and transmitting activity information generated by users.
[0069] An "analytical processing device" is a software and hardware system that interprets collected activity information and processes it according to a specific purpose.
[0070] "Augmented reality" refers to a space that enhances the user's experience by overlaying digital information and images onto the real world's field of view.
[0071] "Various displayable information" refers to various types of information presented to the user visually or audibly based on analyzed activity information.
[0072] "Operation information" refers to data that includes actions, commands, and instructions performed by augmented reality users within the augmented reality space.
[0073] A "reflection device" is a device that applies the operation information transmitted by augmented reality users to the virtual reality space and updates the system state.
[0074] A "generative AI model" is an algorithm and software that uses machine learning techniques to analyze data and convert it into natural language or visual information.
[0075] This invention is a system that enables two-way real-time communication between users in virtual reality (VR) and augmented reality (AR) spaces. The entire system consists of a server, a terminal, and a user.
[0076] The server receives user interaction data transmitted from the VR terminal and analyzes it using a generative AI model. The generative AI model converts the user's actions and location information into natural language and visual information. This conversion process can be displayed, for example, as voice commands or text data.
[0077] The terminal collects sensor information from the headset and controllers worn by the VR user, converts it into a standardized data format, and sends it to the server. This data includes the user's hand movements and field of view, which forms the basis of the VR experience.
[0078] Users (AR users) can perceive changes in the VR space by viewing information displayed using AR devices. The AR device sends feedback to the server through voice commands and gestures, which is then reflected in the VR space. This process allows the AR user's instructions to influence the VR user's view and experience in real time.
[0079] A concrete example is a virtual museum tour in entertainment or education. VR users can move around within the virtual museum and observe the exhibits. AR users can experience the same tour in the real world and provide supplementary information through their AR devices. An example of a prompt would be a scenario where a VR user is exploring underwater ruins, and the AR user can add information such as "there is a shark behind the shipwreck."
[0080] This system allows users to leverage each other's interactions to create new experiential value, learning and having fun together in the space between reality and virtual reality.
[0081] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0082] Step 1:
[0083] The device collects information about the actions and location of VR users in the virtual reality space using a headset and controllers. These actions include pointing at and moving to virtual objects. The collected data is input as sensor information and converted into a standardized data format. This conversion prepares the data for consistent transmission to the server.
[0084] Step 2:
[0085] The server receives standardized data transmitted from the terminal and analyzes it using a generative AI model. The data provided as input is converted into natural language or visual information through the generative AI model. As a concrete example of this analysis, the hand movements of a VR user are translated into the action "the user selected an object." The analyzed information is prepared as output in an appropriate format for AR users.
[0086] Step 3:
[0087] The user (AR user) views the analyzed information sent from the server using an AR device. The AR device receives this information and presents it visually to the user. The AR user provides feedback from their environment using voice commands and gestures. This feedback data is standardized again by the device and converted into input data to be sent to the server.
[0088] Step 4:
[0089] The server receives and analyzes feedback from the AR device. Based on the input feedback, data calculations are performed to update objects and situations in the virtual reality space. For example, if an AR user instructs "move this object," that instruction is reflected in the VR space, and the object actually moves.
[0090] Step 5:
[0091] The user (VR user) experiences the VR space updated on the server and interacts again with the new situation that incorporates the feedback from the AR user. This allows the virtual environment to continuously update itself, maintaining real-time and two-way communication between users.
[0092] (Application Example 1)
[0093] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0094] The challenge lies in achieving more interactive, real-time, two-way communication between users in virtual and extended environments. In particular, there is a need for a system where information within the virtual environment is intuitively understood in the extended environment, and operations in the extended environment are immediately reflected in the virtual environment.
[0095] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0096] In this invention, the server includes means for collecting individual behavioral data in a virtual environment, means for analyzing the collected behavioral data and visualizing it as information in an extended environment, and means for acquiring operation data from users of the extended environment and reflecting the acquired operation data in the virtual environment. This integrates the virtual environment and the extended environment, enabling real-time bidirectional communication between users.
[0097] A "virtual environment" is a synthetic space created using digital technology that is different from reality, and it is a space in which users can have an immersive experience both visually and interactively.
[0098] "Behavioral data" refers to information that describes a series of actions and operations performed by an individual within a virtual and augmented environment, including actions, location, and voice.
[0099] An "augmented environment" is an environment created by overlaying digital information onto information from the real world, allowing users to bring virtual objects and information into their real-world field of vision.
[0100] "Visualizing information" means converting data and interactions into a form that is easily recognizable by users, and expressing them in the form of shapes, graphs, text, and other visual media.
[0101] "Operation data" refers to various types of information that users of the extended environment input via their devices, including gestures, voice input, and selection operations.
[0102] "To be integrated" refers to a state in which different environments or systems cooperate with each other to achieve unified functionality.
[0103] To implement this invention, the user wears smart glasses and browses products and services in a digital space. However, the store staff wear a virtual reality headset and provide customer service from a separate virtual space. The server has the function of collecting user behavior data from the smart glasses and converting it into a standardized data format. This data is processed on the server and parsed as natural language by a generative AI model (e.g., OpenAI®'s GPT-4®).
[0104] The server reconstructs these analysis results so that the user can visually obtain the information within the augmented environment, and outputs it to the smart glasses. In addition, gestures and voices made by the augmented reality user are sent to the server as operation data, which is analyzed by the generating AI model and immediately reflected in the virtual reality environment.
[0105] As a concrete example, suppose a customer is walking around a virtual store and notices a particular brand of shoes. The server acquires the customer's gaze information, uses a generative AI model to translate detailed information and history of the shoes into natural language, and displays it on the smart glasses. The store clerk can then review this information in VR and offer further advice or suggest other products. An example of a prompt might be, "When the customer looks at a particular pair of shoes, please provide relevant information."
[0106] In this way, users of servers, virtual environments, and extended environments can all achieve real-time, two-way communication, enabling efficient service delivery.
[0107] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0108] Step 1:
[0109] The device (smart glasses) captures the user's movements and gaze through cameras and sensors. This input data serves as basic information for identifying which object the user is focusing on. This information is collected as a digital signal and converted into a structured data format.
[0110] Step 2:
[0111] The device transfers structured behavioral data to the server. The server receives this data and analyzes it using a generative AI model. The analysis expresses information about the products the user is focusing their attention on in natural language. This output includes which items are attracting attention, as well as background information and whether similar products exist.
[0112] Step 3:
[0113] The server sends the analysis results back to the terminal. The terminal then performs data conversion to visualize the information on a digital display. This output is the information the user receives visually through smart glasses. The visualized information includes product details and recommended products.
[0114] Step 4:
[0115] If the user requests further information through gestures or voice, the device sends this interaction to the server as new input data. The server then analyzes this input again and generates new information corresponding to the interaction. This generated information is immediately reflected in the augmented reality space.
[0116] Step 5:
[0117] The server reflects augmented reality changes in the virtual reality environment. Within the virtual reality environment, the store clerk performs further tasks based on the user's requests and the products they are interested in. This output becomes an information environment shared by the user and the store clerk. Using a generative AI model, prompts are used to execute instructions such as, "If the customer looks at a specific pair of shoes, please provide relevant information."
[0118] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0119] This invention is a system that combines an emotion recognition engine to enhance interaction between virtual reality (VR) and augmented reality (AR) environments. By incorporating reactions that correspond to the user's emotional state, this system provides a more immersive experience.
[0120] The user (VR user) uses a VR device to perform activities within a virtual environment. The device collects interaction data from the VR user's actions and analyzes the user's emotions via an emotion engine. Based on this emotion analysis, the virtual reality environment generates feedback and content that corresponds to the user's emotions.
[0121] The user (AR user) uses an AR device to visualize the VR user's interaction information and emotional state transmitted from the server. Based on this information, the AR user provides feedback to the VR environment using gestures and voice commands. The AR device sends this feedback to the server, which generates instructions to update the VR environment.
[0122] The server uses an emotion engine to monitor the emotional states of both users and dynamically control elements in the VR and AR environments. For example, if a VR user is feeling anxious, the server adjusts the VR environment to present relaxing sounds and images.
[0123] For example, when a VR user encounters a scene that evokes awe, the emotion engine detects this reaction, and the server reflects an amplified effect in the VR environment. Simultaneously, the AR user can visualize this emotional change and provide appropriate feedback, thereby building a collaborative relationship with the VR user.
[0124] Thus, the present invention enables sophisticated interactions that respond to the user's emotions, providing an immersive experience that further transcends the boundaries between virtual and reality.
[0125] The following describes the processing flow.
[0126] Step 1:
[0127] The user (VR user) wears a VR headset and interacts within the VR environment. The device acquires interaction data such as the user's movements, gaze, and voice, and transmits it to the emotion engine in real time.
[0128] Step 2:
[0129] The emotion engine analyzes the acquired data to identify the user's emotional state. Based on these results, it generates instructions to adjust the content and effects in the VR environment.
[0130] Step 3:
[0131] The server updates the VR environment based on instructions received from the emotion engine. For example, if the user is feeling anxious, the server inserts calming background sounds and visuals.
[0132] Step 4:
[0133] The user (AR user) visualizes the VR user's emotional state and interaction information through the AR device. The device continuously updates this information in real time.
[0134] Step 5:
[0135] AR users provide appropriate feedback through gestures and voice commands based on the visualized information. This input is immediately sent to the server via the device.
[0136] Step 6:
[0137] The server receives and analyzes feedback from the AR user. Based on this analysis, it generates additional instructions that affect the VR environment and sends them to the device.
[0138] Step 7:
[0139] The terminal reflects new instructions from the server to the VR device, further updating the VR user's virtual environment. This allows the user to have a realistic and emotionally rich experience.
[0140] (Example 2)
[0141] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0142] Modern virtual environments suffer from a lack of feedback that reflects user emotional changes, resulting in a limited sense of immersion in the user experience. Furthermore, it is difficult to smoothly visualize the user's emotional state within the virtual environment in the augmented environment. This, in turn, restricts interaction between the augmented and virtual environments.
[0143] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0144] In this invention, the server includes means for acquiring data on the user's state from a measuring device, means for classifying the acquired data into emotional states using an analysis device, and means for applying changes generated based on the analysis results to the virtual environment. This makes it possible to dynamically adjust the virtual environment according to the user's emotions and provide a deeper sense of immersion.
[0145] A "user" refers to an entity that performs operations or experiences within a virtual or extended environment, and whose actions and emotional states are acquired and analyzed by the system.
[0146] "Status data" refers to information including user movements, actions, and physiological or psychological indicators, and is used to understand the user's current situation and reactions.
[0147] A "measuring device" refers to equipment or sensors that physically or digitally collect data about the user's state, enabling the system to measure the user's behavior and reactions.
[0148] An "analysis device" refers to a program or device that processes acquired data to identify the user's emotional state, and the virtual environment is adjusted based on the results.
[0149] "Emotional state" is a classification of a user's psychological response, and includes emotional categories such as joy, surprise, and relaxation.
[0150] A "virtual environment" refers to an artificially generated digital space in which users immerse themselves and interact with others.
[0151] An "extended environment" refers to a technology that provides users with visual and auditory information by overlaying digital information onto the physical world, enabling collaboration with virtual environments.
[0152] "Dynamic control" refers to real-time changes or adaptations, meaning a process in which system settings and content are adjusted according to the user's emotions and circumstances.
[0153] This invention provides a system that integrates a virtual environment and an extended environment in order to realize advanced interactions based on the user's emotional state.
[0154] The server utilizes measuring devices to acquire data about the user's state, including sensor devices such as VR and AR devices. This data includes physiological or behavioral information such as the user's movements, location, heart rate, and facial expressions. The terminal is responsible for aggregating this data and transmitting it to the server. The server processes the received data using analysis devices and uses a generative AI model to identify the user's emotional state. This model is built, for example, using a machine learning algorithm, and recognizes emotions by comparing them with pre-trained emotional pattern data.
[0155] Based on the results of emotion analysis, the server generates instructions to dynamically adjust the settings of the virtual environment. For example, if the user expresses the emotion of "surprise," the server will enhance the visual effects within the virtual environment and adjust the sound effects accordingly. It can also change the background music to enhance relaxation.
[0156] Meanwhile, users of AR devices receive real-time information on the VR user's emotional state and environmental changes sent from the server. This information is visually presented in the augmented environment, and AR users can interact with the VR environment through voice commands and gestures. The device promptly sends this feedback to the server, ensuring that the environment is updated in response to the user's feedback.
[0157] As a concrete example, by inputting the following prompt into the generating AI model: "Please propose a method to analyze the emotions of a VR user upon entering a new environment in real time and provide optimized sound and visuals based on the results," this system can make the user experience richer and more immersive. In this way, the present invention realizes a new form of interaction between the virtual environment and the real world.
[0158] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0159] Step 1:
[0160] The user puts on a VR device and begins activities within the virtual environment. The terminal detects the user's movements and actions and collects interaction data in real time. Inputs include user movement information, location information, and physiological data obtained from sensors. The terminal sends this as a single dataset to the server. The output is a detailed dataset showing the user's state.
[0161] Step 2:
[0162] The server processes the received interaction data using an analysis device. It uses a generative AI model to analyze the data and identify the user's emotional state. The input is interaction data sent from the VR device. The server feeds this data into the AI model and identifies the emotional category by comparing it with pre-trained emotional patterns. The output is the specific emotional state the user is experiencing (e.g., "joy," "excitement," "relaxation," etc.).
[0163] Step 3:
[0164] The server generates instructions to dynamically adjust the virtual environment settings based on the analyzed emotional state. The input is the results of the emotional analysis. Based on this data, the server issues specific instructions to adjust the lighting, sound, and other visual effects within the virtual environment. The output is the adjusted virtual environment settings.
[0165] Step 4:
[0166] Users of AR devices receive information about the VR user's emotional state and environmental changes sent from the server. Inputs include information about changes in the VR environment and emotional data sent from the server. The device visualizes this information in real time and presents it to the AR user. Outputs are change information that the user can confirm through visual and audio commands.
[0167] Step 5:
[0168] AR users provide feedback using voice commands and gestures based on visualized information. The device sends this feedback information to the server. Inputs include the AR user's gesture data and voice commands. The device quickly transmits this information to the server, providing further information for adjusting the virtual environment. Outputs are the feedback data received by the server.
[0169] In this way, servers, terminals, and users collaborate with each other to achieve sophisticated, emotion-based interactions.
[0170] (Application Example 2)
[0171] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".
[0172] Traditional virtual reality and augmented reality experiences have failed to fully adapt to the user's emotional state, making it difficult to dynamically customize individual experiences. Furthermore, interaction within the virtual environment is limited to sight and sound, hindering user immersion. This resulted in users having to continue the experience while emotionally unresponsive to certain scenes and events, preventing them from achieving deeper emotional impact and immersion.
[0173] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0174] In this invention, the server includes means for collecting user behavior data in a virtual environment, means for analyzing the collected behavior data and converting it into various types of information to be visualized in an extended environment, means for acquiring operation data from users of the extended environment and reflecting the acquired operation data in the virtual environment, and means for analyzing the user's emotional state using an emotion recognition engine and dynamically adjusting the audio and video effects of the virtual environment based on the analysis results. This enables real-time feedback and environmental adjustments in response to the user's emotions towards the content being viewed, allowing for a deeper sense of immersion.
[0175] A "virtual environment" is a digital realm distinct from the real world, constructed using computer technology, within which users can have visual and auditory experiences.
[0176] "User" refers to an individual who experiences and operates the virtual and extended environments through this system.
[0177] "Behavioral data" refers to information about various operations and actions performed by users within a virtual environment, and the system uses this data to analyze user behavior and intentions.
[0178] An "augmented environment" refers to an environment that enhances the user experience by overlaying digital information onto the real world, thereby adding virtual elements to the actual field of view.
[0179] "Visualization" is the process of representing data and information in a way that is easy for users to understand, thereby enabling digital information to be displayed effectively on a screen.
[0180] "Operation data" refers to the records of various operations performed by the user in the extended environment, and this data forms the basis for feedback responses to the virtual environment.
[0181] An "emotion recognition engine" is a software or hardware component that analyzes a user's facial expressions, voice tone, and other behavioral patterns to understand their emotional state.
[0182] "Audio and video effects" refers to sound and visual effects used to enrich the experience and enhance immersion within a virtual environment.
[0183] "Dynamic adjustment" refers to a process in which the behavior and settings of a system are automatically and in real time changed in response to specific user conditions or changes in the environment.
[0184] The system for realizing this application consists of a server, a terminal equipped with a VR / AR device, and an emotion recognition engine. The server collects user behavior data and analyzes the user's emotional state in real time through the emotion recognition engine. Based on this analysis, the server dynamically adjusts the audio and video effects of the virtual environment. The server also receives operation data from the augmented environment user and reflects it as feedback in the virtual environment.
[0185] The device uses a smartphone or head-mounted display (e.g., a typical VR headset) to provide the user with visual and auditory information. This allows the user to immerse themselves in the virtual environment and have a personalized experience with the content they view. The device also features cameras and sensors to detect the user's facial expressions and actions for an emotion recognition engine.
[0186] For example, if a user expresses surprise during a movie viewing, the server can enhance the movie's soundtrack and visual effects accordingly. This allows users to have a more immersive experience. Furthermore, interactive features are enabled that allow users to share the same experience and exchange opinions based on their emotions.
[0187] In relation to the generative AI model, this system analyzes emotions using prompts such as, "If the user is surprised by a specific scene in a movie, dynamically change the effects to match the emotion," and adjusts various elements within the virtual environment accordingly.
[0188] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0189] Step 1:
[0190] The device collects behavioral and emotional data of the user during their VR / AR experience in real time. This process involves capturing facial expressions and movements using cameras and sensors, which are then used as initial data. The input consists of the user's movements and facial expressions, while the output is corresponding behavioral and emotional data.
[0191] Step 2:
[0192] The server receives the collected behavioral and emotional data and analyzes it using an emotion recognition engine. This extracts the user's emotional state (e.g., surprise, joy, fear). The input is the raw data collected by the terminal, and the output is the user's emotional state information obtained through the analysis.
[0193] Step 3:
[0194] The server uses a generative AI model to select appropriate feedback based on the analysis results and prepares corresponding audio and video effects. The prompt message "When the user is surprised in a specific scene of the movie, dynamically change the effects to match the emotion" is used to determine the direction of the analysis. The input is emotional state information, and the output is the selected feedback content.
[0195] Step 4:
[0196] The server sends the prepared feedback to the terminal, instructing it to dynamically adjust the audio and video within the virtual environment. The terminal receives this instruction and adjusts the effects on the content being played in real time. The input is the selected feedback, and the output is the playback of the adjusted content.
[0197] Step 5:
[0198] Users experience newly adjusted audio and video effects in a virtual environment, creating an immersive experience. If the user's emotions change, the process returns to step 1 and is repeated. The input is the adjusted content, and the output is the user's experience.
[0199] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0200] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0201] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0202] [Second Embodiment]
[0203] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0204] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0205] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0206] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0207] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0208] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0209] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0210] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0211] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0212] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0213] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0214] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0215] This invention provides a system that enables real-time information sharing and bidirectional interaction between a virtual reality (VR) environment and an augmented reality (AR) environment. Specific embodiments of the system are described below.
[0216] The user (VR user) wears a VR headset and controllers and interacts in a virtual reality environment. The VR user's movements and location information are transmitted to the VR terminal via wireless and wired connections, where interaction data is collected. The terminal converts this data into a standardized data format and sends it to the server.
[0217] After receiving this data, the server analyzes it using a generative AI model. This analysis converts user interactions in the VR environment into natural language or visual information, processing it into a format that is easy for AR users to understand.
[0218] The user (AR user) operates the AR device to visualize and confirm information sent from the server. The AR user can provide feedback to the VR environment through their own interface, inputting information for example through gestures or voice input. The AR device then standardizes this input data and sends it back to the server.
[0219] The server receives operation data from the AR user and reflects it in the VR environment. As a result, objects and situations in the virtual reality environment are updated based on the AR user's actions. For example, if the AR user directs an object in the VR environment with a gesture, the object will move in the VR user's field of view accordingly.
[0220] Thus, the system of the present invention makes it possible for VR users and AR users to simultaneously experience the same virtual environment and realize two-way communication in real time. Users can leverage each other's interactions and collaborate to obtain an experience that transcends the boundaries between virtual and reality.
[0221] The following describes the processing flow.
[0222] Step 1:
[0223] The user (VR user) engages in activities within a virtual reality environment using a VR device. The terminal acquires interaction data such as the user's location and gestures from the VR device and temporarily stores this data.
[0224] Step 2:
[0225] The terminal processes the stored interaction data in real time and converts it into a standardized data format. This standardized data is then prepared for transmission to the server.
[0226] Step 3:
[0227] The terminal sends the converted interaction data to the server. The server receives this data and begins the analysis process.
[0228] Step 4:
[0229] The server uses a generated AI model based on the received data to analyze the VR user's interactions. Based on the analysis results, it converts the data into visual or linguistic data.
[0230] Step 5:
[0231] The server sends the converted visual or linguistic data to the device (AR user side).
[0232] Step 6:
[0233] The user (AR user) receives information transmitted from the server using an AR device and confirms the visualized VR interaction. Based on this information, the user provides feedback to the virtual reality environment through gestures and voice.
[0234] Step 7:
[0235] The device (AR user side) acquires user feedback information and standardizes that data for transmission to the server.
[0236] Step 8:
[0237] The server receives feedback data from AR users and performs analysis. Based on the analysis results, it generates data that instructs the server to make necessary changes to the VR environment.
[0238] Step 9:
[0239] The server sends the generated instruction data to the terminal (VR user side) and prepares to update objects and actions in the virtual reality environment.
[0240] Step 10:
[0241] The terminal (VR user side) receives instruction data from the server and updates the VR environment. As a result, the user (VR user) experiences virtual reality that reflects the latest environmental changes.
[0242] (Example 1)
[0243] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0244] There is a need for technologies that enable users to share information bidirectionally and engage in real-time, collaborative interactions in both virtual and augmented reality spaces. However, current technologies face challenges in smoothly visualizing and reflecting information due to delays in information analysis and compatibility issues.
[0245] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0246] In this invention, the server includes a device for collecting user activity information in a virtual reality space, a processing device for analyzing the collected activity information and converting it into various types of information that can be displayed in an augmented reality space, a device for acquiring operation information from augmented reality users and reflecting the acquired operation information in the virtual reality space, and a device for analyzing received data using a generative AI model. This enables users to share information with each other in real time and apply timely and appropriate feedback in their respective real spaces.
[0247] A "virtual reality space" is an artificial three-dimensional environment created by a computer that users can experience through their sight and hearing.
[0248] "User activity information" refers to data about the actions, location, and other interactions that users take in the virtual reality space.
[0249] A "collection device" is a hardware and software system for detecting, storing, and transmitting activity information generated by users.
[0250] An "analytical processing device" is a software and hardware system that interprets collected activity information and processes it according to a specific purpose.
[0251] "Augmented reality" refers to a space that enhances the user's experience by overlaying digital information and images onto the real world's field of view.
[0252] "Various displayable information" refers to various types of information presented to the user visually or audibly based on analyzed activity information.
[0253] "Operation information" refers to data that includes actions, commands, and instructions performed by augmented reality users within the augmented reality space.
[0254] A "reflection device" is a device that applies the operation information transmitted by augmented reality users to the virtual reality space and updates the system state.
[0255] A "generative AI model" is an algorithm and software that uses machine learning techniques to analyze data and convert it into natural language or visual information.
[0256] This invention is a system that enables two-way real-time communication between users in virtual reality (VR) and augmented reality (AR) spaces. The entire system consists of a server, a terminal, and a user.
[0257] The server receives user interaction data transmitted from the VR terminal and analyzes it using a generative AI model. The generative AI model converts the user's actions and location information into natural language and visual information. This conversion process can be displayed, for example, as voice commands or text data.
[0258] The terminal collects sensor information from the headset and controllers worn by the VR user, converts it into a standardized data format, and sends it to the server. This data includes the user's hand movements and field of view, which forms the basis of the VR experience.
[0259] Users (AR users) can perceive changes in the VR space by viewing information displayed using AR devices. The AR device sends feedback to the server through voice commands and gestures, which is then reflected in the VR space. This process allows the AR user's instructions to influence the VR user's view and experience in real time.
[0260] A concrete example is a virtual museum tour in entertainment or education. VR users can move around within the virtual museum and observe the exhibits. AR users can experience the same tour in the real world and provide supplementary information through their AR devices. An example of a prompt would be a scenario where a VR user is exploring underwater ruins, and the AR user can add information such as "there is a shark behind the shipwreck."
[0261] This system allows users to leverage each other's interactions to create new experiential value, learning and having fun together in the space between reality and virtual reality.
[0262] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0263] Step 1:
[0264] The device collects information about the actions and location of VR users in the virtual reality space using a headset and controllers. These actions include pointing at and moving to virtual objects. The collected data is input as sensor information and converted into a standardized data format. This conversion prepares the data for consistent transmission to the server.
[0265] Step 2:
[0266] The server receives standardized data transmitted from the terminal and analyzes it using a generative AI model. The data provided as input is converted into natural language or visual information through the generative AI model. As a concrete example of this analysis, the hand movements of a VR user are translated into the action "the user selected an object." The analyzed information is prepared as output in an appropriate format for AR users.
[0267] Step 3:
[0268] The user (AR user) views the analyzed information sent from the server using an AR device. The AR device receives this information and presents it visually to the user. The AR user provides feedback from their environment using voice commands and gestures. This feedback data is standardized again by the device and converted into input data to be sent to the server.
[0269] Step 4:
[0270] The server receives and analyzes feedback from the AR device. Based on the input feedback, data calculations are performed to update objects and situations in the virtual reality space. For example, if an AR user instructs "move this object," that instruction is reflected in the VR space, and the object actually moves.
[0271] Step 5:
[0272] The user (VR user) experiences the VR space updated on the server and interacts again with the new situation that incorporates the feedback from the AR user. This allows the virtual environment to continuously update itself, maintaining real-time and two-way communication between users.
[0273] (Application Example 1)
[0274] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0275] The challenge lies in achieving more interactive, real-time, two-way communication between users in virtual and extended environments. In particular, there is a need for a system where information within the virtual environment is intuitively understood in the extended environment, and operations in the extended environment are immediately reflected in the virtual environment.
[0276] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0277] In this invention, the server includes means for collecting individual behavioral data in a virtual environment, means for analyzing the collected behavioral data and visualizing it as information in an extended environment, and means for acquiring operation data from users of the extended environment and reflecting the acquired operation data in the virtual environment. This integrates the virtual environment and the extended environment, enabling real-time bidirectional communication between users.
[0278] A "virtual environment" is a synthetic space created using digital technology that is different from reality, and it is a space in which users can have an immersive experience both visually and interactively.
[0279] "Behavioral data" refers to information that describes a series of actions and operations performed by an individual within a virtual and augmented environment, including actions, location, and voice.
[0280] The "extended environment" is an environment realized by overlaying digital information on real-world information, and is a space where users can bring virtual objects and information into their real field of vision.
[0281] "Visualize as information" means converting data and interactions into a form that is easy for users to visually recognize, and is expressed in forms such as graphics, graphs, and text.
[0282] "Operation data" refers to various types of information input by users of the extended environment through devices, and includes gestures, voice inputs, selection operations, etc.
[0283] "Integrated" refers to a state where different environments and systems cooperate with each other to realize an integrated function.
[0284] To implement this invention, the user wears smart glasses and browses products and services in the digital space. However, the store clerk wears a virtual reality headset and serves customers from another virtual space. The server has the function of collecting the user's motion data from the smart glasses and converting it into a standardized data format. This data is processed by the server and analyzed as natural language by a generative AI model (e.g., GPT-4 of OpenAI).
[0285] The server reconstructs this analysis result so that the user can visually obtain information within the extended environment and outputs it to the smart glasses. Also, gestures and voices made by the extended reality user are sent to the server as operation data, and after being analyzed by the generative AI model, they are immediately reflected within the virtual reality environment.
[0286] As a specific example, suppose a customer notices shoes of a particular brand while walking around in a virtual store. The server acquires the customer's line-of-sight information, naturalizes the detailed information and history of the shoes using a generative AI model, and displays it on smart glasses. The store clerk can check the information in VR and offer further advice or recommend other products. An example of a prompt sentence is "When the customer looks at a specific pair of shoes, please present the relevant information."
[0287] In this way, users of the server, virtual environment, and augmented environment can achieve two-way communication in real time together, enabling efficient service provision.
[0288] The flow of the specific process in Application Example 1 will be described using FIG. 12.
[0289] Step 1:
[0290] The terminal (smart glasses) captures the user's actions and line of sight through cameras and sensors. This input data serves as basic information for identifying which object the user is focusing on. This information is collected as a digital signal and converted into a structured data format.
[0291] Step 2:
[0292] The terminal transfers the structured action data to the server. The server receives this data and performs analysis using a generative AI model. Through the analysis, information about the product on which the user's line of sight is concentrated is expressed in natural language. This output includes which item is being focused on, as well as its background information and whether there are similar products.
[0293] Step 3:
[0294] The server sends the analysis result back to the terminal. At this time, the terminal performs data conversion for visualizing the information on a digital display. This output is the information that the user receives visually through the smart glasses. The visualized information includes details about the product and recommended products, etc.
[0295] Step 4:
[0296] If the user requests further information through gestures or voice, the device sends this interaction to the server as new input data. The server then analyzes this input again and generates new information corresponding to the interaction. This generated information is immediately reflected in the augmented reality space.
[0297] Step 5:
[0298] The server reflects augmented reality changes in the virtual reality environment. Within the virtual reality environment, the store clerk performs further tasks based on the user's requests and the products they are interested in. This output becomes an information environment shared by the user and the store clerk. Using a generative AI model, prompts are used to execute instructions such as, "If the customer looks at a specific pair of shoes, please provide relevant information."
[0299] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0300] This invention is a system that combines an emotion recognition engine to enhance interaction between virtual reality (VR) and augmented reality (AR) environments. By incorporating reactions that correspond to the user's emotional state, this system provides a more immersive experience.
[0301] The user (VR user) uses a VR device to perform activities within a virtual environment. The device collects interaction data from the VR user's actions and analyzes the user's emotions via an emotion engine. Based on this emotion analysis, the virtual reality environment generates feedback and content that corresponds to the user's emotions.
[0302] The user (AR user) uses an AR device to visualize the interaction information and emotional state of the VR user sent from the server. Based on this information, the AR user provides feedback to the VR environment through gestures and voice commands. The AR terminal sends this feedback to the server, and the server generates instructions for updating the VR environment.
[0303] The server uses an emotion engine to monitor the emotional states of both users and dynamically control each element in the VR and AR environments. For example, when the VR user is tense, the server adjusts the VR environment to present sounds and videos with a relaxing effect.
[0304] As a specific example, when the VR user faces a scene of amazement, the emotion engine detects the reaction, and the server reflects the amplified effect in the VR environment. At the same time, the AR user visualizes the emotional change and can establish a cooperation system with the VR user by providing appropriate feedback.
[0305] Thus, according to the present invention, advanced interaction according to the user's emotion becomes possible, providing an immersive experience that further transcends the boundary between virtual and reality.
[0306] The following describes the processing flow.
[0307] Step 1:
[0308] The user (VR user) wears a VR headset and interacts within the VR environment. The terminal acquires interaction data such as the user's actions, gaze, and voice, and sends it to the emotion engine in real time.
[0309] [[ID=二七]] Step 2:
[0310] The emotion engine analyzes the acquired data to identify the user's emotional state. Based on this result, it generates instructions for adjusting the content and effects in the VR environment.
[0311] Step 3:
[0312] The server updates the VR environment based on instructions received from the emotion engine. For example, if the user is feeling anxious, the server inserts calming background sounds and visuals.
[0313] Step 4:
[0314] The user (AR user) visualizes the VR user's emotional state and interaction information through the AR device. The device continuously updates this information in real time.
[0315] Step 5:
[0316] AR users provide appropriate feedback through gestures and voice commands based on the visualized information. This input is immediately sent to the server via the device.
[0317] Step 6:
[0318] The server receives and analyzes feedback from the AR user. Based on this analysis, it generates additional instructions that affect the VR environment and sends them to the device.
[0319] Step 7:
[0320] The terminal reflects new instructions from the server to the VR device, further updating the VR user's virtual environment. This allows the user to have a realistic and emotionally rich experience.
[0321] (Example 2)
[0322] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0323] Modern virtual environments suffer from a lack of feedback that reflects user emotional changes, resulting in a limited sense of immersion in the user experience. Furthermore, it is difficult to smoothly visualize the user's emotional state within the virtual environment in the augmented environment. This, in turn, restricts interaction between the augmented and virtual environments.
[0324] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0325] In this invention, the server includes means for acquiring data on the user's state from a measuring device, means for classifying the acquired data into emotional states using an analysis device, and means for applying changes generated based on the analysis results to the virtual environment. This makes it possible to dynamically adjust the virtual environment according to the user's emotions and provide a deeper sense of immersion.
[0326] A "user" refers to an entity that performs operations or experiences within a virtual or extended environment, and whose actions and emotional states are acquired and analyzed by the system.
[0327] "Status data" refers to information including user movements, actions, and physiological or psychological indicators, and is used to understand the user's current situation and reactions.
[0328] A "measuring device" refers to equipment or sensors that physically or digitally collect data about the user's state, enabling the system to measure the user's behavior and reactions.
[0329] An "analysis device" refers to a program or device that processes acquired data to identify the user's emotional state, and the virtual environment is adjusted based on the results.
[0330] "Emotional state" is a classification of a user's psychological response, and includes emotional categories such as joy, surprise, and relaxation.
[0331] A "virtual environment" refers to an artificially generated digital space in which users immerse themselves and interact with others.
[0332] An "extended environment" refers to a technology that provides users with visual and auditory information by overlaying digital information onto the physical world, enabling collaboration with virtual environments.
[0333] "Dynamic control" refers to real-time changes or adaptations, meaning a process in which system settings and content are adjusted according to the user's emotions and circumstances.
[0334] This invention provides a system that integrates a virtual environment and an extended environment in order to realize advanced interactions based on the user's emotional state.
[0335] The server utilizes measuring devices to acquire data about the user's state, including sensor devices such as VR and AR devices. This data includes physiological or behavioral information such as the user's movements, location, heart rate, and facial expressions. The terminal is responsible for aggregating this data and transmitting it to the server. The server processes the received data using analysis devices and uses a generative AI model to identify the user's emotional state. This model is built, for example, using a machine learning algorithm, and recognizes emotions by comparing them with pre-trained emotional pattern data.
[0336] Based on the results of emotion analysis, the server generates instructions to dynamically adjust the settings of the virtual environment. For example, if the user expresses the emotion of "surprise," the server will enhance the visual effects within the virtual environment and adjust the sound effects accordingly. It can also change the background music to enhance relaxation.
[0337] Meanwhile, users of AR devices receive real-time information on the VR user's emotional state and environmental changes sent from the server. This information is visually presented in the augmented environment, and AR users can interact with the VR environment through voice commands and gestures. The device promptly sends this feedback to the server, ensuring that the environment is updated in response to the user's feedback.
[0338] As a concrete example, by inputting the following prompt into the generating AI model: "Please propose a method to analyze the emotions of a VR user upon entering a new environment in real time and provide optimized sound and visuals based on the results," this system can make the user experience richer and more immersive. In this way, the present invention realizes a new form of interaction between the virtual environment and the real world.
[0339] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0340] Step 1:
[0341] The user puts on a VR device and begins activities within the virtual environment. The terminal detects the user's movements and actions and collects interaction data in real time. Inputs include user movement information, location information, and physiological data obtained from sensors. The terminal sends this as a single dataset to the server. The output is a detailed dataset showing the user's state.
[0342] Step 2:
[0343] The server processes the received interaction data using an analysis device. It uses a generative AI model to analyze the data and identify the user's emotional state. The input is interaction data sent from the VR device. The server feeds this data into the AI model and identifies the emotional category by comparing it with pre-trained emotional patterns. The output is the specific emotional state the user is experiencing (e.g., "joy," "excitement," "relaxation," etc.).
[0344] Step 3:
[0345] The server generates instructions to dynamically adjust the virtual environment settings based on the analyzed emotional state. The input is the results of the emotional analysis. Based on this data, the server issues specific instructions to adjust the lighting, sound, and other visual effects within the virtual environment. The output is the adjusted virtual environment settings.
[0346] Step 4:
[0347] Users of AR devices receive information about the VR user's emotional state and environmental changes sent from the server. Inputs include information about changes in the VR environment and emotional data sent from the server. The device visualizes this information in real time and presents it to the AR user. Outputs are change information that the user can confirm through visual and audio commands.
[0348] Step 5:
[0349] AR users provide feedback using voice commands and gestures based on visualized information. The device sends this feedback information to the server. Inputs include the AR user's gesture data and voice commands. The device quickly transmits this information to the server, providing further information for adjusting the virtual environment. Outputs are the feedback data received by the server.
[0350] In this way, servers, terminals, and users collaborate with each other to achieve sophisticated, emotion-based interactions.
[0351] (Application Example 2)
[0352] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the smart glasses 214 as the "terminal".
[0353] Traditional virtual reality and augmented reality experiences have failed to fully adapt to the user's emotional state, making it difficult to dynamically customize individual experiences. Furthermore, interaction within the virtual environment is limited to sight and sound, hindering user immersion. This resulted in users having to continue the experience while emotionally unresponsive to certain scenes and events, preventing them from achieving deeper emotional impact and immersion.
[0354] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0355] In this invention, the server includes means for collecting user behavior data in a virtual environment, means for analyzing the collected behavior data and converting it into various types of information to be visualized in an extended environment, means for acquiring operation data from users of the extended environment and reflecting the acquired operation data in the virtual environment, and means for analyzing the user's emotional state using an emotion recognition engine and dynamically adjusting the audio and video effects of the virtual environment based on the analysis results. This enables real-time feedback and environmental adjustments in response to the user's emotions towards the content being viewed, allowing for a deeper sense of immersion.
[0356] A "virtual environment" is a digital realm distinct from the real world, constructed using computer technology, within which users can have visual and auditory experiences.
[0357] "User" refers to an individual who experiences and operates the virtual and extended environments through this system.
[0358] "Behavioral data" refers to information about various operations and actions performed by users within a virtual environment, and the system uses this data to analyze user behavior and intentions.
[0359] An "augmented environment" refers to an environment that enhances the user experience by overlaying digital information onto the real world, thereby adding virtual elements to the actual field of view.
[0360] "Visualization" is the process of representing data and information in a way that is easy for users to understand, thereby enabling digital information to be displayed effectively on a screen.
[0361] "Operation data" refers to the records of various operations performed by the user in the extended environment, and this data forms the basis for feedback responses to the virtual environment.
[0362] An "emotion recognition engine" is a software or hardware component that analyzes a user's facial expressions, voice tone, and other behavioral patterns to understand their emotional state.
[0363] "Audio and video effects" refers to sound and visual effects used to enrich the experience and enhance immersion within a virtual environment.
[0364] "Dynamic adjustment" refers to a process in which the behavior and settings of a system are automatically and in real time changed in response to specific user conditions or changes in the environment.
[0365] The system for realizing this application consists of a server, a terminal equipped with a VR / AR device, and an emotion recognition engine. The server collects user behavior data and analyzes the user's emotional state in real time through the emotion recognition engine. Based on this analysis, the server dynamically adjusts the audio and video effects of the virtual environment. The server also receives operation data from the augmented environment user and reflects it as feedback in the virtual environment.
[0366] The device uses a smartphone or head-mounted display (e.g., a typical VR headset) to provide the user with visual and auditory information. This allows the user to immerse themselves in the virtual environment and have a personalized experience with the content they view. The device also features cameras and sensors to detect the user's facial expressions and actions for an emotion recognition engine.
[0367] For example, if a user expresses surprise during a movie viewing, the server can enhance the movie's soundtrack and visual effects accordingly. This allows users to have a more immersive experience. Furthermore, interactive features are enabled that allow users to share the same experience and exchange opinions based on their emotions.
[0368] In relation to the generative AI model, this system analyzes emotions using prompts such as, "If the user is surprised by a specific scene in a movie, dynamically change the effects to match the emotion," and adjusts various elements within the virtual environment accordingly.
[0369] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0370] Step 1:
[0371] The device collects behavioral and emotional data of the user during their VR / AR experience in real time. This process involves capturing facial expressions and movements using cameras and sensors, which are then used as initial data. The input consists of the user's movements and facial expressions, while the output is corresponding behavioral and emotional data.
[0372] Step 2:
[0373] The server receives the collected behavioral and emotional data and analyzes it using an emotion recognition engine. This extracts the user's emotional state (e.g., surprise, joy, fear). The input is the raw data collected by the terminal, and the output is the user's emotional state information obtained through the analysis.
[0374] Step 3:
[0375] The server uses a generative AI model to select appropriate feedback based on the analysis results and prepares corresponding audio and video effects. The prompt message "When the user is surprised in a specific scene of the movie, dynamically change the effects to match the emotion" is used to determine the direction of the analysis. The input is emotional state information, and the output is the selected feedback content.
[0376] Step 4:
[0377] The server sends the prepared feedback to the terminal, instructing it to dynamically adjust the audio and video within the virtual environment. The terminal receives this instruction and adjusts the effects on the content being played in real time. The input is the selected feedback, and the output is the playback of the adjusted content.
[0378] Step 5:
[0379] Users experience newly adjusted audio and video effects in a virtual environment, creating an immersive experience. If the user's emotions change, the process returns to step 1 and is repeated. The input is the adjusted content, and the output is the user's experience.
[0380] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0381] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0382] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0383] [Third Embodiment]
[0384] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0385] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0386] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0387] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0388] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0389] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0390] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0391] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0392] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0393] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0394] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0395] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0396] This invention provides a system that enables real-time information sharing and bidirectional interaction between a virtual reality (VR) environment and an augmented reality (AR) environment. Specific embodiments of the system are described below.
[0397] The user (VR user) wears a VR headset and controllers and interacts in a virtual reality environment. The VR user's movements and location information are transmitted to the VR terminal via wireless and wired connections, where interaction data is collected. The terminal converts this data into a standardized data format and sends it to the server.
[0398] After receiving this data, the server analyzes it using a generative AI model. This analysis converts user interactions in the VR environment into natural language or visual information, processing it into a format that is easy for AR users to understand.
[0399] The user (AR user) operates the AR device to visualize and confirm information sent from the server. The AR user can provide feedback to the VR environment through their own interface, inputting information for example through gestures or voice input. The AR device then standardizes this input data and sends it back to the server.
[0400] The server receives operation data from the AR user and reflects it in the VR environment. As a result, objects and situations in the virtual reality environment are updated based on the AR user's actions. For example, if the AR user directs an object in the VR environment with a gesture, the object will move in the VR user's field of view accordingly.
[0401] Thus, the system of the present invention makes it possible for VR users and AR users to simultaneously experience the same virtual environment and realize two-way communication in real time. Users can leverage each other's interactions and collaborate to obtain an experience that transcends the boundaries between virtual and reality.
[0402] The following describes the processing flow.
[0403] Step 1:
[0404] The user (VR user) engages in activities within a virtual reality environment using a VR device. The terminal acquires interaction data such as the user's location and gestures from the VR device and temporarily stores this data.
[0405] Step 2:
[0406] The terminal processes the stored interaction data in real time and converts it into a standardized data format. This standardized data is then prepared for transmission to the server.
[0407] Step 3:
[0408] The terminal sends the converted interaction data to the server. The server receives this data and begins the analysis process.
[0409] Step 4:
[0410] The server uses a generated AI model based on the received data to analyze the VR user's interactions. Based on the analysis results, it converts the data into visual or linguistic data.
[0411] Step 5:
[0412] The server sends the converted visual or linguistic data to the device (AR user side).
[0413] Step 6:
[0414] The user (AR user) receives information transmitted from the server using an AR device and confirms the visualized VR interaction. Based on this information, the user provides feedback to the virtual reality environment through gestures and voice.
[0415] Step 7:
[0416] The device (AR user side) acquires user feedback information and standardizes that data for transmission to the server.
[0417] Step 8:
[0418] The server receives feedback data from AR users and performs analysis. Based on the analysis results, it generates data that instructs the server to make necessary changes to the VR environment.
[0419] Step 9:
[0420] The server sends the generated instruction data to the terminal (VR user side) and prepares to update objects and actions in the virtual reality environment.
[0421] Step 10:
[0422] The terminal (VR user side) receives instruction data from the server and updates the VR environment. As a result, the user (VR user) experiences virtual reality that reflects the latest environmental changes.
[0423] (Example 1)
[0424] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0425] There is a need for technologies that enable users to share information bidirectionally and engage in real-time, collaborative interactions in both virtual and augmented reality spaces. However, current technologies face challenges in smoothly visualizing and reflecting information due to delays in information analysis and compatibility issues.
[0426] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0427] In this invention, the server includes a device for collecting user activity information in a virtual reality space, a processing device for analyzing the collected activity information and converting it into various types of information that can be displayed in an augmented reality space, a device for acquiring operation information from augmented reality users and reflecting the acquired operation information in the virtual reality space, and a device for analyzing received data using a generative AI model. This enables users to share information with each other in real time and apply timely and appropriate feedback in their respective real spaces.
[0428] A "virtual reality space" is an artificial three-dimensional environment created by a computer that users can experience through their sight and hearing.
[0429] "User activity information" refers to data about the actions, location, and other interactions that users take in the virtual reality space.
[0430] A "collection device" is a hardware and software system for detecting, storing, and transmitting activity information generated by users.
[0431] An "analytical processing device" is a software and hardware system that interprets collected activity information and processes it according to a specific purpose.
[0432] "Augmented reality" refers to a space that enhances the user's experience by overlaying digital information and images onto the real world's field of view.
[0433] "Various displayable information" refers to various types of information presented to the user visually or audibly based on analyzed activity information.
[0434] "Operation information" refers to data that includes actions, commands, and instructions performed by augmented reality users within the augmented reality space.
[0435] A "reflection device" is a device that applies the operation information transmitted by augmented reality users to the virtual reality space and updates the system state.
[0436] A "generative AI model" is an algorithm and software that uses machine learning techniques to analyze data and convert it into natural language or visual information.
[0437] This invention is a system that enables two-way real-time communication between users in virtual reality (VR) and augmented reality (AR) spaces. The entire system consists of a server, a terminal, and a user.
[0438] The server receives user interaction data transmitted from the VR terminal and analyzes it using a generative AI model. The generative AI model converts the user's actions and location information into natural language and visual information. This conversion process can be displayed, for example, as voice commands or text data.
[0439] The terminal collects sensor information from the headset and controllers worn by the VR user, converts it into a standardized data format, and sends it to the server. This data includes the user's hand movements and field of view, which forms the basis of the VR experience.
[0440] Users (AR users) can perceive changes in the VR space by viewing information displayed using AR devices. The AR device sends feedback to the server through voice commands and gestures, which is then reflected in the VR space. This process allows the AR user's instructions to influence the VR user's view and experience in real time.
[0441] A concrete example is a virtual museum tour in entertainment or education. VR users can move around within the virtual museum and observe the exhibits. AR users can experience the same tour in the real world and provide supplementary information through their AR devices. An example of a prompt would be a scenario where a VR user is exploring underwater ruins, and the AR user can add information such as "there is a shark behind the shipwreck."
[0442] This system allows users to leverage each other's interactions to create new experiential value, learning and having fun together in the space between reality and virtual reality.
[0443] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0444] Step 1:
[0445] The device collects information about the actions and location of VR users in the virtual reality space using a headset and controllers. These actions include pointing at and moving to virtual objects. The collected data is input as sensor information and converted into a standardized data format. This conversion prepares the data for consistent transmission to the server.
[0446] Step 2:
[0447] The server receives standardized data transmitted from the terminal and analyzes it using a generative AI model. The data provided as input is converted into natural language or visual information through the generative AI model. As a concrete example of this analysis, the hand movements of a VR user are translated into the action "the user selected an object." The analyzed information is prepared as output in an appropriate format for AR users.
[0448] Step 3:
[0449] The user (AR user) views the analyzed information sent from the server using an AR device. The AR device receives this information and presents it visually to the user. The AR user provides feedback from their environment using voice commands and gestures. This feedback data is standardized again by the device and converted into input data to be sent to the server.
[0450] Step 4:
[0451] The server receives and analyzes feedback from the AR device. Based on the input feedback, data calculations are performed to update objects and situations in the virtual reality space. For example, if an AR user instructs "move this object," that instruction is reflected in the VR space, and the object actually moves.
[0452] Step 5:
[0453] The user (VR user) experiences the VR space updated on the server and interacts again with the new situation that incorporates the feedback from the AR user. This allows the virtual environment to continuously update itself, maintaining real-time and two-way communication between users.
[0454] (Application Example 1)
[0455] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0456] The challenge lies in achieving more interactive, real-time, two-way communication between users in virtual and extended environments. In particular, there is a need for a system where information within the virtual environment is intuitively understood in the extended environment, and operations in the extended environment are immediately reflected in the virtual environment.
[0457] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0458] In this invention, the server includes means for collecting individual behavioral data in a virtual environment, means for analyzing the collected behavioral data and visualizing it as information in an extended environment, and means for acquiring operation data from users of the extended environment and reflecting the acquired operation data in the virtual environment. This integrates the virtual environment and the extended environment, enabling real-time bidirectional communication between users.
[0459] A "virtual environment" is a synthetic space created using digital technology that is different from reality, and it is a space in which users can have an immersive experience both visually and interactively.
[0460] "Behavioral data" refers to information that describes a series of actions and operations performed by an individual within a virtual and augmented environment, including actions, location, and voice.
[0461] An "augmented environment" is an environment created by overlaying digital information onto information from the real world, allowing users to bring virtual objects and information into their real-world field of vision.
[0462] "Visualizing information" means converting data and interactions into a form that is easily recognizable by users, and expressing them in the form of shapes, graphs, text, and other visual media.
[0463] "Operation data" refers to various types of information that users of the extended environment input via their devices, including gestures, voice input, and selection operations.
[0464] "To be integrated" refers to a state in which different environments or systems cooperate with each other to achieve unified functionality.
[0465] To implement this invention, the user wears smart glasses and browses products and services in a digital space. However, the store clerk wears a virtual reality headset and provides customer service from a separate virtual space. The server has the function of collecting user behavior data from the smart glasses and converting it into a standardized data format. This data is processed on the server and parsed as natural language by a generative AI model (e.g., OpenAI's GPT-4).
[0466] The server reconstructs these analysis results so that the user can visually obtain the information within the augmented environment, and outputs it to the smart glasses. In addition, gestures and voices made by the augmented reality user are sent to the server as operation data, which is analyzed by the generating AI model and immediately reflected in the virtual reality environment.
[0467] As a concrete example, suppose a customer is walking around a virtual store and notices a particular brand of shoes. The server acquires the customer's gaze information, uses a generative AI model to translate detailed information and history of the shoes into natural language, and displays it on the smart glasses. The store clerk can then review this information in VR and offer further advice or suggest other products. An example of a prompt might be, "When the customer looks at a particular pair of shoes, please provide relevant information."
[0468] In this way, users of servers, virtual environments, and extended environments can all achieve real-time, two-way communication, enabling efficient service delivery.
[0469] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0470] Step 1:
[0471] The device (smart glasses) captures the user's movements and gaze through cameras and sensors. This input data serves as basic information for identifying which object the user is focusing on. This information is collected as a digital signal and converted into a structured data format.
[0472] Step 2:
[0473] The device transfers structured behavioral data to the server. The server receives this data and analyzes it using a generative AI model. The analysis expresses information about the products the user is focusing their attention on in natural language. This output includes which items are attracting attention, as well as background information and whether similar products exist.
[0474] Step 3:
[0475] The server sends the analysis results back to the terminal. The terminal then performs data conversion to visualize the information on a digital display. This output is the information the user receives visually through smart glasses. The visualized information includes product details and recommended products.
[0476] Step 4:
[0477] If the user requests further information through gestures or voice, the device sends this interaction to the server as new input data. The server then analyzes this input again and generates new information corresponding to the interaction. This generated information is immediately reflected in the augmented reality space.
[0478] Step 5:
[0479] The server reflects augmented reality changes in the virtual reality environment. Within the virtual reality environment, the store clerk performs further tasks based on the user's requests and the products they are interested in. This output becomes an information environment shared by the user and the store clerk. Using a generative AI model, prompts are used to execute instructions such as, "If the customer looks at a specific pair of shoes, please provide relevant information."
[0480] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0481] This invention is a system that combines an emotion recognition engine to enhance interaction between virtual reality (VR) and augmented reality (AR) environments. By incorporating reactions that correspond to the user's emotional state, this system provides a more immersive experience.
[0482] The user (VR user) uses a VR device to perform activities within a virtual environment. The device collects interaction data from the VR user's actions and analyzes the user's emotions via an emotion engine. Based on this emotion analysis, the virtual reality environment generates feedback and content that corresponds to the user's emotions.
[0483] The user (AR user) uses an AR device to visualize the VR user's interaction information and emotional state transmitted from the server. Based on this information, the AR user provides feedback to the VR environment using gestures and voice commands. The AR device sends this feedback to the server, which generates instructions to update the VR environment.
[0484] The server uses an emotion engine to monitor the emotional states of both users and dynamically control elements in the VR and AR environments. For example, if a VR user is feeling anxious, the server adjusts the VR environment to present relaxing sounds and images.
[0485] For example, when a VR user encounters a scene that evokes awe, the emotion engine detects this reaction, and the server reflects an amplified effect in the VR environment. Simultaneously, the AR user can visualize this emotional change and provide appropriate feedback, thereby building a collaborative relationship with the VR user.
[0486] Thus, the present invention enables sophisticated interactions that respond to the user's emotions, providing an immersive experience that further transcends the boundaries between virtual and reality.
[0487] The following describes the processing flow.
[0488] Step 1:
[0489] The user (VR user) wears a VR headset and interacts within the VR environment. The device acquires interaction data such as the user's movements, gaze, and voice, and transmits it to the emotion engine in real time.
[0490] Step 2:
[0491] The emotion engine analyzes the acquired data to identify the user's emotional state. Based on these results, it generates instructions to adjust the content and effects in the VR environment.
[0492] Step 3:
[0493] The server updates the VR environment based on instructions received from the emotion engine. For example, if the user is feeling anxious, the server inserts calming background sounds and visuals.
[0494] Step 4:
[0495] The user (AR user) visualizes the VR user's emotional state and interaction information through the AR device. The device continuously updates this information in real time.
[0496] Step 5:
[0497] AR users provide appropriate feedback through gestures and voice commands based on the visualized information. This input is immediately sent to the server via the device.
[0498] Step 6:
[0499] The server receives and analyzes feedback from the AR user. Based on this analysis, it generates additional instructions that affect the VR environment and sends them to the device.
[0500] Step 7:
[0501] The terminal reflects new instructions from the server to the VR device, further updating the VR user's virtual environment. This allows the user to have a realistic and emotionally rich experience.
[0502] (Example 2)
[0503] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0504] Modern virtual environments suffer from a lack of feedback that reflects user emotional changes, resulting in a limited sense of immersion in the user experience. Furthermore, it is difficult to smoothly visualize the user's emotional state within the virtual environment in the augmented environment. This, in turn, restricts interaction between the augmented and virtual environments.
[0505] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0506] In this invention, the server includes means for acquiring data on the user's state from a measuring device, means for classifying the acquired data into emotional states using an analysis device, and means for applying changes generated based on the analysis results to the virtual environment. This makes it possible to dynamically adjust the virtual environment according to the user's emotions and provide a deeper sense of immersion.
[0507] A "user" refers to an entity that performs operations or experiences within a virtual or extended environment, and whose actions and emotional states are acquired and analyzed by the system.
[0508] "Status data" refers to information including user movements, actions, and physiological or psychological indicators, and is used to understand the user's current situation and reactions.
[0509] A "measuring device" refers to equipment or sensors that physically or digitally collect data about the user's state, enabling the system to measure the user's behavior and reactions.
[0510] An "analysis device" refers to a program or device that processes acquired data to identify the user's emotional state, and the virtual environment is adjusted based on the results.
[0511] "Emotional state" is a classification of a user's psychological response, and includes emotional categories such as joy, surprise, and relaxation.
[0512] A "virtual environment" refers to an artificially generated digital space in which users immerse themselves and interact with others.
[0513] An "extended environment" refers to a technology that provides users with visual and auditory information by overlaying digital information onto the physical world, enabling collaboration with virtual environments.
[0514] "Dynamic control" refers to real-time changes or adaptations, meaning a process in which system settings and content are adjusted according to the user's emotions and circumstances.
[0515] This invention provides a system that integrates a virtual environment and an extended environment in order to realize advanced interactions based on the user's emotional state.
[0516] The server utilizes measuring devices to acquire data about the user's state, including sensor devices such as VR and AR devices. This data includes physiological or behavioral information such as the user's movements, location, heart rate, and facial expressions. The terminal is responsible for aggregating this data and transmitting it to the server. The server processes the received data using analysis devices and uses a generative AI model to identify the user's emotional state. This model is built, for example, using a machine learning algorithm, and recognizes emotions by comparing them with pre-trained emotional pattern data.
[0517] Based on the results of emotion analysis, the server generates instructions to dynamically adjust the settings of the virtual environment. For example, if the user expresses the emotion of "surprise," the server will enhance the visual effects within the virtual environment and adjust the sound effects accordingly. It can also change the background music to enhance relaxation.
[0518] Meanwhile, users of AR devices receive real-time information on the VR user's emotional state and environmental changes sent from the server. This information is visually presented in the augmented environment, and AR users can interact with the VR environment through voice commands and gestures. The device promptly sends this feedback to the server, ensuring that the environment is updated in response to the user's feedback.
[0519] As a concrete example, by inputting the following prompt into the generating AI model: "Please propose a method to analyze the emotions of a VR user upon entering a new environment in real time and provide optimized sound and visuals based on the results," this system can make the user experience richer and more immersive. In this way, the present invention realizes a new form of interaction between the virtual environment and the real world.
[0520] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0521] Step 1:
[0522] The user puts on a VR device and begins activities within the virtual environment. The terminal detects the user's movements and actions and collects interaction data in real time. Inputs include user movement information, location information, and physiological data obtained from sensors. The terminal sends this as a single dataset to the server. The output is a detailed dataset showing the user's state.
[0523] Step 2:
[0524] The server processes the received interaction data using an analysis device. It uses a generative AI model to analyze the data and identify the user's emotional state. The input is interaction data sent from the VR device. The server feeds this data into the AI model and identifies the emotional category by comparing it with pre-trained emotional patterns. The output is the specific emotional state the user is experiencing (e.g., "joy," "excitement," "relaxation," etc.).
[0525] Step 3:
[0526] The server generates instructions to dynamically adjust the virtual environment settings based on the analyzed emotional state. The input is the results of the emotional analysis. Based on this data, the server issues specific instructions to adjust the lighting, sound, and other visual effects within the virtual environment. The output is the adjusted virtual environment settings.
[0527] Step 4:
[0528] Users of AR devices receive information about the VR user's emotional state and environmental changes sent from the server. Inputs include information about changes in the VR environment and emotional data sent from the server. The device visualizes this information in real time and presents it to the AR user. Outputs are change information that the user can confirm through visual and audio commands.
[0529] Step 5:
[0530] AR users provide feedback using voice commands and gestures based on visualized information. The device sends this feedback information to the server. Inputs include the AR user's gesture data and voice commands. The device quickly transmits this information to the server, providing further information for adjusting the virtual environment. Outputs are the feedback data received by the server.
[0531] In this way, servers, terminals, and users collaborate with each other to achieve sophisticated, emotion-based interactions.
[0532] (Application Example 2)
[0533] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0534] Traditional virtual reality and augmented reality experiences have failed to fully adapt to the user's emotional state, making it difficult to dynamically customize individual experiences. Furthermore, interaction within the virtual environment is limited to sight and sound, hindering user immersion. This resulted in users having to continue the experience while emotionally unresponsive to certain scenes and events, preventing them from achieving deeper emotional impact and immersion.
[0535] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0536] In this invention, the server includes means for collecting user behavior data in a virtual environment, means for analyzing the collected behavior data and converting it into various types of information to be visualized in an extended environment, means for acquiring operation data from users of the extended environment and reflecting the acquired operation data in the virtual environment, and means for analyzing the user's emotional state using an emotion recognition engine and dynamically adjusting the audio and video effects of the virtual environment based on the analysis results. This enables real-time feedback and environmental adjustments in response to the user's emotions towards the content being viewed, allowing for a deeper sense of immersion.
[0537] A "virtual environment" is a digital realm distinct from the real world, constructed using computer technology, within which users can have visual and auditory experiences.
[0538] "User" refers to an individual who experiences and operates the virtual and extended environments through this system.
[0539] "Behavioral data" refers to information about various operations and actions performed by users within a virtual environment, and the system uses this data to analyze user behavior and intentions.
[0540] An "augmented environment" refers to an environment that enhances the user experience by overlaying digital information onto the real world, thereby adding virtual elements to the actual field of view.
[0541] "Visualization" is the process of representing data and information in a way that is easy for users to understand, thereby enabling digital information to be displayed effectively on a screen.
[0542] "Operation data" refers to the records of various operations performed by the user in the extended environment, and this data forms the basis for feedback responses to the virtual environment.
[0543] An "emotion recognition engine" is a software or hardware component that analyzes a user's facial expressions, voice tone, and other behavioral patterns to understand their emotional state.
[0544] "Audio and video effects" refers to sound and visual effects used to enrich the experience and enhance immersion within a virtual environment.
[0545] "Dynamic adjustment" refers to a process in which the behavior and settings of a system are automatically and in real time changed in response to specific user conditions or changes in the environment.
[0546] The system for realizing this application consists of a server, a terminal equipped with a VR / AR device, and an emotion recognition engine. The server collects user behavior data and analyzes the user's emotional state in real time through the emotion recognition engine. Based on this analysis, the server dynamically adjusts the audio and video effects of the virtual environment. The server also receives operation data from the augmented environment user and reflects it as feedback in the virtual environment.
[0547] The device uses a smartphone or head-mounted display (e.g., a typical VR headset) to provide the user with visual and auditory information. This allows the user to immerse themselves in the virtual environment and have a personalized experience with the content they view. The device also features cameras and sensors to detect the user's facial expressions and actions for an emotion recognition engine.
[0548] For example, if a user expresses surprise during a movie viewing, the server can enhance the movie's soundtrack and visual effects accordingly. This allows users to have a more immersive experience. Furthermore, interactive features are enabled that allow users to share the same experience and exchange opinions based on their emotions.
[0549] In relation to the generative AI model, this system analyzes emotions using prompts such as, "If the user is surprised by a specific scene in a movie, dynamically change the effects to match the emotion," and adjusts various elements within the virtual environment accordingly.
[0550] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0551] Step 1:
[0552] The device collects behavioral and emotional data of the user during their VR / AR experience in real time. This process involves capturing facial expressions and movements using cameras and sensors, which are then used as initial data. The input consists of the user's movements and facial expressions, while the output is corresponding behavioral and emotional data.
[0553] Step 2:
[0554] The server receives the collected behavioral and emotional data and analyzes it using an emotion recognition engine. This extracts the user's emotional state (e.g., surprise, joy, fear). The input is the raw data collected by the terminal, and the output is the user's emotional state information obtained through the analysis.
[0555] Step 3:
[0556] The server uses a generative AI model to select appropriate feedback based on the analysis results and prepares corresponding audio and video effects. The prompt message "When the user is surprised in a specific scene of the movie, dynamically change the effects to match the emotion" is used to determine the direction of the analysis. The input is emotional state information, and the output is the selected feedback content.
[0557] Step 4:
[0558] The server sends the prepared feedback to the terminal, instructing it to dynamically adjust the audio and video within the virtual environment. The terminal receives this instruction and adjusts the effects on the content being played in real time. The input is the selected feedback, and the output is the playback of the adjusted content.
[0559] Step 5:
[0560] Users experience newly adjusted audio and video effects in a virtual environment, creating an immersive experience. If the user's emotions change, the process returns to step 1 and is repeated. The input is the adjusted content, and the output is the user's experience.
[0561] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0562] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0563] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[0564] [Fourth Embodiment]
[0565] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0566] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0567] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0568] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0569] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0570] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0571] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0572] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0573] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0574] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0575] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0576] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0577] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0578] This invention provides a system that enables real-time information sharing and bidirectional interaction between a virtual reality (VR) environment and an augmented reality (AR) environment. Specific embodiments of the system are described below.
[0579] The user (VR user) wears a VR headset and controllers and interacts in a virtual reality environment. The VR user's movements and location information are transmitted to the VR terminal via wireless and wired connections, where interaction data is collected. The terminal converts this data into a standardized data format and sends it to the server.
[0580] After receiving this data, the server analyzes it using a generative AI model. This analysis converts user interactions in the VR environment into natural language or visual information, processing it into a format that is easy for AR users to understand.
[0581] The user (AR user) operates the AR device to visualize and confirm information sent from the server. The AR user can provide feedback to the VR environment through their own interface, inputting information for example through gestures or voice input. The AR device then standardizes this input data and sends it back to the server.
[0582] The server receives operation data from the AR user and reflects it in the VR environment. As a result, objects and situations in the virtual reality environment are updated based on the AR user's actions. For example, if the AR user directs an object in the VR environment with a gesture, the object will move in the VR user's field of view accordingly.
[0583] Thus, the system of the present invention makes it possible for VR users and AR users to simultaneously experience the same virtual environment and realize two-way communication in real time. Users can leverage each other's interactions and collaborate to obtain an experience that transcends the boundaries between virtual and reality.
[0584] The following describes the processing flow.
[0585] Step 1:
[0586] The user (VR user) engages in activities within a virtual reality environment using a VR device. The terminal acquires interaction data such as the user's location and gestures from the VR device and temporarily stores this data.
[0587] Step 2:
[0588] The terminal processes the stored interaction data in real time and converts it into a standardized data format. This standardized data is then prepared for transmission to the server.
[0589] Step 3:
[0590] The terminal sends the converted interaction data to the server. The server receives this data and begins the analysis process.
[0591] Step 4:
[0592] The server uses a generated AI model based on the received data to analyze the VR user's interactions. Based on the analysis results, it converts the data into visual or linguistic data.
[0593] Step 5:
[0594] The server sends the converted visual or linguistic data to the device (AR user side).
[0595] Step 6:
[0596] The user (AR user) receives information transmitted from the server using an AR device and confirms the visualized VR interaction. Based on this information, the user provides feedback to the virtual reality environment through gestures and voice.
[0597] Step 7:
[0598] The device (AR user side) acquires user feedback information and standardizes that data for transmission to the server.
[0599] Step 8:
[0600] The server receives feedback data from AR users and performs analysis. Based on the analysis results, it generates data that instructs the server to make necessary changes to the VR environment.
[0601] Step 9:
[0602] The server sends the generated instruction data to the terminal (VR user side) and prepares to update objects and actions in the virtual reality environment.
[0603] Step 10:
[0604] The terminal (VR user side) receives instruction data from the server and updates the VR environment. As a result, the user (VR user) experiences virtual reality that reflects the latest environmental changes.
[0605] (Example 1)
[0606] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0607] There is a need for technologies that enable users to share information bidirectionally and engage in real-time, collaborative interactions in both virtual and augmented reality spaces. However, current technologies face challenges in smoothly visualizing and reflecting information due to delays in information analysis and compatibility issues.
[0608] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0609] In this invention, the server includes a device for collecting user activity information in a virtual reality space, a processing device for analyzing the collected activity information and converting it into various types of information that can be displayed in an augmented reality space, a device for acquiring operation information from augmented reality users and reflecting the acquired operation information in the virtual reality space, and a device for analyzing received data using a generative AI model. This enables users to share information with each other in real time and apply timely and appropriate feedback in their respective real spaces.
[0610] A "virtual reality space" is an artificial three-dimensional environment created by a computer that users can experience through their sight and hearing.
[0611] "User activity information" refers to data about the actions, location, and other interactions that users take in the virtual reality space.
[0612] A "collection device" is a hardware and software system for detecting, storing, and transmitting activity information generated by users.
[0613] An "analytical processing device" is a software and hardware system that interprets collected activity information and processes it according to a specific purpose.
[0614] "Augmented reality" refers to a space that enhances the user's experience by overlaying digital information and images onto the real world's field of view.
[0615] "Various displayable information" refers to various types of information presented to the user visually or audibly based on analyzed activity information.
[0616] "Operation information" refers to data that includes actions, commands, and instructions performed by augmented reality users within the augmented reality space.
[0617] A "reflection device" is a device that applies the operation information transmitted by augmented reality users to the virtual reality space and updates the system state.
[0618] A "generative AI model" is an algorithm and software that uses machine learning techniques to analyze data and convert it into natural language or visual information.
[0619] This invention is a system that enables two-way real-time communication between users in virtual reality (VR) and augmented reality (AR) spaces. The entire system consists of a server, a terminal, and a user.
[0620] The server receives user interaction data transmitted from the VR terminal and analyzes it using a generative AI model. The generative AI model converts the user's actions and location information into natural language and visual information. This conversion process can be displayed, for example, as voice commands or text data.
[0621] The terminal collects sensor information from the headset and controllers worn by the VR user, converts it into a standardized data format, and sends it to the server. This data includes the user's hand movements and field of view, which forms the basis of the VR experience.
[0622] Users (AR users) can perceive changes in the VR space by viewing information displayed using AR devices. The AR device sends feedback to the server through voice commands and gestures, which is then reflected in the VR space. This process allows the AR user's instructions to influence the VR user's view and experience in real time.
[0623] A concrete example is a virtual museum tour in entertainment or education. VR users can move around within the virtual museum and observe the exhibits. AR users can experience the same tour in the real world and provide supplementary information through their AR devices. An example of a prompt would be a scenario where a VR user is exploring underwater ruins, and the AR user can add information such as "there is a shark behind the shipwreck."
[0624] This system allows users to leverage each other's interactions to create new experiential value, learning and having fun together in the space between reality and virtual reality.
[0625] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0626] Step 1:
[0627] The device collects information about the actions and location of VR users in the virtual reality space using a headset and controllers. These actions include pointing at and moving to virtual objects. The collected data is input as sensor information and converted into a standardized data format. This conversion prepares the data for consistent transmission to the server.
[0628] Step 2:
[0629] The server receives standardized data transmitted from the terminal and analyzes it using a generative AI model. The data provided as input is converted into natural language or visual information through the generative AI model. As a concrete example of this analysis, the hand movements of a VR user are translated into the action "the user selected an object." The analyzed information is prepared as output in an appropriate format for AR users.
[0630] Step 3:
[0631] The user (AR user) views the analyzed information sent from the server using an AR device. The AR device receives this information and presents it visually to the user. The AR user provides feedback from their environment using voice commands and gestures. This feedback data is standardized again by the device and converted into input data to be sent to the server.
[0632] Step 4:
[0633] The server receives and analyzes feedback from the AR device. Based on the input feedback, data calculations are performed to update objects and situations in the virtual reality space. For example, if an AR user instructs "move this object," that instruction is reflected in the VR space, and the object actually moves.
[0634] Step 5:
[0635] The user (VR user) experiences the VR space updated on the server and interacts again with the new situation that incorporates the feedback from the AR user. This allows the virtual environment to continuously update itself, maintaining real-time and two-way communication between users.
[0636] (Application Example 1)
[0637] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0638] The challenge lies in achieving more interactive, real-time, two-way communication between users in virtual and extended environments. In particular, there is a need for a system where information within the virtual environment is intuitively understood in the extended environment, and operations in the extended environment are immediately reflected in the virtual environment.
[0639] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0640] In this invention, the server includes means for collecting individual behavioral data in a virtual environment, means for analyzing the collected behavioral data and visualizing it as information in an extended environment, and means for acquiring operation data from users of the extended environment and reflecting the acquired operation data in the virtual environment. This integrates the virtual environment and the extended environment, enabling real-time bidirectional communication between users.
[0641] A "virtual environment" is a synthetic space created using digital technology that is different from reality, and it is a space in which users can have an immersive experience both visually and interactively.
[0642] "Behavioral data" refers to information that describes a series of actions and operations performed by an individual within a virtual and augmented environment, including actions, location, and voice.
[0643] An "augmented environment" is an environment created by overlaying digital information onto information from the real world, allowing users to bring virtual objects and information into their real-world field of vision.
[0644] "Visualizing information" means converting data and interactions into a form that is easily recognizable by users, and expressing them in the form of shapes, graphs, text, and other visual media.
[0645] "Operation data" refers to various types of information that users of the extended environment input via their devices, including gestures, voice input, and selection operations.
[0646] "To be integrated" refers to a state in which different environments or systems cooperate with each other to achieve unified functionality.
[0647] To implement this invention, the user wears smart glasses and browses products and services in a digital space. However, the store clerk wears a virtual reality headset and provides customer service from a separate virtual space. The server has the function of collecting user behavior data from the smart glasses and converting it into a standardized data format. This data is processed on the server and parsed as natural language by a generative AI model (e.g., OpenAI's GPT-4).
[0648] The server reconstructs these analysis results so that the user can visually obtain the information within the augmented environment, and outputs it to the smart glasses. In addition, gestures and voices made by the augmented reality user are sent to the server as operation data, which is analyzed by the generating AI model and immediately reflected in the virtual reality environment.
[0649] As a concrete example, suppose a customer is walking around a virtual store and notices a particular brand of shoes. The server acquires the customer's gaze information, uses a generative AI model to translate detailed information and history of the shoes into natural language, and displays it on the smart glasses. The store clerk can then review this information in VR and offer further advice or suggest other products. An example of a prompt might be, "When the customer looks at a particular pair of shoes, please provide relevant information."
[0650] In this way, users of servers, virtual environments, and extended environments can all achieve real-time, two-way communication, enabling efficient service delivery.
[0651] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0652] Step 1:
[0653] The device (smart glasses) captures the user's movements and gaze through cameras and sensors. This input data serves as basic information for identifying which object the user is focusing on. This information is collected as a digital signal and converted into a structured data format.
[0654] Step 2:
[0655] The device transfers structured behavioral data to the server. The server receives this data and analyzes it using a generative AI model. The analysis expresses information about the products the user is focusing their attention on in natural language. This output includes which items are attracting attention, as well as background information and whether similar products exist.
[0656] Step 3:
[0657] The server sends the analysis results back to the terminal. The terminal then performs data conversion to visualize the information on a digital display. This output is the information the user receives visually through smart glasses. The visualized information includes product details and recommended products.
[0658] Step 4:
[0659] If the user requests further information through gestures or voice, the device sends this interaction to the server as new input data. The server then analyzes this input again and generates new information corresponding to the interaction. This generated information is immediately reflected in the augmented reality space.
[0660] Step 5:
[0661] The server reflects augmented reality changes in the virtual reality environment. Within the virtual reality environment, the store clerk performs further tasks based on the user's requests and the products they are interested in. This output becomes an information environment shared by the user and the store clerk. Using a generative AI model, prompts are used to execute instructions such as, "If the customer looks at a specific pair of shoes, please provide relevant information."
[0662] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0663] This invention is a system that combines an emotion recognition engine to enhance interaction between virtual reality (VR) and augmented reality (AR) environments. By incorporating reactions that correspond to the user's emotional state, this system provides a more immersive experience.
[0664] The user (VR user) uses a VR device to perform activities within a virtual environment. The device collects interaction data from the VR user's actions and analyzes the user's emotions via an emotion engine. Based on this emotion analysis, the virtual reality environment generates feedback and content that corresponds to the user's emotions.
[0665] The user (AR user) uses an AR device to visualize the VR user's interaction information and emotional state transmitted from the server. Based on this information, the AR user provides feedback to the VR environment using gestures and voice commands. The AR device sends this feedback to the server, which generates instructions to update the VR environment.
[0666] The server uses an emotion engine to monitor the emotional states of both users and dynamically control elements in the VR and AR environments. For example, if a VR user is feeling anxious, the server adjusts the VR environment to present relaxing sounds and images.
[0667] For example, when a VR user encounters a scene that evokes awe, the emotion engine detects this reaction, and the server reflects an amplified effect in the VR environment. Simultaneously, the AR user can visualize this emotional change and provide appropriate feedback, thereby building a collaborative relationship with the VR user.
[0668] Thus, the present invention enables sophisticated interactions that respond to the user's emotions, providing an immersive experience that further transcends the boundaries between virtual and reality.
[0669] The following describes the processing flow.
[0670] Step 1:
[0671] The user (VR user) wears a VR headset and interacts within the VR environment. The device acquires interaction data such as the user's movements, gaze, and voice, and transmits it to the emotion engine in real time.
[0672] Step 2:
[0673] The emotion engine analyzes the acquired data to identify the user's emotional state. Based on these results, it generates instructions to adjust the content and effects in the VR environment.
[0674] Step 3:
[0675] The server updates the VR environment based on instructions received from the emotion engine. For example, if the user is feeling anxious, the server inserts calming background sounds and visuals.
[0676] Step 4:
[0677] The user (AR user) visualizes the VR user's emotional state and interaction information through the AR device. The device continuously updates this information in real time.
[0678] Step 5:
[0679] AR users provide appropriate feedback through gestures and voice commands based on the visualized information. This input is immediately sent to the server via the device.
[0680] Step 6:
[0681] The server receives and analyzes feedback from the AR user. Based on this analysis, it generates additional instructions that affect the VR environment and sends them to the device.
[0682] Step 7:
[0683] The terminal reflects new instructions from the server to the VR device, further updating the VR user's virtual environment. This allows the user to have a realistic and emotionally rich experience.
[0684] (Example 2)
[0685] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0686] Modern virtual environments suffer from a lack of feedback that reflects user emotional changes, resulting in a limited sense of immersion in the user experience. Furthermore, it is difficult to smoothly visualize the user's emotional state within the virtual environment in the augmented environment. This, in turn, restricts interaction between the augmented and virtual environments.
[0687] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0688] In this invention, the server includes means for acquiring data on the user's state from a measuring device, means for classifying the acquired data into emotional states using an analysis device, and means for applying changes generated based on the analysis results to the virtual environment. This makes it possible to dynamically adjust the virtual environment according to the user's emotions and provide a deeper sense of immersion.
[0689] A "user" refers to an entity that performs operations or experiences within a virtual or extended environment, and whose actions and emotional states are acquired and analyzed by the system.
[0690] "Status data" refers to information including user movements, actions, and physiological or psychological indicators, and is used to understand the user's current situation and reactions.
[0691] A "measuring device" refers to equipment or sensors that physically or digitally collect data about the user's state, enabling the system to measure the user's behavior and reactions.
[0692] An "analysis device" refers to a program or device that processes acquired data to identify the user's emotional state, and the virtual environment is adjusted based on the results.
[0693] "Emotional state" is a classification of a user's psychological response, and includes emotional categories such as joy, surprise, and relaxation.
[0694] A "virtual environment" refers to an artificially generated digital space in which users immerse themselves and interact with others.
[0695] An "extended environment" refers to a technology that provides users with visual and auditory information by overlaying digital information onto the physical world, enabling collaboration with virtual environments.
[0696] "Dynamic control" refers to real-time changes or adaptations, meaning a process in which system settings and content are adjusted according to the user's emotions and circumstances.
[0697] This invention provides a system that integrates a virtual environment and an extended environment in order to realize advanced interactions based on the user's emotional state.
[0698] The server utilizes measuring devices to acquire data about the user's state, including sensor devices such as VR and AR devices. This data includes physiological or behavioral information such as the user's movements, location, heart rate, and facial expressions. The terminal is responsible for aggregating this data and transmitting it to the server. The server processes the received data using analysis devices and uses a generative AI model to identify the user's emotional state. This model is built, for example, using a machine learning algorithm, and recognizes emotions by comparing them with pre-trained emotional pattern data.
[0699] Based on the results of emotion analysis, the server generates instructions to dynamically adjust the settings of the virtual environment. For example, if the user expresses the emotion of "surprise," the server will enhance the visual effects within the virtual environment and adjust the sound effects accordingly. It can also change the background music to enhance relaxation.
[0700] Meanwhile, users of AR devices receive real-time information on the VR user's emotional state and environmental changes sent from the server. This information is visually presented in the augmented environment, and AR users can interact with the VR environment through voice commands and gestures. The device promptly sends this feedback to the server, ensuring that the environment is updated in response to the user's feedback.
[0701] As a concrete example, by inputting the following prompt into the generating AI model: "Please propose a method to analyze the emotions of a VR user upon entering a new environment in real time and provide optimized sound and visuals based on the results," this system can make the user experience richer and more immersive. In this way, the present invention realizes a new form of interaction between the virtual environment and the real world.
[0702] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0703] Step 1:
[0704] The user puts on a VR device and begins activities within the virtual environment. The terminal detects the user's movements and actions and collects interaction data in real time. Inputs include user movement information, location information, and physiological data obtained from sensors. The terminal sends this as a single dataset to the server. The output is a detailed dataset showing the user's state.
[0705] Step 2:
[0706] The server processes the received interaction data using an analysis device. It uses a generative AI model to analyze the data and identify the user's emotional state. The input is interaction data sent from the VR device. The server feeds this data into the AI model and identifies the emotional category by comparing it with pre-trained emotional patterns. The output is the specific emotional state the user is experiencing (e.g., "joy," "excitement," "relaxation," etc.).
[0707] Step 3:
[0708] The server generates instructions to dynamically adjust the virtual environment settings based on the analyzed emotional state. The input is the results of the emotional analysis. Based on this data, the server issues specific instructions to adjust the lighting, sound, and other visual effects within the virtual environment. The output is the adjusted virtual environment settings.
[0709] Step 4:
[0710] Users of AR devices receive information about the VR user's emotional state and environmental changes sent from the server. Inputs include information about changes in the VR environment and emotional data sent from the server. The device visualizes this information in real time and presents it to the AR user. Outputs are change information that the user can confirm through visual and audio commands.
[0711] Step 5:
[0712] AR users provide feedback using voice commands and gestures based on visualized information. The device sends this feedback information to the server. Inputs include the AR user's gesture data and voice commands. The device quickly transmits this information to the server, providing further information for adjusting the virtual environment. Outputs are the feedback data received by the server.
[0713] In this way, servers, terminals, and users collaborate with each other to achieve sophisticated, emotion-based interactions.
[0714] (Application Example 2)
[0715] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0716] Traditional virtual reality and augmented reality experiences have failed to fully adapt to the user's emotional state, making it difficult to dynamically customize individual experiences. Furthermore, interaction within the virtual environment is limited to sight and sound, hindering user immersion. This resulted in users having to continue the experience while emotionally unresponsive to certain scenes and events, preventing them from achieving deeper emotional impact and immersion.
[0717] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0718] In this invention, the server includes means for collecting user behavior data in a virtual environment, means for analyzing the collected behavior data and converting it into various types of information to be visualized in an extended environment, means for acquiring operation data from users of the extended environment and reflecting the acquired operation data in the virtual environment, and means for analyzing the user's emotional state using an emotion recognition engine and dynamically adjusting the audio and video effects of the virtual environment based on the analysis results. This enables real-time feedback and environmental adjustments in response to the user's emotions towards the content being viewed, allowing for a deeper sense of immersion.
[0719] A "virtual environment" is a digital realm distinct from the real world, constructed using computer technology, within which users can have visual and auditory experiences.
[0720] "User" refers to an individual who experiences and operates the virtual and extended environments through this system.
[0721] "Behavioral data" refers to information about various operations and actions performed by users within a virtual environment, and the system uses this data to analyze user behavior and intentions.
[0722] An "augmented environment" refers to an environment that enhances the user experience by overlaying digital information onto the real world, thereby adding virtual elements to the actual field of view.
[0723] "Visualization" is the process of representing data and information in a way that is easy for users to understand, thereby enabling digital information to be displayed effectively on a screen.
[0724] "Operation data" refers to the records of various operations performed by the user in the extended environment, and this data forms the basis for feedback responses to the virtual environment.
[0725] An "emotion recognition engine" is a software or hardware component that analyzes a user's facial expressions, voice tone, and other behavioral patterns to understand their emotional state.
[0726] "Audio and video effects" refers to sound and visual effects used to enrich the experience and enhance immersion within a virtual environment.
[0727] "Dynamic adjustment" refers to a process in which the behavior and settings of a system are automatically and in real time changed in response to specific user conditions or changes in the environment.
[0728] The system for realizing this application consists of a server, a terminal equipped with a VR / AR device, and an emotion recognition engine. The server collects user behavior data and analyzes the user's emotional state in real time through the emotion recognition engine. Based on this analysis, the server dynamically adjusts the audio and video effects of the virtual environment. The server also receives operation data from the augmented environment user and reflects it as feedback in the virtual environment.
[0729] The device uses a smartphone or head-mounted display (e.g., a typical VR headset) to provide the user with visual and auditory information. This allows the user to immerse themselves in the virtual environment and have a personalized experience with the content they view. The device also features cameras and sensors to detect the user's facial expressions and actions for an emotion recognition engine.
[0730] For example, if a user expresses surprise during a movie viewing, the server can enhance the movie's soundtrack and visual effects accordingly. This allows users to have a more immersive experience. Furthermore, interactive features are enabled that allow users to share the same experience and exchange opinions based on their emotions.
[0731] In relation to the generative AI model, this system analyzes emotions using prompts such as, "If the user is surprised by a specific scene in a movie, dynamically change the effects to match the emotion," and adjusts various elements within the virtual environment accordingly.
[0732] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0733] Step 1:
[0734] The device collects behavioral and emotional data of the user during their VR / AR experience in real time. This process involves capturing facial expressions and movements using cameras and sensors, which are then used as initial data. The input consists of the user's movements and facial expressions, while the output is corresponding behavioral and emotional data.
[0735] Step 2:
[0736] The server receives the collected behavioral and emotional data and analyzes it using an emotion recognition engine. This extracts the user's emotional state (e.g., surprise, joy, fear). The input is the raw data collected by the terminal, and the output is the user's emotional state information obtained through the analysis.
[0737] Step 3:
[0738] The server uses a generative AI model to select appropriate feedback based on the analysis results and prepares corresponding audio and video effects. The prompt message "When the user is surprised in a specific scene of the movie, dynamically change the effects to match the emotion" is used to determine the direction of the analysis. The input is emotional state information, and the output is the selected feedback content.
[0739] Step 4:
[0740] The server sends the prepared feedback to the terminal, instructing it to dynamically adjust the audio and video within the virtual environment. The terminal receives this instruction and adjusts the effects on the content being played in real time. The input is the selected feedback, and the output is the playback of the adjusted content.
[0741] Step 5:
[0742] Users experience newly adjusted audio and video effects in a virtual environment, creating an immersive experience. If the user's emotions change, the process returns to step 1 and is repeated. The input is the adjusted content, and the output is the user's experience.
[0743] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0744] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0745] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[0746] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0747] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[0748] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[0749] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[0750] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[0751] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[0752] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[0753] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[0754] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[0755] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0756] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0757] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[0758] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[0759] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[0760] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[0761] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0762] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0763] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted as being incorporated by reference.
[0764] The following is further disclosed regarding the embodiments described above.
[0765] (Claim 1)
[0766] A means for collecting user interaction data in a virtual reality environment,
[0767] A means of analyzing the collected interaction data and converting it into various types of information that can be visualized in an augmented reality environment,
[0768] A means for acquiring operation data from augmented reality users and reflecting the acquired operation data in the virtual reality environment,
[0769] A system that includes this.
[0770] (Claim 2)
[0771] The system according to claim 1, characterized by comprising means for analyzing the operation data of an augmented reality user and executing a process to change the behavior of an object in a virtual reality environment.
[0772] (Claim 3)
[0773] The system according to claim 1, characterized in that virtual reality user environment update information is delivered to augmented reality users in real time.
[0774] "Example 1"
[0775] (Claim 1)
[0776] A device for collecting user activity information in a virtual reality space,
[0777] A processing device that analyzes collected activity information and converts it into various types of information that can be displayed in an augmented reality space,
[0778] A device that acquires operation information from augmented reality users and reflects the acquired operation information in the virtual reality space,
[0779] A device that analyzes received data using a generative AI model,
[0780] ...
[0781] A system that includes this.
[0782] (Claim 2)
[0783] The system according to claim 1, characterized by comprising a device for analyzing the operation information of an augmented reality user and executing a process to change the movement of an object in a virtual reality space.
[0784] (Claim 3)
[0785] The system according to claim 1, characterized in that spatial update information of a virtual reality user is immediately transmitted to an augmented reality user.
[0786] "Application Example 1"
[0787] (Claim 1)
[0788] A means of collecting individual behavioral data in a virtual environment,
[0789] A means of analyzing collected behavioral data and visualizing it as information within an augmented environment,
[0790] A means of acquiring operation data from users of the extended environment and reflecting the acquired operation data in the virtual environment,
[0791] A means of providing information about objects in a virtual environment based on user instructions within the extended environment,
[0792] A system that includes this.
[0793] (Claim 2)
[0794] The system according to claim 1, characterized by comprising means for analyzing user operation data of an extended environment and executing processing to change the behavior of components in the virtual environment.
[0795] (Claim 3)
[0796] The system according to claim 1, characterized in that environment update information of virtual environment users is delivered to extended environment users in real time.
[0797] "Example 2 of combining an emotion engine"
[0798] (Claim 1)
[0799] A means for acquiring data about the user's state from a measuring device,
[0800] A means of classifying acquired data into emotional states using an analysis device,
[0801] A means of applying the changes generated based on the analysis results to the virtual environment,
[0802] A means of presenting visualized information in an extended environment,
[0803] A means of obtaining information from extended users and transmitting it to the virtual environment,
[0804] A system that includes this.
[0805] (Claim 2)
[0806] The system according to claim 1, characterized by comprising means for dynamically controlling elements within a virtual environment based on analyzed data.
[0807] (Claim 3)
[0808] The system according to claim 1, characterized in that the collected data and analysis results are provided to extended users in real time.
[0809] "Application example 2 when combining with an emotional engine"
[0810] (Claim 1)
[0811] A means of collecting user behavior data in a virtual environment,
[0812] A means of analyzing collected behavioral data and converting it into diverse information that can be visualized in an augmented environment,
[0813] A means of acquiring operation data from users of the extended environment and reflecting the acquired operation data in the virtual environment,
[0814] A means for analyzing the user's emotional state using an emotion recognition engine and dynamically adjusting the audio and video effects of the virtual environment based on the analysis results,
[0815] A system that includes this.
[0816] (Claim 2)
[0817] The system according to claim 1, characterized by comprising means for analyzing operation data of users of the extended environment and executing processing to change the behavior of an object in the virtual environment.
[0818] (Claim 3)
[0819] The system according to claim 1, characterized in that environment update information of virtual environment users is delivered to extended environment users in real time. [Explanation of symbols]
[0820] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. A means for collecting user interaction data in a virtual reality environment, A means of analyzing the collected interaction data and converting it into various types of information that can be visualized in an augmented reality environment, A means for acquiring operation data from augmented reality users and reflecting the acquired operation data in the virtual reality environment, A system that includes this.
2. The system according to claim 1, characterized by comprising means for analyzing the operation data of an augmented reality user and executing a process to change the behavior of an object in a virtual reality environment.
3. The system according to claim 1, characterized in that virtual reality user environment update information is delivered to augmented reality users in real time.