System
The 3D concierge system integrates VR and generative AI to provide immediate, realistic responses to user queries, addressing the inefficiency of traditional search methods and lack of character interaction.
Patent Information
- Application Number
- JP2024121569
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-26
- Publication Date
- 2026-02-05
AI Technical Summary
People face challenges in obtaining quick and intuitive answers to trivial questions, such as cooking instructions or weather forecasts, often relying on inefficient smartphone searches, and there are no systems to interact with favorite characters from anime or games.
A 3D concierge system combining VR technology and generative AI, allowing users to input questions via voice or text, process them through natural language, and receive immediate, realistic 3D responses from a virtual concierge.
Enables efficient and intuitive interaction with a virtual concierge that provides immediate, detailed answers in a realistic virtual environment, enhancing user experience.
Smart Images

Figure 2026019821000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] Many people today face the challenge of not having someone nearby who can quickly answer their small questions. For example, it can be difficult to get instant answers to trivial questions, such as cooking instructions or weather forecast information. In such cases, people often use smartphones or computers to search for information, but this is inefficient and unintuitive. Furthermore, anime and game fans often want answers from their favorite characters, but there are no systems in place to accommodate this. To solve this problem, the introduction of VR technology, generative AI, and even 3D concierges is required. [Means for solving the problem]
[0005] The present invention solves the above-mentioned problems by providing a 3D concierge system that combines VR technology and generative AI. This system includes a means for a user to wear a virtual reality device, a means for collecting user-input voice or text, a means for analyzing the collected user question data and performing natural language processing, a means for inputting the analyzed data into a generative AI model to generate an appropriate response, and a means for displaying the generated response data in the form of a 3D concierge in a virtual reality space. Furthermore, the 3D concierge responds to the user with natural movements and voice, giving the user an experience that feels as if they are conversing with a real person. Furthermore, by converting the user's voice input into text in real time and sending it to a generative AI model, appropriate answers can be obtained immediately. In this way, a system can be provided that efficiently and intuitively responds to small questions and inquiries in everyday life.
[0006] A "virtual reality device" is a device that allows a user to experience a virtual reality space through their senses, such as sight and hearing.
[0007] A "user" is a person who wears a virtual reality device and enters a question into the concierge system.
[0008] "Voice or text" refers to the input format of the question that the user asks the concierge system, and is either voice input or text input.
[0009] "Means of collection" refers to the technical means for obtaining voice or text data input by the user.
[0010] "Means of analysis" means the technical means for understanding and analyzing collected voice or text data through natural language processing.
[0011] "Natural language processing" is a technology that allows computers to understand and analyze human language.
[0012] A "generative artificial intelligence model" is an artificial intelligence program that generates appropriate responses based on input data.
[0013] "Means for generating a response" refers to the technical means of creating an appropriate response using a generative artificial intelligence model based on the analyzed data.
[0014] A "virtual reality space" is a three-dimensional virtual environment that a user experiences visually and aurally through a virtual reality device.
[0015] A "concierge" is a virtual person or character that responds to the user's questions within the virtual reality space.
[0016] "Means for three-dimensional display" means technical means for visually displaying the generated response data as a three-dimensional object in a virtual reality space.
[0017] "Natural movements and voice" means that the concierge in the virtual reality space moves and speaks like a real person.
[0018] "Means for converting text in real time" refers to technical means for instantly converting a user's voice input into character data. [Brief explanation of the drawings]
[0019] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5]FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0020] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0021] First, the terms used in the following description will be explained.
[0022] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0023] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0024] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0025] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0026] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0027] [First embodiment]
[0028] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0029] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0030] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0031] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0032] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0033] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0034] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0035] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0036] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0037] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0038] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0039] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0040] The present invention provides a 3D concierge system that combines VR technology and generative AI. In this system, a user wears a virtual reality device, asks questions by voice or text, and an appropriate response is displayed in 3D. Specific embodiments for implementing the present invention are described below.
[0041] Program processing
[0042] The system consists of three main components: users, terminals, and servers, each of which plays a specific role and works in tandem.
[0043] User operations
[0044] 1. Booting the system
[0045] The user puts on the virtual reality device and launches a dedicated concierge application, at which point a concierge appears in the virtual space.
[0046] 2. Enter your question
[0047] The user asks a question by voice or text, for example, while cooking, "Tell me how to make tomato sauce."
[0048] Questions are entered into the system through the microphone or virtual keyboard of the virtual reality device.
[0049] Terminal handling
[0050] 3. Speech Recognition and Text Conversion
[0051] The user's voice input is recorded in real time by the terminal and converted into text data by voice recognition software.
[0052] The converted text data is packed into packets as question data and sent to the server.
[0053] 4. Receiving response data
[0054] Receives response data sent from the server, checks and converts the format.
[0055] The concierge then formats the responses appropriately so that they can be given to the user using natural movements and voice.
[0056] 5. Rendering 3D Objects
[0057] The device renders the received response data as a three-dimensional object.
[0058] A concierge in the virtual space presents the generated answer to the user visually and audibly.
[0059] Server Processing
[0060] 6. Receiving and analyzing query data
[0061] The server receives the query data sent from the terminal.
[0062] The received data is pre-processed for natural language processing, including tokenization and normalization.
[0063] 7. Answer generation using generative AI
[0064] The natural language processed data is input into a generative artificial intelligence model, which generates an appropriate answer.
[0065] For example, if you input "How to make tomato sauce," the generative AI will generate the answer, "Here's how to make tomato sauce: First, chop the tomatoes. Next, heat the olive oil in a frying pan, add the garlic and fry until fragrant. Then, adjust the flavor with the chopped tomatoes, salt and pepper, and simmer."
[0066] 8. Sending response data
[0067] The generated response data is appropriately formatted and transmitted to the terminal.
[0068] Specific examples
[0069] For example, the specific operations and processing flow when a user asks "Please tell me how to make curry" while cooking will be shown.
[0070] User
[0071] Put on the virtual reality device and say, "Tell me how to make curry."
[0072] Terminal
[0073] The user's speech is recorded and converted into text data using voice recognition software, such as "Please tell me how to make curry."
[0074] The converted text data is sent to the server.
[0075] server
[0076] It receives text data such as "Please tell me how to make curry" and performs natural language processing.
[0077] Using generative AI, the answer is generated: "Here's how to make curry: First, fry the onions. Next, add the meat, then the vegetables and simmer. Finally, add the curry roux and it's done."
[0078] The generated response data is transmitted to the terminal.
[0079] Terminal
[0080] Based on the received response data, it is rendered as a 3D object.
[0081] The concierge in the virtual space responds to the user using natural movements and voice, saying, "Here's how to make curry: First, fry the onions. Next, add the meat, then the vegetables and simmer. Finally, add the curry roux and it's done."
[0082] In this way, users can have a virtual reality experience that feels like they are having a conversation with a real person. The system responds immediately to detailed questions, enabling intuitive and efficient information acquisition.
[0083] The processing flow will be explained below.
[0084] Step 1:
[0085] The user puts on the virtual reality device and presses the start button on the device to launch the dedicated concierge application, which causes a concierge to appear in the virtual space.
[0086] Step 2:
[0087] The user enters a question. The user speaks into the VR device's microphone, such as "What's the weather like in Tokyo today?", or enters the question as text using the virtual keyboard.
[0088] Step 3:
[0089] The device collects voice data. The user's speech is recorded by the VR device's microphone and converted into text data in real time.
[0090] Step 4:
[0091] The device sends the text data to the server. The converted text data is packed into packets and sent to the server via the Internet.
[0092] Step 5:
[0093] The server receives the text data. The server receives and checks the data sent from the terminal.
[0094] Step 6:
[0095] The server analyzes the text data received, performs preprocessing for natural language processing, and tokenizes and normalizes the text data.
[0096] Step 7:
[0097] The server inputs the data into a generative AI model, which then inputs the analyzed text data into the model to generate an appropriate response.
[0098] Step 8:
[0099] The server formats the generated response data, and validates the generated response data and converts it into a format for display within the virtual reality space.
[0100] Step 9:
[0101] The server sends the formatted response data to the terminal, which then packages the response data into packets and sends them over the Internet to the terminal.
[0102] Step 10:
[0103] The terminal receives the response data. The terminal receives the response data sent from the server and checks its format.
[0104] Step 11:
[0105] The device renders the response data as a 3D object, and processes the received data so that the 3D concierge can respond to the user in a virtual space using natural movements and voice.
[0106] Step 12:
[0107] The user confirms the answer from the Moving Concierge. The concierge in the virtual reality space responds to the user, saying, "The weather in Tokyo is sunny today." The user confirms this visually and audibly.
[0108] In this way, each step works in cooperation to realize a system that allows a user to ask a question in a virtual reality space and receive an appropriate response immediately.
[0109] Example 1
[0110] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0111] In conventional virtual reality systems, it has been difficult to obtain natural responses in real time when a user asks a question. In particular, the processing required to provide detailed and appropriate responses to a user's question is complex, resulting in a non-intuitive user experience. Furthermore, conventional systems have a problem in that the quality of the visual and audio representation of the responses is low, resulting in a lack of realism.
[0112] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0113] In this invention, the server includes a means for a user to wear a virtual reality device, a means for collecting voice or text input by the user, a means for analyzing the collected user question data, a means for natural language processing the analyzed data, a means for inputting the processed natural language data into an artificial intelligence model for generating an appropriate response, a means for three-dimensionally displaying the generated response data in the form of an agent in a virtual reality space, and a means for presenting the generated response data in visual and audio formats, thereby enabling the user to have a more realistic interactive experience.
[0114] "User" refers to a person who uses the system.
[0115] "Virtual reality device" refers to equipment used by a user to immerse themselves in a virtual reality environment, including, but not limited to, a headset and a handheld controller.
[0116] "Voice or text collection means" refers to devices or software that capture user-input voice or text data. Voice input is accomplished via a microphone, and text input is accomplished via a virtual keyboard or voice recognition software.
[0117] "Question data" refers to data including the question entered by the user.
[0118] "Analysis tools" refers to software and algorithms used to understand the collected question data and perform processes such as parsing, tokenization, and normalization.
[0119] "Natural language processing" refers to the technology that enables a machine to understand and appropriately process input human language. Specifically, it includes tokenization, grammatical analysis, and semantic analysis.
[0120] "Generative artificial intelligence model" refers to a machine learning algorithm or artificial intelligence technique that generates an appropriate response based on natural language processed data. Specific examples include generative AI models.
[0121] "Response data" refers to data containing answers generated by an artificial intelligence model.
[0122] "Means for three-dimensional display" refers to devices and software for displaying the generated response data in three-dimensional space, specifically including a 3D engine for visually displaying the response to an agent (concierge) in a virtual reality space.
[0123] "Means for visual and audio presentation" refers to devices and software for presenting the generated response data to the user in visual animation and audio, including, for example, speech synthesis software that converts text into speech and animation software that controls the behavior of an agent.
[0124] This invention relates to a 3D concierge system that combines virtual reality (VR) technology and generative artificial intelligence (AI). Specifically, the system provides a user with a virtual reality device, asks questions by voice or text, and displays appropriate responses in 3D. The system consists of three main components: a user, a terminal, and a server.
[0125] System configuration
[0126] User operations
[0127] The user wears a virtual reality device (such as Oculus Rift or HTC Vive). Through this device, the user can launch a dedicated concierge application and ask questions to a three-dimensional agent that appears in the virtual space. Questions can be asked by voice input or text input. Voice input uses a microphone, and text input uses a virtual keyboard.
[0128] Terminal handling
[0129] Speech recognition and text conversion
[0130] The device records the user's voice input in real time and converts it into text using speech recognition software such as the Google Speech-to-Text API or Microsoft Azure Speech Service. For example, if a user says, "Tell me how to make tomato sauce," this will be converted into text.
[0131] Sending text data
[0132] The converted text data is packaged into packets as question data and sent to the server via the HTTPS protocol.
[0133] Receiving and processing response data
[0134] Receives the response data sent from the server, checks and converts the format. Specifically, it converts text data into audio data and prepares it into data for setting the motion of a 3D agent.
[0135] Rendering 3D objects
[0136] The device uses a 3D engine such as Unity or Unreal Engine to display the virtual space, and the 3D agent uses natural movements (such as hand gestures and facial expressions) to present the generated responses visually and audibly.
[0137] Server Processing
[0138] Receiving and analyzing question data
[0139] The server receives the query data sent from the terminal. This data is received through a web server (such as Apache or Nginx). The received data is analyzed by a server-side application written in Python or Java. Natural language processing such as tokenization and normalization is performed in the initial analysis stage.
[0140] Answer generation using generative AI
[0141] The analyzed data is input into a generative AI model (e.g., OpenAI GPT-3). The generative AI model generates an appropriate response based on a prompt (e.g., "Please tell me how to make tomato sauce.") For example, it might generate an answer like, "Here's how to make tomato sauce: First, chop the tomatoes. Next, heat olive oil in a frying pan, add the garlic, and fry until fragrant. Next, add the chopped tomatoes, salt, and pepper to taste, and simmer."
[0142] Sending response data
[0143] The generated response data is formatted in a format such as JSON and sent to the terminal via the HTTPS protocol.
[0144] Specific examples
[0145] For example, the specific operations and processing flow when a user asks "Please tell me how to make curry" while cooking will be described.
[0146] 1. Users
[0147] Put on the VR device and say, "Please tell me how to make curry."
[0148] 2. Terminal
[0149] The user's speech is recorded and converted into text data using voice recognition software, such as "Please tell me how to make curry."
[0150] The converted text data is sent to the server.
[0151] 3. Server
[0152] It receives text data such as "Please tell me how to make curry" and performs natural language processing.
[0153] Using generative AI, the answer is generated: "Here's how to make curry: First, fry the onions. Next, add the meat, then the vegetables and simmer. Finally, add the curry roux and it's done."
[0154] The generated response data is transmitted to the terminal.
[0155] 4. Terminal
[0156] Based on the received response data, it is rendered as a 3D object.
[0157] The agent in the virtual space responds to the user using natural movements and voice, saying, "Here's how to make curry: First, fry the onions. Next, add the meat, then the vegetables and simmer. Finally, add the curry roux and it's done."
[0158] In this way, users can have a virtual reality experience that feels like they are having a conversation with a real person. The system responds immediately to detailed questions, enabling intuitive and efficient information acquisition.
[0159] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0160] Step 1:
[0161] The user puts on the virtual reality device and launches a dedicated concierge application. The input is the user putting on the virtual reality device and launching the application, and the output is the agent being ready to appear in the virtual space. Specifically, the user puts on the headset and controllers and clicks on the application icon to launch it.
[0162] Step 2:
[0163] The user asks a question by voice or text. The input is the user's speech or text input, and the output is the data of the user's question. Specifically, the user speaks "Tell me how to make tomato sauce" or enters similar text on a virtual keyboard.
[0164] Step 3:
[0165] The device records the user's voice input in real time and converts it into text data using the Google Speech-to-Text API or Microsoft Azure Speech Service. The input is the user's voice data and the output is text data. Specifically, the voice recognition software analyzes the voice data and converts it into text such as "Please tell me how to make tomato sauce."
[0166] Step 4:
[0167] The terminal assembles the converted text data into packets as question data and sends them to the server via the HTTPS protocol. The input is the text data converted from the voice, and the output is the question data sent to the server. Specifically, the terminal encodes and sends the data to the server.
[0168] Step 5:
[0169] The server receives the query data sent from the terminal. The input is the query data sent from the terminal, and the output is raw data for analysis. In concrete terms, the web server (e.g., Apache or Nginx) passes the received data to the application server.
[0170] Step 6:
[0171] The server analyzes the received question data using natural language processing (NLP). The input is the received question data, and the output is tokenized and normalized data. Specifically, an NLP engine implemented in Python, Java, or other languages tokenizes the data and performs grammatical and semantic analysis.
[0172] Step 7:
[0173] The server inputs the preprocessed data into a generative artificial intelligence model (e.g., OpenAI GPT-3) to generate an appropriate answer. The input is natural language processed data, and the output is the generated response data. Specifically, it inputs a prompt sentence (e.g., "Please tell me how to make tomato sauce.") into GPT-3 and obtains the generated response (e.g., "Here's how to make tomato sauce. First, chop the tomatoes...").
[0174] Step 8:
[0175] The server formats the generated response data in a format such as JSON and sends it to the terminal via the HTTPS protocol. The input is the generated response data, and the output is the response data sent to the terminal. Specifically, the server formats, encodes, and sends the data.
[0176] Step 9:
[0177] The terminal receives the response data sent from the server and checks and converts the format. The input is the response data sent from the server, and the output is renderable data. Specific operations include decoding the received data and generating speech synthesis and action generation data for the agent.
[0178] Step 10:
[0179] The device uses a 3D engine such as Unity or Unreal Engine to display the agent in a virtual space. The input is renderable data, and the output is the agent's response, presented visually and audibly. Specifically, the agent explains, with hand gestures and facial expressions, "Here's how to make tomato sauce. First, chop the tomatoes..."
[0180] (Application example 1)
[0181] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0182] In recent years, there have been many attempts to utilize technology to improve the customer experience in brick-and-mortar stores. However, current systems often make it difficult for users to quickly obtain detailed product information, especially when visiting a store for the first time or in an unfamiliar product category. Furthermore, when customers are hesitant to directly interact with store staff or when the store is crowded, quickly obtaining information becomes even more difficult. Therefore, there is a need for a system that allows users to easily obtain product information by voice and presents that information visually and audibly.
[0183] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0184] In this invention, the server includes: means for a user to wear a virtual reality device; means for collecting voice or text input by the user; means for analyzing the collected user question data and performing natural language processing; means for inputting the analyzed data into a generative artificial intelligence model and generating an appropriate response; means for three-dimensionally displaying the generated response data in the form of a concierge in a virtual reality space; means for forming a user interface for providing product information in a physical store; and means for generating answers to the user's voice questions to provide product information and presenting them to the user by voice, thereby enabling users to quickly and easily obtain product information in a physical store and receive that information visually and audibly.
[0185] A "user" is a person who uses the system to obtain information.
[0186] A "virtual reality device" is a device worn by a user to access a virtual reality space that includes vision and sound.
[0187] A "voice or text collection means" is any device or software that collects user input in the form of voice or text.
[0188] "Question data" is data including a question entered by a user.
[0189] "Natural language processing" is the process of analyzing collected question data and converting it into a format that a computer can understand.
[0190] A "generative artificial intelligence model" is an artificial intelligence model that inputs natural language processed data and generates an appropriate response.
[0191] "Response data" is data that includes an answer generated by a generative artificial intelligence model.
[0192] The "means for three-dimensional display" refers to a device or software for displaying the response data in the form of a three-dimensional concierge in a virtual reality space.
[0193] A "brick and mortar store" is a retail outlet that exists in a physical location.
[0194] "User interface" means an interface through which a user inputs information and interacts with the concierge system.
[0195] "Product information" is detailed information about the products being sold.
[0196] The "means for providing by voice" refers to a device or software for presenting the generated response data to the user in voice format.
[0197] The present invention provides a system that allows users to acquire product information in a physical store using a virtual reality device. The system displays appropriate responses to questions in the form of a three-dimensional concierge and presents them to the user via audio.
[0198] System configuration
[0199] The main components of the system are:
[0200] 1. Users
[0201] 2. Terminal
[0202] 3. Server
[0203] User operations
[0204] The user puts on the virtual reality device and asks a question about a product in the store by voice, for example, "What are the ingredients in this product?" This voice is transmitted to the terminal through the microphone of the virtual reality device.
[0205] Terminal handling
[0206] The device processes the user's voice input in the following steps:
[0207] 1. Use speech recognition software (e.g., SpeechRecognition library, Google Text-to-Speech) to convert the audio data into text data.
[0208] 2. The converted text data is packed into packets and sent to the server.
[0209] 3. Receive the response data from the server and format it appropriately.
[0210] 4. Render the generated answer as a 3D object and a virtual concierge will present the answer visually and audibly.
[0211] Server Processing
[0212] The server processes the query data in the following steps:
[0213] 1. Receive text data sent from the device.
[0214] 2. Perform natural language processing (NLP) and analyze the data.
[0215] 3. Use a generative AI model (e.g., GPT-3) to generate an appropriate answer. For example, if a user asks, "What are the ingredients in this product?", the generative AI will generate an answer such as, "This product's ingredients include water, glycerin, parabens, vitamin E, etc."
[0216] 4. The generated response data is sent to the terminal.
[0217] Hardware and software used
[0218] Hardware: Virtual reality devices (smart glasses, head-mounted displays, etc.), microphones, smartphones
[0219] Software: SpeechRecognition Library, Google Text-to-Speech (gTTS), OpenAI API
[0220] Specific examples
[0221] For example, if a user is browsing products in a physical store and asks through smart glasses, "What are the ingredients in this product?", the voice is sent to the device via a microphone. The device converts the voice into text, and this text data is sent to the server. The server uses a generative AI model to generate an answer, returning, for example, "The ingredients in this product include water, glycerin, parabens, vitamin E, etc." to the device. The device then provides this answer to the user as a 3D display and audio.
[0222] Prompt Sentence Examples
[0223] An example of a prompt sentence is, "A user asks a question about a product in a store. Question: 'What are the ingredients in this product?' Please provide an easy-to-understand answer to this question." By inputting this prompt sentence into a generative AI model, an appropriate answer will be generated.
[0224] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0225] Step 1:
[0226] The user wears a virtual reality device and asks a question about a product by voice. This voice is collected by the built-in microphone of the virtual reality device. The input is the user's voice, and the output is voice data.
[0227] Step 2:
[0228] The device uses speech recognition software (e.g., SpeechRecognition library, Google Text-to-Speech) to convert the collected voice data into text data. Here, speech is converted into text, such as "Please tell me the ingredients of this product." The input is voice data, and the output is text data.
[0229] Step 3:
[0230] The terminal packs the converted text data into packets and sends them to the server. The input is text data and the output is packet data.
[0231] Step 4:
[0232] The server receives packet data sent from the terminal and analyzes the data using natural language processing (NLP), where text data is tokenized and normalized. The input is packet data, and the output is analyzed data.
[0233] Step 5:
[0234] The server inputs the analyzed data into a generative artificial intelligence model (e.g., GPT-3) to generate an appropriate answer. For example, in response to the question, "What are the ingredients of this product?", an answer such as, "This product contains water, glycerin, parabens, vitamin E, etc." is generated. The input is the analyzed data, and the output is the generated answer data.
[0235] Step 6:
[0236] The server sends the generated response data to the terminal, where the input is the generated response data and the output is the reformatted response data.
[0237] Step 7:
[0238] The device then formats the answer data received from the server and renders it as a 3D object, where a concierge in a virtual reality space presents the generated answer visually and audibly. The input is the reformatted answer data, and the output is 3D visual and audio data.
[0239] Step 8:
[0240] The user receives visual and audio responses from the concierge in the virtual reality space and confirms detailed product information. The input is 3D visual and audio data, and the output is the user's understanding and acquisition of product information.
[0241] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0242] The present invention combines a 3D concierge system that combines VR technology and generative AI with an emotion engine that recognizes the user's emotions. This system allows the user to wear a virtual reality device, ask questions by voice or text, and display responses in 3D, while simultaneously generating appropriate responses based on the user's emotions. Specific embodiments for implementing the present invention are described below.
[0243] Program processing
[0244] The system consists of three main components: the user, the device, and the server, among which an emotion engine has been newly incorporated. Each plays a specific role and works in conjunction with each other.
[0245] User operations
[0246] 1. Booting the system
[0247] The user puts on the virtual reality device and launches a dedicated concierge application, at which point a concierge appears in the virtual space.
[0248] 2. Enter your question
[0249] The user asks a question by voice or text. For example, the user says, "What's the weather like in Tokyo today?"
[0250] Terminal handling
[0251] 3. Speech Recognition and Text Conversion
[0252] The user's voice input is recorded in real time by the terminal and converted into text data by voice recognition software.
[0253] 4. Emotion recognition
[0254] The device's built-in emotion engine analyzes the user's voice tone and facial expressions to generate emotion data, which identifies the user's emotional state (e.g., joy, sadness, anger, etc.).
[0255] 5. Sending Question Data and Emotion Data
[0256] The converted text question data and the analyzed emotion data are packaged into packets and sent to the server.
[0257] Server Processing
[0258] 6. Analysis of received data
[0259] The server receives and confirms the question data and emotion data sent from the terminal.
[0260] 7. Response Generation Using Generative AI
[0261] The question data is processed using natural language processing and input into a generative AI model. Emotional data is then incorporated to generate a response that reflects the user's emotions.
[0262] For example, if a user asks in a sad tone, "How's the weather in Tokyo today?", the generative AI will generate an answer that takes emotions into consideration, such as, "The weather in Tokyo is sunny today. It's nice, so why not go outside for a bit to change your mood?"
[0263] 8. Response Data Formatting and Transmission
[0264] The generated response data is formatted and sent to the terminal.
[0265] Terminal Processing (cont.)
[0266] 9. Receiving and displaying response data
[0267] The device receives the response data sent from the server and renders it as a three-dimensional object.
[0268] The concierge in the virtual space responds to the user with natural movements and voice, and also responds according to the user's emotions.
[0269] Specific examples
[0270] For example, the specific operations and processing flow when a user asks, "I'm really tired from work today. Please tell me how to relax" in a stressful daily life is shown below.
[0271] User
[0272] Put on the virtual reality device and say, "I'm really tired from work today. Please tell me how to relax."
[0273] Terminal
[0274] The user's speech is recorded and converted into text using speech recognition software.
[0275] The emotion engine detects fatigue and stress from the user's voice tone and generates emotion data.
[0276] The question data and emotion data are packaged into a packet and sent to the server.
[0277] server
[0278] Receives question data and emotion data and performs natural language processing.
[0279] Using generative AI, in response to the question, "I'm really tired from work today. Please tell me how to relax," the system generates a response that takes emotions into consideration, such as, "To relax, why not try taking some slow, deep breaths or taking a warm bath? I also recommend listening to your favorite music."
[0280] The generated response data is formatted and sent to the terminal.
[0281] Terminal
[0282] Based on the received response data, the concierge in the virtual space responds to the user with natural movements and voice.
[0283] The concierge tells the user, "To relax, why not try taking some slow, deep breaths or taking a warm bath? We also recommend listening to your favorite music."
[0284] In this way, by having each step work in conjunction with one another, a system is realized that allows a user to ask questions in a virtual reality space and receive answers that correspond to their emotions.
[0285] The processing flow will be explained below.
[0286] Step 1:
[0287] The user puts on the virtual reality device, turns on the device, and launches the dedicated concierge application, which displays the concierge in the virtual space.
[0288] Step 2:
[0289] The user types a question. The user can either speak aloud, such as "Work today is really tiring. What are some ways to relax?" or type the question as text using the virtual keyboard.
[0290] Step 3:
[0291] The device collects voice data: the user's speech is recorded by the VR device's microphone and sent to the voice recognition software.
[0292] Step 4:
[0293] The device converts the voice data into text data. Voice recognition software analyzes the voice data in real time and generates text data such as, "I'm really tired from work today. Please tell me how to relax."
[0294] Step 5:
[0295] The device collects the user's voice tone and facial expressions, and the emotion engine analyzes the user's voice tone and facial expressions to identify the user's emotional state.
[0296] Step 6:
[0297] The terminal generates emotion data, and based on the collected emotion information, generates emotion data indicating that the user is tired.
[0298] Step 7:
[0299] The device sends the question data and emotion data to the server, which then assembles the converted text data and generated emotion data into packets and sends them to the server via the Internet.
[0300] Step 8:
[0301] The server receives the question data and emotion data, analyzes the received data, and prepares it for processing.
[0302] Step 9:
[0303] The server performs natural language processing, tokenizing and normalizing the question data and converting it into a format that can be input to a generative AI model.
[0304] Step 10:
[0305] The server generates responses using a generative artificial intelligence model. It inputs question data and emotion data into the model and generates an appropriate response.
[0306] Step 11:
[0307] The server formats the response data, examining the generated response and formatting it in a format that can be displayed within the virtual reality space.
[0308] Step 12:
[0309] The server sends the formatted response data to the terminal, which then packages the response data into packets and sends them to the terminal over the Internet.
[0310] Step 13:
[0311] The terminal receives the response data. The terminal receives the response data sent from the server and checks its format.
[0312] Step 14:
[0313] The device renders the response data as a 3D object, and processes the received data so that the concierge can respond in a virtual space using natural movements and voice.
[0314] Step 15:
[0315] The user confirms the concierge's response by visually and audibly hearing the concierge in the virtual reality space reply to the user, "To relax, why not try some slow, deep breathing or a warm bath? We also recommend listening to your favorite music."
[0316] In this way, each step works in conjunction to realize a system that allows a user to ask questions in a virtual reality space and receive appropriate answers based on their emotions.
[0317] Example 2
[0318] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0319] Conventional virtual reality concierge systems can generate responses to questions entered by users, but they are unable to provide appropriate responses that take the user's emotions into consideration. This limits the user experience and prevents the user from receiving the personalized responses they desire.
[0320] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0321] In this invention, the server includes a means for allowing a user to wear a virtual reality device, a means for collecting voice or text input by the user, a built-in emotion engine for generating emotion data from the collected voice data, and a means for inputting the analyzed data into a generative AI model to generate an appropriate response, thereby enabling a response that takes into account the user's emotions.
[0322] A "virtual reality device" is a device that allows a user to experience a virtual space through sight and sound.
[0323] "Voice or text collection means" refers to hardware and software components for collecting user-spoken voice or text information.
[0324] "Means for analyzing question data and performing natural language processing" refers to technology for analyzing the content of questions entered by users and processing them in an appropriate manner.
[0325] A "generative artificial intelligence model" is a set of artificial intelligence algorithms that perform natural language responses and other generative tasks based on given data.
[0326] "Means for three-dimensionally displaying response data in the form of a concierge in a virtual reality space" refers to a method that utilizes three-dimensional graphics technology to communicate the generated response to the user visually and audibly.
[0327] The "emotion engine" is an analysis device and a group of algorithms for extracting emotional data from a user's tone of voice and facial expressions.
[0328] The "means for generating a response that takes emotion into consideration" refers to a process and technology for generating a response that is appropriate to a user's emotion based on the user's emotion data.
[0329] This invention is a 3D concierge system that combines virtual reality technology and generative artificial intelligence, and further incorporates an emotion engine that recognizes the user's emotions, thereby providing responses that take the user's emotions into consideration. The system allows the user to wear a virtual reality device, input questions by voice or text, and displays the responses in 3D while simultaneously generating appropriate responses according to the user's emotions.
[0330] This system consists of three main components: the user, the terminal, and the server, each of which plays a specific role and works in conjunction with each other. The emotion engine has been newly incorporated, and the operation of each component is shown below.
[0331] User operations
[0332] First, the user puts on a virtual reality device (e.g., a "VR headset" as it is commonly called) and launches a dedicated concierge application. At this point, a three-dimensional concierge appears in the virtual space. The user asks questions through a voice recognition and text input interface. For example, when the user says, "How is the weather in Tokyo today?", the voice is sent to the device.
[0333] Terminal handling
[0334] The device records the user's voice input in real time and converts it into text data using voice recognition software (e.g., generically called a "voice recognition engine"). The device also has a built-in emotion engine (e.g., generically called an "emotion analysis module") that recognizes the user's emotions and analyzes voice tone and facial expression data to generate emotion data. For example, if a user says in a tired voice, "I'm really tired from work today. Please tell me how to relax," the emotion will be identified as "fatigue." The converted text question data and analyzed emotion data are packaged into a data packet and sent to the server.
[0335] Server Processing
[0336] The server receives the data packet sent from the device and analyzes its contents. It uses a natural language processing (NLP) module (e.g., a generic term "natural language processing engine") to analyze the question text and generate a prompt sentence, which is input to a generative artificial intelligence model (e.g., a generic term "generative AI model"). The generated prompt sentence also includes emotional data. For example, the following content is input to the generative AI model: "A user asks, 'I'm really tired from work today. Please tell me how to relax.' This user is feeling tired. Please suggest relaxation methods to the user that take their emotions into consideration." The generative AI model then generates an appropriate response to the question, such as, "Deep breathing and a warm bath are good ways to relax. I also recommend listening to your favorite music."
[0337] The generated response data is formatted and sent to the device. The device then renders the received response data as a 3D object. The concierge responds to the user in the virtual space using natural movements and voice. For example, if the user asks, "I'm really tired from work today. What are some ways to relax?" the concierge will tell the user, "Deep breathing and a warm bath are good ways to relax. I also recommend listening to your favorite music."
[0338] In this way, a system is realized that allows a user to ask questions in a virtual reality space and receive answers that take emotion into consideration.
[0339] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0340] Step 1:
[0341] The user wears a virtual reality device and launches a dedicated concierge application, which displays a 3D concierge in the virtual space. The input is the virtual reality device worn by the user and the application that operates it, and the output is the display of the 3D concierge.
[0342] Step 2:
[0343] The user inputs a question by voice or text. For example, they might say, "What's the weather like in Tokyo today?" The input is the user's voice or text, which is sent to the device. The output is the voice data sent to the device.
[0344] Step 3:
[0345] The device records the user's voice input in real time and converts it into text data using voice recognition software. The voice recognition software used is a general voice recognition engine. The input for text conversion is the recorded voice data, and the output is text data. Specifically, the voice signal is frequency analyzed and converted into text based on a language model.
[0346] Step 4:
[0347] The device's emotion engine analyzes the user's voice tone and facial expressions to generate emotion data. The emotion engine uses a general emotion analysis module to analyze voice intonation and facial muscle movements. The input is the user's voice tone and facial expression data, and the output is emotion data.
[0348] Step 5:
[0349] The terminal assembles the converted text question data and the analyzed emotion data into packets and sends them to the server. The input is text data and emotion data, and the output is the data packet sent to the server.
[0350] Step 6:
[0351] The server receives data packets sent from the terminal and analyzes their contents. The input is the received data packet, and the output is the analyzed question data and emotion data. The server checks the consistency of the data and requests a retransmission if there is an inconsistency.
[0352] Step 7:
[0353] The server uses a natural language processing engine to analyze the question text, generate a prompt sentence, and input it to the generative artificial intelligence model. The generative artificial intelligence model uses a general generative AI model. The input is question data and emotional data, and the output is response data that takes the emotion into consideration. For example, the prompt sentence input is, "The user asked, 'I'm really tired from work today. Please tell me how to relax.' This user is feeling tired. Please suggest a relaxation method for the user that takes their emotions into consideration."
[0354] Step 8:
[0355] The server formats the generated response data and sends it to the terminal. The input is the generated response data and the output is the formatted data sent to the terminal.
[0356] Step 9:
[0357] The device renders the received response data as a 3D object, using a 3D graphics engine. The input is the response data received from the server, and the output is a response made by the concierge in the virtual space using natural movements and voice. The concierge naturally tells the user, "To relax, deep breathing and a warm bath are good ways to do it. I also recommend listening to your favorite music."
[0358] (Application example 2)
[0359] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0360] Conventional systems using virtual reality technology have the ability to provide natural responses to user questions, but they are unable to provide individual responses based on the user's emotions. This limits the user's satisfaction and the quality of the experience. The objective of this invention is to provide a more personalized experience by recognizing the user's emotions and providing appropriate responses based on those emotions.
[0361] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0362] In this invention, the server includes a means for incorporating an emotion engine that recognizes and analyzes the user's emotional state, a means for inputting the data into a generative AI model to generate an appropriate response, and a means for displaying the generated response data in three dimensions in the form of a concierge in a virtual reality space, thereby enabling personalized responses according to the user's emotions.
[0363] A "virtual reality device" is a device worn by a user to immerse themselves in a virtual reality space, and refers to devices such as head-mounted displays and VR goggles.
[0364] "Voice or text collection means" refers to an input device such as a microphone or keyboard for capturing voice or text data input by a user in real time.
[0365] "Means for analyzing question data and performing natural language processing" refers to software or algorithms for analyzing collected user questions and semantically analyzing the text using natural language processing technology.
[0366] A "generative artificial intelligence model" is an artificial intelligence model that takes natural language processed data as input and generates appropriate responses, and refers to a model that generates responses based on learned patterns.
[0367] "Means for three-dimensional display" refers to rendering software and display devices for displaying the generated response in three dimensions within a virtual reality space.
[0368] An "emotion engine" refers to hardware or software that analyzes a user's tone of voice, facial expressions, etc., to identify the user's emotional state (e.g., joy, sadness, anger, etc.).
[0369] "Means for generating and displaying a response according to emotions" refers to a series of processes or devices that allow a generative artificial intelligence model to create a response based on the emotional state identified by the emotion engine and to display that response appropriately.
[0370] This invention is a 3D concierge system that combines virtual reality technology and generative artificial intelligence, incorporating an emotion engine that recognizes the user's emotions. With this system, the user puts on a virtual reality device and asks a question. The response to the question is displayed in 3D, and the user can receive an appropriate response according to their emotion.
[0371] The system mainly consists of three main components: users, terminals, and servers.
[0372] User operations
[0373] First, the user puts on a virtual reality device and launches a dedicated concierge application. This refers to a virtual reality device such as a head-mounted display or VR goggles. When the user asks a question by voice or text, the input is collected by the device. For example, the user might say, "Tell me about this product."
[0374] Terminal handling
[0375] The device converts the user's speech into text data using voice recognition software (e.g., Vosk). The device's built-in emotion engine analyzes the user's tone of voice and facial expressions to generate emotion data. The emotion engine uses OpenFace, for example. For example, if a user asks a question in a curious tone, the device identifies the user's curious emotional state.
[0376] Server Processing
[0377] The server receives the question data and emotion data sent from the device and performs natural language processing. The question data is input into a generative artificial intelligence model (such as OpenAI's GPT-3.5 Turbo) to generate a response based on the user's emotion. For example, if a user asks out of curiosity, "Tell me about this product," the server generates a response such as, "This product is new this season, made of silk, and is characterized by its extremely soft feel."
[0378] Terminal Processing (cont.)
[0379] The device receives the response data sent from the server and displays it in 3D as a concierge in the virtual reality space. This display is rendered using a VR library (e.g., VRRenderer). It is also possible to respond by voice using pyttsx3 or similar.
[0380] Prompt Sentence Examples
[0381] For example, the following prompt statement is generated:
[0382] The user was curious and asked: "Tell me about this product."
[0383] As described above, the present invention aims to enable users to ask questions in a virtual reality space and receive personalized responses to those questions based on their emotions, thereby providing a richer experience for the user.
[0384] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0385] Step 1: User interaction
[0386] The user puts on the virtual reality device and launches a dedicated concierge application. A concierge appears in the virtual space, and the user asks a question by voice or text. Input (voice or text) begins. For example, the user might say, "Tell me about this product."
[0387] Step 2: Collecting audio and converting it to text
[0388] The device records the user's speech with a microphone and converts it into text data using voice recognition software (e.g., Vosk). In the process of converting voice data (input) into text data (output), noise removal and voice analysis are performed. Specifically, the recorded voice data is converted into text data such as "Tell me about this product."
[0389] Step 3: Recognize emotions
[0390] The device's built-in emotion engine (e.g., OpenFace) captures the user's voice tone and facial expressions with a camera and analyzes the data. It then identifies the type of emotion and generates emotion data (e.g., curious). The emotion data (output) is packaged into a packet of question data (text data).
[0391] Step 4: Send data
[0392] The terminal assembles the converted question data and the analyzed emotion data into packets and sends them to the server. The input (question data and emotion data) is configured as a data packet to be sent to the server.
[0393] Step 5: Receive and analyze question and sentiment data
[0394] The server receives the question data and emotion data sent from the device and analyzes each data. It uses natural language processing (NLP) to analyze the question data and integrate the emotion data. It then outputs a prompt sentence generated by analyzing the data.
[0395] Step 6: Generate a response
[0396] The server inputs the prompt into a generative artificial intelligence model (e.g., GPT-3.5 Turbo) to generate an appropriate response. In this scenario, if the user's emotion of curiosity is identified, a detailed response tailored to that interest is generated. For example, a response such as "This product is new this season, made of silk, and is characterized by its extremely soft feel" may be output.
[0397] Step 7: Format and send response data
[0398] The server formats the generated response data and sends it as a data packet to the terminal. The response data (output) is now properly formatted and ready to be sent.
[0399] Step 8: Receiving response data and displaying it in 3D
[0400] The device receives the response data sent from the server and displays it in 3D as a concierge in the virtual space. Using VR rendering software (e.g., VRRenderer), the concierge responds with natural movements and voice. Specifically, the concierge explains, "This product is new this season, made of silk, and is extremely soft to the touch."
[0401] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0402] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0403] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0404] [Second embodiment]
[0405] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0406] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0407] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0408] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0409] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0410] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0411] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0412] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0413] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0414] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0415] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0416] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0417] The present invention provides a 3D concierge system that combines VR technology and generative AI. In this system, a user wears a virtual reality device, asks questions by voice or text, and an appropriate response is displayed in 3D. Specific embodiments for implementing the present invention are described below.
[0418] Program processing
[0419] The system consists of three main components: users, terminals, and servers, each of which plays a specific role and works in tandem.
[0420] User operations
[0421] 1. Booting the system
[0422] The user puts on the virtual reality device and launches a dedicated concierge application, at which point a concierge appears in the virtual space.
[0423] 2. Enter your question
[0424] The user asks a question by voice or text, for example, while cooking, "Tell me how to make tomato sauce."
[0425] Questions are entered into the system through the microphone or virtual keyboard of the virtual reality device.
[0426] Terminal handling
[0427] 3. Speech Recognition and Text Conversion
[0428] The user's voice input is recorded in real time by the terminal and converted into text data by voice recognition software.
[0429] The converted text data is packed into packets as question data and sent to the server.
[0430] 4. Receiving response data
[0431] Receives response data sent from the server, checks and converts the format.
[0432] The concierge then formats the responses appropriately so that they can be given to the user using natural movements and voice.
[0433] 5. Rendering 3D Objects
[0434] The device renders the received response data as a three-dimensional object.
[0435] A concierge in the virtual space presents the generated answer to the user visually and audibly.
[0436] Server Processing
[0437] 6. Receiving and analyzing query data
[0438] The server receives the query data sent from the terminal.
[0439] The received data is pre-processed for natural language processing, including tokenization and normalization.
[0440] 7. Answer generation using generative AI
[0441] The natural language processed data is input into a generative artificial intelligence model, which generates an appropriate answer.
[0442] For example, if you input "How to make tomato sauce," the generative AI will generate the answer, "Here's how to make tomato sauce: First, chop the tomatoes. Next, heat the olive oil in a frying pan, add the garlic and fry until fragrant. Then, adjust the flavor with the chopped tomatoes, salt and pepper, and simmer."
[0443] 8. Sending response data
[0444] The generated response data is appropriately formatted and transmitted to the terminal.
[0445] Specific examples
[0446] For example, the specific operations and processing flow when a user asks "Please tell me how to make curry" while cooking will be shown.
[0447] User
[0448] Put on the virtual reality device and say, "Tell me how to make curry."
[0449] Terminal
[0450] The user's speech is recorded and converted into text data using voice recognition software, such as "Please tell me how to make curry."
[0451] The converted text data is sent to the server.
[0452] server
[0453] It receives text data such as "Please tell me how to make curry" and performs natural language processing.
[0454] Using generative AI, the answer is generated: "Here's how to make curry: First, fry the onions. Next, add the meat, then the vegetables and simmer. Finally, add the curry roux and it's done."
[0455] The generated response data is transmitted to the terminal.
[0456] Terminal
[0457] Based on the received response data, it is rendered as a 3D object.
[0458] The concierge in the virtual space responds to the user using natural movements and voice, saying, "Here's how to make curry: First, fry the onions. Next, add the meat, then the vegetables and simmer. Finally, add the curry roux and it's done."
[0459] In this way, users can have a virtual reality experience that feels like they are having a conversation with a real person. The system responds immediately to detailed questions, enabling intuitive and efficient information acquisition.
[0460] The processing flow will be explained below.
[0461] Step 1:
[0462] The user puts on the virtual reality device and presses the start button on the device to launch the dedicated concierge application, which causes a concierge to appear in the virtual space.
[0463] Step 2:
[0464] The user enters a question. The user speaks into the VR device's microphone, such as "What's the weather like in Tokyo today?", or enters the question as text using the virtual keyboard.
[0465] Step 3:
[0466] The device collects voice data. The user's speech is recorded by the VR device's microphone and converted into text data in real time.
[0467] Step 4:
[0468] The device sends the text data to the server. The converted text data is packed into packets and sent to the server via the Internet.
[0469] Step 5:
[0470] The server receives the text data. The server receives and checks the data sent from the terminal.
[0471] Step 6:
[0472] The server analyzes the text data received, performs preprocessing for natural language processing, and tokenizes and normalizes the text data.
[0473] Step 7:
[0474] The server inputs the data into a generative AI model, which then inputs the analyzed text data into the model to generate an appropriate response.
[0475] Step 8:
[0476] The server formats the generated response data, and examines the generated response data and converts it into a format for display within the virtual reality space.
[0477] Step 9:
[0478] The server sends the formatted response data to the terminal, which then packages the response data into packets and sends them over the Internet to the terminal.
[0479] Step 10:
[0480] The terminal receives the response data. The terminal receives the response data sent from the server and checks its format.
[0481] Step 11:
[0482] The device renders the response data as a 3D object, and processes the received data so that the 3D concierge can respond to the user in a virtual space using natural movements and voice.
[0483] Step 12:
[0484] The user confirms the answer from the Moving Concierge. The concierge in the virtual reality space responds to the user, saying, "The weather in Tokyo is sunny today." The user confirms this visually and audibly.
[0485] In this way, each step works in cooperation to realize a system that allows a user to ask a question in a virtual reality space and receive an appropriate response immediately.
[0486] Example 1
[0487] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0488] In conventional virtual reality systems, it has been difficult to obtain natural responses in real time when a user asks a question. In particular, the processing required to provide detailed and appropriate responses to a user's question is complex, resulting in a non-intuitive user experience. Furthermore, conventional systems have a problem in that the quality of the visual and audio representation of the responses is low, resulting in a lack of realism.
[0489] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0490] In this invention, the server includes a means for a user to wear a virtual reality device, a means for collecting voice or text input by the user, a means for analyzing the collected user question data, a means for natural language processing the analyzed data, a means for inputting the processed natural language data into an artificial intelligence model for generating an appropriate response, a means for three-dimensionally displaying the generated response data in the form of an agent in a virtual reality space, and a means for presenting the generated response data in visual and audio formats, thereby enabling the user to have a more realistic interactive experience.
[0491] "User" refers to a person who uses the system.
[0492] "Virtual reality device" refers to equipment used by a user to immerse themselves in a virtual reality environment, including, but not limited to, a headset and a handheld controller.
[0493] "Voice or text collection means" refers to devices or software that capture user-input voice or text data. Voice input is accomplished via a microphone, and text input is accomplished via a virtual keyboard or voice recognition software.
[0494] "Question data" refers to data including the question entered by the user.
[0495] "Analysis tools" refers to software and algorithms used to understand the collected question data and perform processes such as parsing, tokenization, and normalization.
[0496] "Natural language processing" refers to the technology that enables a machine to understand and appropriately process input human language. Specifically, it includes tokenization, grammatical analysis, and semantic analysis.
[0497] "Generative artificial intelligence model" refers to a machine learning algorithm or artificial intelligence technique that generates an appropriate response based on natural language processed data. Specific examples include generative AI models.
[0498] "Response data" refers to data containing answers generated by an artificial intelligence model.
[0499] "Means for three-dimensional display" refers to devices and software for displaying the generated response data in three-dimensional space, specifically including a 3D engine for visually displaying the response to an agent (concierge) in a virtual reality space.
[0500] "Means for visual and audio presentation" refers to devices and software for presenting the generated response data to the user in visual animation and audio, including, for example, speech synthesis software that converts text into speech and animation software that controls the behavior of an agent.
[0501] This invention relates to a 3D concierge system that combines virtual reality (VR) technology and generative artificial intelligence (AI). Specifically, the system provides a user with a virtual reality device, asks questions by voice or text, and displays appropriate responses in 3D. The system consists of three main components: a user, a terminal, and a server.
[0502] System configuration
[0503] User operations
[0504] The user wears a virtual reality device (such as Oculus Rift or HTC Vive). Through this device, the user can launch a dedicated concierge application and ask questions to a three-dimensional agent that appears in the virtual space. Questions can be asked by voice input or text input. Voice input uses a microphone, and text input uses a virtual keyboard.
[0505] Terminal handling
[0506] Speech recognition and text conversion
[0507] The device records the user's voice input in real time and converts it into text using speech recognition software such as the Google Speech-to-Text API or Microsoft Azure Speech Service. For example, if a user says, "Tell me how to make tomato sauce," this will be converted into text.
[0508] Sending text data
[0509] The converted text data is packaged into packets as question data and sent to the server via the HTTPS protocol.
[0510] Receiving and processing response data
[0511] Receives the response data sent from the server, checks and converts the format. Specifically, it converts text data into audio data and prepares it into data for setting the motion of a 3D agent.
[0512] Rendering 3D objects
[0513] The device uses a 3D engine such as Unity or Unreal Engine to display the virtual space, and the 3D agent uses natural movements (such as hand gestures and facial expressions) to present the generated responses visually and audibly.
[0514] Server Processing
[0515] Receiving and analyzing question data
[0516] The server receives the query data sent from the terminal. This data is received through a web server (such as Apache or Nginx). The received data is analyzed by a server-side application written in Python or Java. Natural language processing such as tokenization and normalization is performed in the initial analysis stage.
[0517] Answer generation using generative AI
[0518] The analyzed data is input into a generative AI model (e.g., OpenAI GPT-3). The generative AI model generates an appropriate response based on a prompt (e.g., "Please tell me how to make tomato sauce.") For example, it might generate an answer like, "Here's how to make tomato sauce: First, chop the tomatoes. Next, heat olive oil in a frying pan, add the garlic, and fry until fragrant. Next, add the chopped tomatoes, salt, and pepper to taste, and simmer."
[0519] Sending response data
[0520] The generated response data is formatted in a format such as JSON and sent to the terminal via the HTTPS protocol.
[0521] Specific examples
[0522] For example, the specific operations and processing flow when a user asks "Please tell me how to make curry" while cooking will be described.
[0523] 1. Users
[0524] Put on the VR device and say, "Please tell me how to make curry."
[0525] 2. Terminal
[0526] The user's speech is recorded and converted into text data using voice recognition software, such as "Please tell me how to make curry."
[0527] The converted text data is sent to the server.
[0528] 3. Server
[0529] It receives text data such as "Please tell me how to make curry" and performs natural language processing.
[0530] Using generative AI, the answer is generated: "Here's how to make curry: First, fry the onions. Next, add the meat, then the vegetables and simmer. Finally, add the curry roux and it's done."
[0531] The generated response data is transmitted to the terminal.
[0532] 4. Terminal
[0533] Based on the received response data, it is rendered as a 3D object.
[0534] The agent in the virtual space responds to the user using natural movements and voice, saying, "Here's how to make curry: First, fry the onions. Next, add the meat, then the vegetables and simmer. Finally, add the curry roux and it's done."
[0535] In this way, users can have a virtual reality experience that feels like they are having a conversation with a real person. The system responds immediately to detailed questions, enabling intuitive and efficient information acquisition.
[0536] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0537] Step 1:
[0538] The user puts on the virtual reality device and launches a dedicated concierge application. The input is the user putting on the virtual reality device and launching the application, and the output is the agent being ready to appear in the virtual space. Specifically, the user puts on the headset and controllers and clicks on the application icon to launch it.
[0539] Step 2:
[0540] The user asks a question by voice or text. The input is the user's speech or text input, and the output is the data of the user's question. Specifically, the user speaks "Tell me how to make tomato sauce" or enters similar text on a virtual keyboard.
[0541] Step 3:
[0542] The device records the user's voice input in real time and converts it into text data using the Google Speech-to-Text API or Microsoft Azure Speech Service. The input is the user's voice data and the output is text data. Specifically, the voice recognition software analyzes the voice data and converts it into text such as "Please tell me how to make tomato sauce."
[0543] Step 4:
[0544] The terminal assembles the converted text data into packets as question data and sends them to the server via the HTTPS protocol. The input is the text data converted from the voice, and the output is the question data sent to the server. Specifically, the terminal encodes and sends the data to the server.
[0545] Step 5:
[0546] The server receives the query data sent from the terminal. The input is the query data sent from the terminal, and the output is raw data for analysis. In concrete terms, the web server (e.g., Apache or Nginx) passes the received data to the application server.
[0547] Step 6:
[0548] The server analyzes the received question data using natural language processing (NLP). The input is the received question data, and the output is tokenized and normalized data. Specifically, an NLP engine implemented in Python, Java, or other languages tokenizes the data and performs grammatical and semantic analysis.
[0549] Step 7:
[0550] The server inputs the preprocessed data into a generative artificial intelligence model (e.g., OpenAI GPT-3) to generate an appropriate answer. The input is natural language processed data, and the output is the generated response data. Specifically, it inputs a prompt sentence (e.g., "Please tell me how to make tomato sauce.") into GPT-3 and obtains the generated response (e.g., "Here's how to make tomato sauce. First, chop the tomatoes...").
[0551] Step 8:
[0552] The server formats the generated response data in a format such as JSON and sends it to the terminal via the HTTPS protocol. The input is the generated response data, and the output is the response data sent to the terminal. Specifically, the server formats, encodes, and sends the data.
[0553] Step 9:
[0554] The terminal receives the response data sent from the server and checks and converts the format. The input is the response data sent from the server, and the output is renderable data. Specific operations include decoding the received data and generating speech synthesis and action generation data for the agent.
[0555] Step 10:
[0556] The device uses a 3D engine such as Unity or Unreal Engine to display the agent in a virtual space. The input is renderable data, and the output is the agent's response, presented visually and audibly. Specifically, the agent explains, with hand gestures and facial expressions, "Here's how to make tomato sauce. First, chop the tomatoes..."
[0557] (Application example 1)
[0558] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0559] In recent years, there have been many attempts to utilize technology to improve the customer experience in brick-and-mortar stores. However, current systems often make it difficult for users to quickly obtain detailed product information, especially when visiting a store for the first time or in an unfamiliar product category. Furthermore, when customers are hesitant to directly interact with store staff or when the store is crowded, quickly obtaining information becomes even more difficult. Therefore, there is a need for a system that allows users to easily obtain product information by voice and presents that information visually and audibly.
[0560] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0561] In this invention, the server includes: means for a user to wear a virtual reality device; means for collecting voice or text input by the user; means for analyzing the collected user question data and performing natural language processing; means for inputting the analyzed data into a generative artificial intelligence model and generating an appropriate response; means for three-dimensionally displaying the generated response data in the form of a concierge in a virtual reality space; means for forming a user interface for providing product information in a physical store; and means for generating answers to the user's voice questions to provide product information and presenting them to the user by voice, thereby enabling users to quickly and easily obtain product information in a physical store and receive that information visually and audibly.
[0562] A "user" is a person who uses the system to obtain information.
[0563] A "virtual reality device" is a device worn by a user to access a virtual reality space that includes vision and sound.
[0564] A "voice or text collection means" is any device or software that collects user input in the form of voice or text.
[0565] "Question data" is data including a question entered by a user.
[0566] "Natural language processing" is the process of analyzing collected question data and converting it into a format that a computer can understand.
[0567] A "generative artificial intelligence model" is an artificial intelligence model that inputs natural language processed data and generates an appropriate response.
[0568] "Response data" is data that includes an answer generated by a generative artificial intelligence model.
[0569] The "means for three-dimensional display" refers to a device or software for displaying the response data in the form of a three-dimensional concierge in a virtual reality space.
[0570] A "brick and mortar store" is a retail outlet that exists in a physical location.
[0571] "User interface" means an interface through which a user inputs information and interacts with the concierge system.
[0572] "Product information" is detailed information about the products being sold.
[0573] The "means for providing by voice" refers to a device or software for presenting the generated response data to the user in voice format.
[0574] The present invention provides a system that allows users to acquire product information in a physical store using a virtual reality device. The system displays appropriate responses to questions in the form of a three-dimensional concierge and presents them to the user via audio.
[0575] System configuration
[0576] The main components of the system are:
[0577] 1. Users
[0578] 2. Terminal
[0579] 3. Server
[0580] User operations
[0581] The user puts on the virtual reality device and asks a question about a product in the store by voice, for example, "What are the ingredients in this product?" This voice is transmitted to the terminal through the microphone of the virtual reality device.
[0582] Terminal handling
[0583] The device processes the user's voice input in the following steps:
[0584] 1. Use speech recognition software (e.g., SpeechRecognition library, Google Text-to-Speech) to convert the audio data into text data.
[0585] 2. The converted text data is packed into packets and sent to the server.
[0586] 3. Receive the response data from the server and format it appropriately.
[0587] 4. Render the generated answer as a 3D object and a virtual concierge will present the answer visually and audibly.
[0588] Server Processing
[0589] The server processes the query data in the following steps:
[0590] 1. Receive text data sent from the device.
[0591] 2. Perform natural language processing (NLP) and analyze the data.
[0592] 3. Use a generative AI model (e.g., GPT-3) to generate an appropriate answer. For example, if a user asks, "What are the ingredients in this product?", the generative AI will generate an answer such as, "This product's ingredients include water, glycerin, parabens, vitamin E, etc."
[0593] 4. The generated response data is sent to the terminal.
[0594] Hardware and software used
[0595] Hardware: Virtual reality devices (smart glasses, head-mounted displays, etc.), microphones, smartphones
[0596] Software: SpeechRecognition Library, Google Text-to-Speech (gTTS), OpenAI API
[0597] Specific examples
[0598] For example, if a user is browsing products in a physical store and asks through smart glasses, "What are the ingredients in this product?", the voice is sent to the device via a microphone. The device converts the voice into text, and this text data is sent to the server. The server uses a generative AI model to generate an answer, returning, for example, "The ingredients in this product include water, glycerin, parabens, vitamin E, etc." to the device. The device then provides this answer to the user as a 3D display and audio.
[0599] Prompt Sentence Examples
[0600] An example of a prompt sentence is, "A user asks a question about a product in a store. Question: 'What are the ingredients in this product?' Please provide an easy-to-understand answer to this question." By inputting this prompt sentence into a generative AI model, an appropriate answer will be generated.
[0601] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0602] Step 1:
[0603] The user wears a virtual reality device and asks a question about a product by voice. This voice is collected by the built-in microphone of the virtual reality device. The input is the user's voice, and the output is voice data.
[0604] Step 2:
[0605] The device uses speech recognition software (e.g., SpeechRecognition library, Google Text-to-Speech) to convert the collected voice data into text data. Here, speech is converted into text, such as "Please tell me the ingredients of this product." The input is voice data, and the output is text data.
[0606] Step 3:
[0607] The terminal packs the converted text data into packets and sends them to the server. The input is text data and the output is packet data.
[0608] Step 4:
[0609] The server receives packet data sent from the terminal and analyzes the data using natural language processing (NLP), where text data is tokenized and normalized. The input is packet data, and the output is analyzed data.
[0610] Step 5:
[0611] The server inputs the analyzed data into a generative artificial intelligence model (e.g., GPT-3) to generate an appropriate answer. For example, in response to the question, "What are the ingredients of this product?", an answer such as, "This product contains water, glycerin, parabens, vitamin E, etc." is generated. The input is the analyzed data, and the output is the generated answer data.
[0612] Step 6:
[0613] The server sends the generated response data to the terminal, where the input is the generated response data and the output is the reformatted response data.
[0614] Step 7:
[0615] The device then formats the answer data received from the server and renders it as a 3D object, where a concierge in a virtual reality space presents the generated answer visually and audibly. The input is the reformatted answer data, and the output is 3D visual and audio data.
[0616] Step 8:
[0617] The user receives visual and audio responses from the concierge in the virtual reality space and confirms detailed product information. The input is 3D visual and audio data, and the output is the user's understanding and acquisition of product information.
[0618] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0619] The present invention combines a 3D concierge system that combines VR technology and generative AI with an emotion engine that recognizes the user's emotions. This system allows the user to wear a virtual reality device, ask questions by voice or text, and display responses in 3D, while simultaneously generating appropriate responses based on the user's emotions. Specific embodiments for implementing the present invention are described below.
[0620] Program processing
[0621] The system consists of three main components: the user, the device, and the server, among which an emotion engine has been newly incorporated. Each plays a specific role and works in conjunction with each other.
[0622] User operations
[0623] 1. Booting the system
[0624] The user puts on the virtual reality device and launches a dedicated concierge application, at which point a concierge appears in the virtual space.
[0625] 2. Enter your question
[0626] The user asks a question by voice or text. For example, the user says, "What's the weather like in Tokyo today?"
[0627] Terminal handling
[0628] 3. Speech Recognition and Text Conversion
[0629] The user's voice input is recorded in real time by the terminal and converted into text data by voice recognition software.
[0630] 4. Emotion recognition
[0631] The device's built-in emotion engine analyzes the user's voice tone and facial expressions to generate emotion data, which identifies the user's emotional state (e.g., joy, sadness, anger, etc.).
[0632] 5. Sending Question Data and Emotion Data
[0633] The converted text question data and the analyzed emotion data are packaged into packets and sent to the server.
[0634] Server Processing
[0635] 6. Analysis of received data
[0636] The server receives and confirms the question data and emotion data sent from the terminal.
[0637] 7. Response Generation Using Generative AI
[0638] The question data is processed using natural language processing and input into a generative AI model. Emotional data is then incorporated to generate a response that reflects the user's emotions.
[0639] For example, if a user asks in a sad tone, "How's the weather in Tokyo today?", the generative AI will generate an answer that takes emotions into consideration, such as, "The weather in Tokyo is sunny today. It's nice, so why not go outside for a bit to change your mood?"
[0640] 8. Response Data Formatting and Transmission
[0641] The generated response data is formatted and sent to the terminal.
[0642] Terminal Processing (cont.)
[0643] 9. Receiving and displaying response data
[0644] The device receives the response data sent from the server and renders it as a three-dimensional object.
[0645] The concierge in the virtual space responds to the user with natural movements and voice, and also responds according to the user's emotions.
[0646] Specific examples
[0647] For example, the specific operations and processing flow when a user asks, "I'm really tired from work today. Please tell me how to relax" in a stressful daily life is shown below.
[0648] User
[0649] Put on the virtual reality device and say, "I'm really tired from work today. Please tell me how to relax."
[0650] Terminal
[0651] The user's speech is recorded and converted into text using speech recognition software.
[0652] The emotion engine detects fatigue and stress from the user's voice tone and generates emotion data.
[0653] The question data and emotion data are packaged into a packet and sent to the server.
[0654] server
[0655] Receives question data and emotion data and performs natural language processing.
[0656] Using generative AI, in response to the question, "I'm really tired from work today. Please tell me how to relax," the system generates a response that takes emotions into consideration, such as, "To relax, why not try taking some slow, deep breaths or taking a warm bath? I also recommend listening to your favorite music."
[0657] The generated response data is formatted and sent to the terminal.
[0658] Terminal
[0659] Based on the received response data, the concierge in the virtual space responds to the user with natural movements and voice.
[0660] The concierge tells the user, "To relax, why not try taking some slow, deep breaths or taking a warm bath? We also recommend listening to your favorite music."
[0661] In this way, by having each step work in conjunction with one another, a system is realized that allows a user to ask questions in a virtual reality space and receive answers that correspond to their emotions.
[0662] The processing flow will be explained below.
[0663] Step 1:
[0664] The user puts on the virtual reality device, turns on the device, and launches the dedicated concierge application, which displays the concierge in the virtual space.
[0665] Step 2:
[0666] The user types a question. The user can either speak aloud, such as "Work today is really tiring. What are some ways to relax?" or type the question as text using the virtual keyboard.
[0667] Step 3:
[0668] The device collects voice data: the user's speech is recorded by the VR device's microphone and sent to the voice recognition software.
[0669] Step 4:
[0670] The device converts the voice data into text data. Voice recognition software analyzes the voice data in real time and generates text data such as, "I'm really tired from work today. Please tell me how to relax."
[0671] Step 5:
[0672] The device collects the user's voice tone and facial expressions, and the emotion engine analyzes the user's voice tone and facial expressions to identify the user's emotional state.
[0673] Step 6:
[0674] The terminal generates emotion data, and based on the collected emotion information, generates emotion data indicating that the user is tired.
[0675] Step 7:
[0676] The device sends the question data and emotion data to the server, which then assembles the converted text data and generated emotion data into packets and sends them to the server via the Internet.
[0677] Step 8:
[0678] The server receives the question data and emotion data, analyzes the received data, and prepares it for processing.
[0679] Step 9:
[0680] The server performs natural language processing, tokenizing and normalizing the question data and converting it into a format that can be input to a generative AI model.
[0681] Step 10:
[0682] The server generates responses using a generative artificial intelligence model. It inputs question data and emotion data into the model and generates an appropriate response.
[0683] Step 11:
[0684] The server formats the response data, examining the generated response and formatting it in a format that can be displayed within the virtual reality space.
[0685] Step 12:
[0686] The server sends the formatted response data to the terminal, which then packages the response data into packets and sends them to the terminal over the Internet.
[0687] Step 13:
[0688] The terminal receives the response data. The terminal receives the response data sent from the server and checks its format.
[0689] Step 14:
[0690] The device renders the response data as a 3D object, and processes the received data so that the concierge can respond in a virtual space using natural movements and voice.
[0691] Step 15:
[0692] The user confirms the concierge's response by visually and audibly hearing the concierge in the virtual reality space reply to the user, "To relax, why not try some slow, deep breathing or a warm bath? We also recommend listening to your favorite music."
[0693] In this way, each step works in conjunction to realize a system that allows a user to ask questions in a virtual reality space and receive appropriate answers based on their emotions.
[0694] Example 2
[0695] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0696] Conventional virtual reality concierge systems can generate responses to questions entered by users, but they are unable to provide appropriate responses that take the user's emotions into consideration. This limits the user experience and prevents the user from receiving the personalized responses they desire.
[0697] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0698] In this invention, the server includes a means for allowing a user to wear a virtual reality device, a means for collecting voice or text input by the user, a built-in emotion engine for generating emotion data from the collected voice data, and a means for inputting the analyzed data into a generative AI model to generate an appropriate response, thereby enabling a response that takes into account the user's emotions.
[0699] A "virtual reality device" is a device that allows a user to experience a virtual space through sight and sound.
[0700] "Voice or text collection means" refers to hardware and software components for collecting user-spoken voice or text information.
[0701] "Means for analyzing question data and performing natural language processing" refers to technology for analyzing the content of questions entered by users and processing them in an appropriate manner.
[0702] A "generative artificial intelligence model" is a set of artificial intelligence algorithms that perform natural language responses and other generative tasks based on given data.
[0703] "Means for three-dimensionally displaying response data in the form of a concierge in a virtual reality space" refers to a method that utilizes three-dimensional graphics technology to communicate the generated response to the user visually and audibly.
[0704] The "emotion engine" is an analysis device and a group of algorithms for extracting emotional data from a user's tone of voice and facial expressions.
[0705] The "means for generating a response that takes emotion into consideration" refers to a process and technology for generating a response that is appropriate to a user's emotion based on the user's emotion data.
[0706] This invention is a 3D concierge system that combines virtual reality technology and generative artificial intelligence, and further incorporates an emotion engine that recognizes the user's emotions, thereby providing responses that take the user's emotions into consideration. The system allows the user to wear a virtual reality device, input questions by voice or text, and displays the responses in 3D while simultaneously generating appropriate responses according to the user's emotions.
[0707] This system consists of three main components: the user, the terminal, and the server, each of which plays a specific role and works in conjunction with each other. The emotion engine has been newly incorporated, and the operation of each component is shown below.
[0708] User operations
[0709] First, the user puts on a virtual reality device (e.g., a "VR headset" as it is commonly called) and launches a dedicated concierge application. At this point, a three-dimensional concierge appears in the virtual space. The user asks questions through a voice recognition and text input interface. For example, when the user says, "How is the weather in Tokyo today?", the voice is sent to the device.
[0710] Terminal handling
[0711] The device records the user's voice input in real time and converts it into text data using voice recognition software (e.g., generically called a "voice recognition engine"). The device also has a built-in emotion engine (e.g., generically called an "emotion analysis module") that recognizes the user's emotions and analyzes voice tone and facial expression data to generate emotion data. For example, if a user says in a tired voice, "I'm really tired from work today. Please tell me how to relax," the emotion will be identified as "fatigue." The converted text question data and analyzed emotion data are packaged into a data packet and sent to the server.
[0712] Server Processing
[0713] The server receives the data packet sent from the device and analyzes its contents. It uses a natural language processing (NLP) module (e.g., a generic term "natural language processing engine") to analyze the question text and generate a prompt sentence, which is input to a generative artificial intelligence model (e.g., a generic term "generative AI model"). The generated prompt sentence also includes emotional data. For example, the following content is input to the generative AI model: "A user asks, 'I'm really tired from work today. Please tell me how to relax.' This user is feeling tired. Please suggest relaxation methods to the user that take their emotions into consideration." The generative AI model then generates an appropriate response to the question, such as, "Deep breathing and a warm bath are good ways to relax. I also recommend listening to your favorite music."
[0714] The generated response data is formatted and sent to the device. The device then renders the received response data as a 3D object. The concierge responds to the user in the virtual space using natural movements and voice. For example, if the user asks, "I'm really tired from work today. What are some ways to relax?" the concierge will tell the user, "Deep breathing and a warm bath are good ways to relax. I also recommend listening to your favorite music."
[0715] In this way, a system is realized that allows a user to ask questions in a virtual reality space and receive answers that take emotion into consideration.
[0716] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0717] Step 1:
[0718] The user wears a virtual reality device and launches a dedicated concierge application, which displays a 3D concierge in the virtual space. The input is the virtual reality device worn by the user and the application that operates it, and the output is the display of the 3D concierge.
[0719] Step 2:
[0720] The user inputs a question by voice or text. For example, they might say, "What's the weather like in Tokyo today?" The input is the user's voice or text, which is sent to the device. The output is the voice data sent to the device.
[0721] Step 3:
[0722] The device records the user's voice input in real time and converts it into text data using voice recognition software. The voice recognition software used is a general voice recognition engine. The input for text conversion is the recorded voice data, and the output is text data. Specifically, the voice signal is frequency analyzed and converted into text based on a language model.
[0723] Step 4:
[0724] The device's emotion engine analyzes the user's voice tone and facial expressions to generate emotion data. The emotion engine uses a general emotion analysis module to analyze voice intonation and facial muscle movements. The input is the user's voice tone and facial expression data, and the output is emotion data.
[0725] Step 5:
[0726] The terminal assembles the converted text question data and the analyzed emotion data into packets and sends them to the server. The input is text data and emotion data, and the output is the data packet sent to the server.
[0727] Step 6:
[0728] The server receives data packets sent from the terminal and analyzes their contents. The input is the received data packet, and the output is the analyzed question data and emotion data. The server checks the consistency of the data and requests a retransmission if there is an inconsistency.
[0729] Step 7:
[0730] The server uses a natural language processing engine to analyze the question text, generate a prompt sentence, and input it to the generative artificial intelligence model. The generative artificial intelligence model uses a general generative AI model. The input is question data and emotional data, and the output is response data that takes the emotion into consideration. For example, the prompt sentence input is, "The user asked, 'I'm really tired from work today. Please tell me how to relax.' This user is feeling tired. Please suggest a relaxation method for the user that takes their emotions into consideration."
[0731] Step 8:
[0732] The server formats the generated response data and sends it to the terminal. The input is the generated response data and the output is the formatted data sent to the terminal.
[0733] Step 9:
[0734] The device renders the received response data as a 3D object, using a 3D graphics engine. The input is the response data received from the server, and the output is a response made by the concierge in the virtual space using natural movements and voice. The concierge naturally tells the user, "To relax, deep breathing and a warm bath are good ways to do it. I also recommend listening to your favorite music."
[0735] (Application example 2)
[0736] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0737] Conventional systems using virtual reality technology have the ability to provide natural responses to user questions, but they are unable to provide individual responses based on the user's emotions. This limits the user's satisfaction and the quality of the experience. The objective of this invention is to provide a more personalized experience by recognizing the user's emotions and providing appropriate responses based on those emotions.
[0738] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0739] In this invention, the server includes a means for incorporating an emotion engine that recognizes and analyzes the user's emotional state, a means for inputting the data into a generative AI model to generate an appropriate response, and a means for displaying the generated response data in three dimensions in the form of a concierge in a virtual reality space, thereby enabling personalized responses according to the user's emotions.
[0740] A "virtual reality device" is a device worn by a user to immerse themselves in a virtual reality space, and refers to devices such as head-mounted displays and VR goggles.
[0741] "Voice or text collection means" refers to an input device such as a microphone or keyboard for capturing voice or text data input by a user in real time.
[0742] "Means for analyzing question data and performing natural language processing" refers to software or algorithms for analyzing collected user questions and semantically analyzing the text using natural language processing technology.
[0743] A "generative artificial intelligence model" is an artificial intelligence model that takes natural language processed data as input and generates appropriate responses, and refers to a model that generates responses based on learned patterns.
[0744] "Means for three-dimensional display" refers to rendering software and display devices for displaying the generated response in three dimensions within a virtual reality space.
[0745] An "emotion engine" refers to hardware or software that analyzes a user's tone of voice, facial expressions, etc., to identify the user's emotional state (e.g., joy, sadness, anger, etc.).
[0746] "Means for generating and displaying a response according to emotions" refers to a series of processes or devices that allow a generative artificial intelligence model to create a response based on the emotional state identified by the emotion engine and to display that response appropriately.
[0747] This invention is a 3D concierge system that combines virtual reality technology and generative artificial intelligence, incorporating an emotion engine that recognizes the user's emotions. With this system, the user puts on a virtual reality device and asks a question. The response to the question is displayed in 3D, and the user can receive an appropriate response according to their emotion.
[0748] The system mainly consists of three main components: users, terminals, and servers.
[0749] User operations
[0750] First, the user puts on a virtual reality device and launches a dedicated concierge application. This refers to a virtual reality device such as a head-mounted display or VR goggles. When the user asks a question by voice or text, the input is collected by the device. For example, the user might say, "Tell me about this product."
[0751] Terminal handling
[0752] The device converts the user's speech into text data using voice recognition software (e.g., Vosk). The device's built-in emotion engine analyzes the user's tone of voice and facial expressions to generate emotion data. The emotion engine uses OpenFace, for example. For example, if a user asks a question in a curious tone, the device identifies the user's curious emotional state.
[0753] Server Processing
[0754] The server receives the question data and emotion data sent from the device and performs natural language processing. The question data is input into a generative artificial intelligence model (such as OpenAI's GPT-3.5 Turbo) to generate a response based on the user's emotion. For example, if a user asks out of curiosity, "Tell me about this product," the server generates a response such as, "This product is new this season, made of silk, and is characterized by its extremely soft feel."
[0755] Terminal Processing (cont.)
[0756] The device receives the response data sent from the server and displays it in 3D as a concierge in the virtual reality space. This display is rendered using a VR library (e.g., VRRenderer). It is also possible to respond by voice using pyttsx3 or similar.
[0757] Prompt Sentence Examples
[0758] For example, the following prompt statement is generated:
[0759] The user was curious and asked: "Tell me about this product."
[0760] As described above, the present invention aims to enable users to ask questions in a virtual reality space and receive personalized responses to those questions based on their emotions, thereby providing a richer experience for the user.
[0761] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0762] Step 1: User interaction
[0763] The user puts on the virtual reality device and launches a dedicated concierge application. A concierge appears in the virtual space, and the user asks a question by voice or text. Input (voice or text) begins. For example, the user might say, "Tell me about this product."
[0764] Step 2: Collecting audio and converting it to text
[0765] The device records the user's speech with a microphone and converts it into text data using voice recognition software (e.g., Vosk). In the process of converting voice data (input) into text data (output), noise removal and voice analysis are performed. Specifically, the recorded voice data is converted into text data such as "Tell me about this product."
[0766] Step 3: Recognize emotions
[0767] The device's built-in emotion engine (e.g., OpenFace) captures the user's voice tone and facial expressions with a camera and analyzes the data. It then identifies the type of emotion and generates emotion data (e.g., curious). The emotion data (output) is packaged into a packet of question data (text data).
[0768] Step 4: Send data
[0769] The terminal assembles the converted question data and the analyzed emotion data into packets and sends them to the server. The input (question data and emotion data) is configured as a data packet to be sent to the server.
[0770] Step 5: Receive and analyze question and sentiment data
[0771] The server receives the question data and emotion data sent from the device and analyzes each data. It uses natural language processing (NLP) to analyze the question data and integrate the emotion data. It then outputs a prompt sentence generated by analyzing the data.
[0772] Step 6: Generate a response
[0773] The server inputs the prompt into a generative artificial intelligence model (e.g., GPT-3.5 Turbo) to generate an appropriate response. In this scenario, if the user's emotion of curiosity is identified, a detailed response tailored to that interest is generated. For example, a response such as "This product is new this season, made of silk, and is characterized by its extremely soft feel" may be output.
[0774] Step 7: Format and send response data
[0775] The server formats the generated response data and sends it as a data packet to the terminal. The response data (output) is now properly formatted and ready to be sent.
[0776] Step 8: Receiving response data and displaying it in 3D
[0777] The device receives the response data sent from the server and displays it in 3D as a concierge in the virtual space. Using VR rendering software (e.g., VRRenderer), the concierge responds with natural movements and voice. Specifically, the concierge explains, "This product is new this season, made of silk, and is extremely soft to the touch."
[0778] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0779] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0780] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0781] [Third embodiment]
[0782] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0783] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[0784] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0785] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0786] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0787] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0788] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0789] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0790] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0791] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0792] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0793] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0794] The present invention provides a 3D concierge system that combines VR technology and generative AI. In this system, a user wears a virtual reality device, asks questions by voice or text, and an appropriate response is displayed in 3D. Specific embodiments for implementing the present invention are described below.
[0795] Program processing
[0796] The system consists of three main components: users, terminals, and servers, each of which plays a specific role and works in tandem.
[0797] User operations
[0798] 1. Booting the system
[0799] The user puts on the virtual reality device and launches a dedicated concierge application, at which point a concierge appears in the virtual space.
[0800] 2. Enter your question
[0801] The user asks a question by voice or text, for example, while cooking, "Tell me how to make tomato sauce."
[0802] Questions are entered into the system through the microphone or virtual keyboard of the virtual reality device.
[0803] Terminal handling
[0804] 3. Speech Recognition and Text Conversion
[0805] The user's voice input is recorded in real time by the terminal and converted into text data by voice recognition software.
[0806] The converted text data is packed into packets as question data and sent to the server.
[0807] 4. Receiving response data
[0808] Receives response data sent from the server, checks and converts the format.
[0809] The concierge then formats the responses appropriately so that they can be given to the user using natural movements and voice.
[0810] 5. Rendering 3D Objects
[0811] The device renders the received response data as a three-dimensional object.
[0812] A concierge in the virtual space presents the generated answer to the user visually and audibly.
[0813] Server Processing
[0814] 6. Receiving and analyzing query data
[0815] The server receives the query data sent from the terminal.
[0816] The received data is pre-processed for natural language processing, including tokenization and normalization.
[0817] 7. Answer generation using generative AI
[0818] The natural language processed data is input into a generative artificial intelligence model, which generates an appropriate answer.
[0819] For example, if you input "How to make tomato sauce," the generative AI will generate the answer, "Here's how to make tomato sauce: First, chop the tomatoes. Next, heat the olive oil in a frying pan, add the garlic and fry until fragrant. Then, adjust the flavor with the chopped tomatoes, salt and pepper, and simmer."
[0820] 8. Sending response data
[0821] The generated response data is appropriately formatted and transmitted to the terminal.
[0822] Specific examples
[0823] For example, the specific operations and processing flow when a user asks "Please tell me how to make curry" while cooking will be shown.
[0824] User
[0825] Put on the virtual reality device and say, "Tell me how to make curry."
[0826] Terminal
[0827] The user's speech is recorded and converted into text data using voice recognition software, such as "Please tell me how to make curry."
[0828] The converted text data is sent to the server.
[0829] server
[0830] It receives text data such as "Please tell me how to make curry" and performs natural language processing.
[0831] Using generative AI, the answer is generated: "Here's how to make curry: First, fry the onions. Next, add the meat, then the vegetables and simmer. Finally, add the curry roux and it's done."
[0832] The generated response data is transmitted to the terminal.
[0833] Terminal
[0834] Based on the received response data, it is rendered as a 3D object.
[0835] The concierge in the virtual space responds to the user using natural movements and voice, saying, "Here's how to make curry: First, fry the onions. Next, add the meat, then the vegetables and simmer. Finally, add the curry roux and it's done."
[0836] In this way, users can have a virtual reality experience that feels like they are having a conversation with a real person. The system responds immediately to detailed questions, enabling intuitive and efficient information acquisition.
[0837] The processing flow will be explained below.
[0838] Step 1:
[0839] The user puts on the virtual reality device and presses the start button on the device to launch the dedicated concierge application, which causes a concierge to appear in the virtual space.
[0840] Step 2:
[0841] The user enters a question. The user speaks into the VR device's microphone, such as "What's the weather like in Tokyo today?", or enters the question as text using the virtual keyboard.
[0842] Step 3:
[0843] The device collects voice data. The user's speech is recorded by the VR device's microphone and converted into text data in real time.
[0844] Step 4:
[0845] The device sends the text data to the server. The converted text data is packed into packets and sent to the server via the Internet.
[0846] Step 5:
[0847] The server receives the text data. The server receives and checks the data sent from the terminal.
[0848] Step 6:
[0849] The server analyzes the text data received, performs preprocessing for natural language processing, and tokenizes and normalizes the text data.
[0850] Step 7:
[0851] The server inputs the data into a generative AI model, which then inputs the analyzed text data into the model to generate an appropriate response.
[0852] Step 8:
[0853] The server formats the generated response data, and examines the generated response data and converts it into a format for display within the virtual reality space.
[0854] Step 9:
[0855] The server sends the formatted response data to the terminal, which then packages the response data into packets and sends them over the Internet to the terminal.
[0856] Step 10:
[0857] The terminal receives the response data. The terminal receives the response data sent from the server and checks its format.
[0858] Step 11:
[0859] The device renders the response data as a 3D object, and processes the received data so that the 3D concierge can respond to the user in a virtual space using natural movements and voice.
[0860] Step 12:
[0861] The user confirms the answer from the Moving Concierge. The concierge in the virtual reality space responds to the user, saying, "The weather in Tokyo is sunny today." The user confirms this visually and audibly.
[0862] In this way, each step works in cooperation to realize a system that allows a user to ask a question in a virtual reality space and receive an appropriate response immediately.
[0863] Example 1
[0864] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0865] In conventional virtual reality systems, it has been difficult to obtain natural responses in real time when a user asks a question. In particular, the processing required to provide detailed and appropriate responses to a user's question is complex, resulting in a non-intuitive user experience. Furthermore, conventional systems have a problem in that the quality of the visual and audio representation of the responses is low, resulting in a lack of realism.
[0866] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0867] In this invention, the server includes a means for a user to wear a virtual reality device, a means for collecting voice or text input by the user, a means for analyzing the collected user question data, a means for natural language processing the analyzed data, a means for inputting the processed natural language data into an artificial intelligence model for generating an appropriate response, a means for three-dimensionally displaying the generated response data in the form of an agent in a virtual reality space, and a means for presenting the generated response data in visual and audio formats, thereby enabling the user to have a more realistic interactive experience.
[0868] "User" refers to a person who uses the system.
[0869] "Virtual reality device" refers to a device used to immerse a user in a virtual reality environment, including, but not limited to, a headset and a handheld controller.
[0870] "Voice or text collection means" refers to devices or software that capture user-input voice or text data. Voice input is accomplished via a microphone, and text input is accomplished via a virtual keyboard or voice recognition software.
[0871] "Question data" refers to data including the question entered by the user.
[0872] "Analysis tools" refers to software and algorithms used to understand the collected question data and perform processes such as parsing, tokenization, and normalization.
[0873] "Natural language processing" refers to the technology that enables a machine to understand and appropriately process input human language. Specifically, it includes tokenization, grammatical analysis, and semantic analysis.
[0874] "Generative artificial intelligence model" refers to a machine learning algorithm or artificial intelligence technique that generates an appropriate response based on natural language processed data. Specific examples include generative AI models.
[0875] "Response data" refers to data containing answers generated by an artificial intelligence model.
[0876] "Means for three-dimensional display" refers to devices and software for displaying the generated response data in three-dimensional space, specifically including a 3D engine for visually displaying the response to an agent (concierge) in a virtual reality space.
[0877] "Means for visual and audio presentation" refers to devices and software for presenting the generated response data to the user in visual animation and audio, including, for example, speech synthesis software that converts text into speech and animation software that controls the behavior of an agent.
[0878] This invention relates to a 3D concierge system that combines virtual reality (VR) technology and generative artificial intelligence (AI). Specifically, the system provides a user with a virtual reality device, asks questions by voice or text, and displays appropriate responses in 3D. The system consists of three main components: a user, a terminal, and a server.
[0879] System configuration
[0880] User operations
[0881] The user wears a virtual reality device (such as Oculus Rift or HTC Vive). Through this device, the user can launch a dedicated concierge application and ask questions to a three-dimensional agent that appears in the virtual space. Questions can be asked by voice input or text input. Voice input uses a microphone, and text input uses a virtual keyboard.
[0882] Terminal handling
[0883] Speech recognition and text conversion
[0884] The device records the user's voice input in real time and converts it into text using speech recognition software such as the Google Speech-to-Text API or Microsoft Azure Speech Service. For example, if a user says, "Tell me how to make tomato sauce," this will be converted into text.
[0885] Sending text data
[0886] The converted text data is packaged into packets as question data and sent to the server via the HTTPS protocol.
[0887] Receiving and processing response data
[0888] Receives the response data sent from the server, checks and converts the format. Specifically, it converts text data into audio data and prepares it into data for setting the motion of a 3D agent.
[0889] Rendering 3D objects
[0890] The device uses a 3D engine such as Unity or Unreal Engine to display the virtual space, and the 3D agent uses natural movements (such as hand gestures and facial expressions) to present the generated responses visually and audibly.
[0891] Server Processing
[0892] Receiving and analyzing question data
[0893] The server receives the query data sent from the terminal. This data is received through a web server (such as Apache or Nginx). The received data is analyzed by a server-side application written in Python or Java. Natural language processing such as tokenization and normalization is performed in the initial analysis stage.
[0894] Answer generation using generative AI
[0895] The analyzed data is input into a generative AI model (e.g., OpenAI GPT-3). The generative AI model generates an appropriate response based on a prompt (e.g., "Please tell me how to make tomato sauce.") For example, it might generate an answer like, "Here's how to make tomato sauce: First, chop the tomatoes. Next, heat olive oil in a frying pan, add the garlic, and fry until fragrant. Next, add the chopped tomatoes, salt, and pepper to taste, and simmer."
[0896] Sending response data
[0897] The generated response data is formatted in a format such as JSON and sent to the terminal via the HTTPS protocol.
[0898] Specific examples
[0899] For example, the specific operations and processing flow when a user asks "Please tell me how to make curry" while cooking will be described.
[0900] 1. Users
[0901] Put on the VR device and say, "Please tell me how to make curry."
[0902] 2. Terminal
[0903] The user's speech is recorded and converted into text data using voice recognition software, such as "Please tell me how to make curry."
[0904] The converted text data is sent to the server.
[0905] 3. Server
[0906] It receives text data such as "Please tell me how to make curry" and performs natural language processing.
[0907] Using generative AI, the answer is generated: "Here's how to make curry: First, fry the onions. Next, add the meat, then the vegetables and simmer. Finally, add the curry roux and it's done."
[0908] The generated response data is transmitted to the terminal.
[0909] 4. Terminal
[0910] Based on the received response data, it is rendered as a 3D object.
[0911] The agent in the virtual space responds to the user using natural movements and voice, saying, "Here's how to make curry: First, fry the onions. Next, add the meat, then the vegetables and simmer. Finally, add the curry roux and it's done."
[0912] In this way, users can have a virtual reality experience that feels like they are having a conversation with a real person. The system responds immediately to detailed questions, enabling intuitive and efficient information acquisition.
[0913] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0914] Step 1:
[0915] The user puts on the virtual reality device and launches a dedicated concierge application. The input is the user putting on the virtual reality device and launching the application, and the output is the agent being ready to appear in the virtual space. Specifically, the user puts on the headset and controllers and clicks on the application icon to launch it.
[0916] Step 2:
[0917] The user asks a question by voice or text. The input is the user's speech or text input, and the output is the data of the user's question. Specifically, the user speaks "Tell me how to make tomato sauce" or enters similar text on a virtual keyboard.
[0918] Step 3:
[0919] The device records the user's voice input in real time and converts it into text data using the Google Speech-to-Text API or Microsoft Azure Speech Service. The input is the user's voice data and the output is text data. Specifically, the voice recognition software analyzes the voice data and converts it into text such as "Please tell me how to make tomato sauce."
[0920] Step 4:
[0921] The terminal assembles the converted text data into packets as question data and sends them to the server via the HTTPS protocol. The input is the text data converted from the voice, and the output is the question data sent to the server. Specifically, the terminal encodes and sends the data to the server.
[0922] Step 5:
[0923] The server receives the query data sent from the terminal. The input is the query data sent from the terminal, and the output is raw data for analysis. In concrete terms, the web server (e.g., Apache or Nginx) passes the received data to the application server.
[0924] Step 6:
[0925] The server analyzes the received question data using natural language processing (NLP). The input is the received question data, and the output is tokenized and normalized data. Specifically, an NLP engine implemented in Python, Java, or other languages tokenizes the data and performs grammatical and semantic analysis.
[0926] Step 7:
[0927] The server inputs the preprocessed data into a generative artificial intelligence model (e.g., OpenAI GPT-3) to generate an appropriate answer. The input is natural language processed data, and the output is the generated response data. Specifically, it inputs a prompt sentence (e.g., "Please tell me how to make tomato sauce.") into GPT-3 and obtains the generated response (e.g., "Here's how to make tomato sauce. First, chop the tomatoes...").
[0928] Step 8:
[0929] The server formats the generated response data in a format such as JSON and sends it to the terminal via the HTTPS protocol. The input is the generated response data, and the output is the response data sent to the terminal. Specifically, the server formats, encodes, and sends the data.
[0930] Step 9:
[0931] The terminal receives the response data sent from the server and checks and converts the format. The input is the response data sent from the server, and the output is renderable data. Specific operations include decoding the received data and generating speech synthesis and action generation data for the agent.
[0932] Step 10:
[0933] The device uses a 3D engine such as Unity or Unreal Engine to display the agent in a virtual space. The input is renderable data, and the output is the agent's response, presented visually and audibly. Specifically, the agent explains, with hand gestures and facial expressions, "Here's how to make tomato sauce. First, chop the tomatoes..."
[0934] (Application example 1)
[0935] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0936] In recent years, there have been many attempts to utilize technology to improve the customer experience in brick-and-mortar stores. However, current systems often make it difficult for users to quickly obtain detailed product information, especially when visiting a store for the first time or in an unfamiliar product category. Furthermore, when customers are hesitant to directly interact with store staff or when the store is crowded, quickly obtaining information becomes even more difficult. Therefore, there is a need for a system that allows users to easily obtain product information by voice and presents that information visually and audibly.
[0937] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0938] In this invention, the server includes: means for a user to wear a virtual reality device; means for collecting voice or text input by the user; means for analyzing the collected user question data and performing natural language processing; means for inputting the analyzed data into a generative artificial intelligence model and generating an appropriate response; means for three-dimensionally displaying the generated response data in the form of a concierge in a virtual reality space; means for forming a user interface for providing product information in a physical store; and means for generating answers to the user's voice questions to provide product information and presenting them to the user by voice, thereby enabling users to quickly and easily obtain product information in a physical store and receive that information visually and audibly.
[0939] A "user" is a person who uses the system to obtain information.
[0940] A "virtual reality device" is a device worn by a user to access a virtual reality space that includes vision and sound.
[0941] A "voice or text collection means" is any device or software that collects user input in the form of voice or text.
[0942] "Question data" is data including a question entered by a user.
[0943] "Natural language processing" is the process of analyzing collected question data and converting it into a format that a computer can understand.
[0944] A "generative artificial intelligence model" is an artificial intelligence model that inputs natural language processed data and generates an appropriate response.
[0945] "Response data" is data that includes an answer generated by a generative artificial intelligence model.
[0946] The "means for three-dimensional display" refers to a device or software for displaying the response data in the form of a three-dimensional concierge in a virtual reality space.
[0947] A "brick and mortar store" is a retail outlet that exists in a physical location.
[0948] "User interface" means an interface through which a user inputs information and interacts with the concierge system.
[0949] "Product information" is detailed information about the products being sold.
[0950] The "means for providing by voice" refers to a device or software for presenting the generated response data to the user in voice format.
[0951] The present invention provides a system that allows users to acquire product information in a physical store using a virtual reality device. The system displays appropriate responses to questions in the form of a three-dimensional concierge and presents them to the user via audio.
[0952] System configuration
[0953] The main components of the system are:
[0954] 1. Users
[0955] 2. Terminal
[0956] 3. Server
[0957] User operations
[0958] The user puts on the virtual reality device and asks a question about a product in the store by voice, for example, "What are the ingredients in this product?" This voice is transmitted to the terminal through the microphone of the virtual reality device.
[0959] Terminal handling
[0960] The device processes the user's voice input in the following steps:
[0961] 1. Use speech recognition software (e.g., SpeechRecognition library, Google Text-to-Speech) to convert the audio data into text data.
[0962] 2. The converted text data is packed into packets and sent to the server.
[0963] 3. Receive the response data from the server and format it appropriately.
[0964] 4. Render the generated answer as a 3D object and a virtual concierge will present the answer visually and audibly.
[0965] Server Processing
[0966] The server processes the query data in the following steps:
[0967] 1. Receive text data sent from the device.
[0968] 2. Perform natural language processing (NLP) and analyze the data.
[0969] 3. Use a generative AI model (e.g., GPT-3) to generate an appropriate answer. For example, if a user asks, "What are the ingredients in this product?", the generative AI will generate an answer such as, "This product's ingredients include water, glycerin, parabens, vitamin E, etc."
[0970] 4. The generated response data is sent to the terminal.
[0971] Hardware and software used
[0972] Hardware: Virtual reality devices (smart glasses, head-mounted displays, etc.), microphones, smartphones
[0973] Software: SpeechRecognition Library, Google Text-to-Speech (gTTS), OpenAI API
[0974] Specific examples
[0975] For example, if a user is browsing products in a physical store and asks through smart glasses, "What are the ingredients in this product?", the voice is sent to the device via a microphone. The device converts the voice into text, and this text data is sent to the server. The server uses a generative AI model to generate an answer, returning, for example, "The ingredients in this product include water, glycerin, parabens, vitamin E, etc." to the device. The device then provides this answer to the user as a 3D display and audio.
[0976] Prompt Sentence Examples
[0977] An example of a prompt sentence is, "A user asks a question about a product in a store. Question: 'What are the ingredients in this product?' Please provide an easy-to-understand answer to this question." By inputting this prompt sentence into a generative AI model, an appropriate answer will be generated.
[0978] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0979] Step 1:
[0980] The user wears a virtual reality device and asks a question about a product by voice. This voice is collected by the built-in microphone of the virtual reality device. The input is the user's voice, and the output is voice data.
[0981] Step 2:
[0982] The device uses speech recognition software (e.g., SpeechRecognition library, Google Text-to-Speech) to convert the collected voice data into text data. Here, speech is converted into text, such as "Please tell me the ingredients of this product." The input is voice data, and the output is text data.
[0983] Step 3:
[0984] The terminal packs the converted text data into packets and sends them to the server. The input is text data and the output is packet data.
[0985] Step 4:
[0986] The server receives packet data sent from the terminal and analyzes the data using natural language processing (NLP), where text data is tokenized and normalized. The input is packet data, and the output is analyzed data.
[0987] Step 5:
[0988] The server inputs the analyzed data into a generative artificial intelligence model (e.g., GPT-3) to generate an appropriate answer. For example, in response to the question, "What are the ingredients of this product?", an answer such as, "This product contains water, glycerin, parabens, vitamin E, etc." is generated. The input is the analyzed data, and the output is the generated answer data.
[0989] Step 6:
[0990] The server sends the generated response data to the terminal, where the input is the generated response data and the output is the reformatted response data.
[0991] Step 7:
[0992] The device then formats the answer data received from the server and renders it as a 3D object, where a concierge in a virtual reality space presents the generated answer visually and audibly. The input is the reformatted answer data, and the output is 3D visual and audio data.
[0993] Step 8:
[0994] The user receives visual and audio responses from the concierge in the virtual reality space and confirms detailed product information. The input is 3D visual and audio data, and the output is the user's understanding and acquisition of product information.
[0995] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0996] The present invention combines a 3D concierge system that combines VR technology and generative AI with an emotion engine that recognizes the user's emotions. This system allows the user to wear a virtual reality device, ask questions by voice or text, and display responses in 3D, while simultaneously generating appropriate responses based on the user's emotions. Specific embodiments for implementing the present invention are described below.
[0997] Program processing
[0998] The system consists of three main components: the user, the device, and the server, among which an emotion engine has been newly incorporated. Each plays a specific role and works in conjunction with each other.
[0999] User operations
[1000] 1. Booting the system
[1001] The user puts on the virtual reality device and launches a dedicated concierge application, at which point a concierge appears in the virtual space.
[1002] 2. Enter your question
[1003] The user asks a question by voice or text. For example, the user says, "What's the weather like in Tokyo today?"
[1004] Terminal handling
[1005] 3. Speech Recognition and Text Conversion
[1006] The user's voice input is recorded in real time by the terminal and converted into text data by voice recognition software.
[1007] 4. Emotion recognition
[1008] The device's built-in emotion engine analyzes the user's voice tone and facial expressions to generate emotion data, which identifies the user's emotional state (e.g., joy, sadness, anger, etc.).
[1009] 5. Sending Question Data and Emotion Data
[1010] The converted text question data and the analyzed emotion data are packaged into packets and sent to the server.
[1011] Server Processing
[1012] 6. Analysis of received data
[1013] The server receives and confirms the question data and emotion data sent from the terminal.
[1014] 7. Response Generation Using Generative AI
[1015] The question data is processed using natural language processing and input into a generative AI model. Emotional data is then incorporated to generate a response that reflects the user's emotions.
[1016] For example, if a user asks in a sad tone, "How's the weather in Tokyo today?", the generative AI will generate an answer that takes emotions into consideration, such as, "The weather in Tokyo is sunny today. It's nice, so why not go outside for a bit to change your mood?"
[1017] 8. Response Data Formatting and Transmission
[1018] The generated response data is formatted and sent to the terminal.
[1019] Terminal Processing (cont.)
[1020] 9. Receiving and displaying response data
[1021] The device receives the response data sent from the server and renders it as a three-dimensional object.
[1022] The concierge in the virtual space responds to the user with natural movements and voice, and also responds according to the user's emotions.
[1023] Specific examples
[1024] For example, the specific operations and processing flow when a user asks, "I'm really tired from work today. Please tell me how to relax" in a stressful daily life is shown below.
[1025] User
[1026] Put on the virtual reality device and say, "I'm really tired from work today. Please tell me how to relax."
[1027] Terminal
[1028] The user's speech is recorded and converted into text using speech recognition software.
[1029] The emotion engine detects fatigue and stress from the user's voice tone and generates emotion data.
[1030] The question data and emotion data are packaged into a packet and sent to the server.
[1031] server
[1032] Receives question data and emotion data and performs natural language processing.
[1033] Using generative AI, in response to the question, "I'm really tired from work today. Please tell me how to relax," the system generates a response that takes emotions into consideration, such as, "To relax, why not try taking some slow, deep breaths or taking a warm bath? I also recommend listening to your favorite music."
[1034] The generated response data is formatted and sent to the terminal.
[1035] Terminal
[1036] Based on the received response data, the concierge in the virtual space responds to the user with natural movements and voice.
[1037] The concierge tells the user, "To relax, why not try taking some slow, deep breaths or taking a warm bath? We also recommend listening to your favorite music."
[1038] In this way, by having each step work in conjunction with one another, a system is realized that allows a user to ask questions in a virtual reality space and receive answers that correspond to their emotions.
[1039] The processing flow will be explained below.
[1040] Step 1:
[1041] The user puts on the virtual reality device, turns on the device, and launches the dedicated concierge application, which displays the concierge in the virtual space.
[1042] Step 2:
[1043] The user types a question. The user can either speak aloud, such as "Work today is really tiring. What are some ways to relax?" or type the question as text using the virtual keyboard.
[1044] Step 3:
[1045] The device collects voice data: the user's speech is recorded by the VR device's microphone and sent to the voice recognition software.
[1046] Step 4:
[1047] The device converts the voice data into text data. Voice recognition software analyzes the voice data in real time and generates text data such as, "I'm really tired from work today. Please tell me how to relax."
[1048] Step 5:
[1049] The device collects the user's voice tone and facial expressions, and the emotion engine analyzes the user's voice tone and facial expressions to identify the user's emotional state.
[1050] Step 6:
[1051] The terminal generates emotion data, and based on the collected emotion information, generates emotion data indicating that the user is tired.
[1052] Step 7:
[1053] The device sends the question data and emotion data to the server, which then assembles the converted text data and generated emotion data into packets and sends them to the server via the Internet.
[1054] Step 8:
[1055] The server receives the question data and emotion data, analyzes the received data, and prepares it for processing.
[1056] Step 9:
[1057] The server performs natural language processing, tokenizing and normalizing the question data and converting it into a format that can be input to a generative AI model.
[1058] Step 10:
[1059] The server generates responses using a generative artificial intelligence model. It inputs question data and emotion data into the model and generates an appropriate response.
[1060] Step 11:
[1061] The server formats the response data, examining the generated response and formatting it in a format that can be displayed within the virtual reality space.
[1062] Step 12:
[1063] The server sends the formatted response data to the terminal, which then packages the response data into packets and sends them to the terminal over the Internet.
[1064] Step 13:
[1065] The terminal receives the response data. The terminal receives the response data sent from the server and checks its format.
[1066] Step 14:
[1067] The device renders the response data as a 3D object, and processes the received data so that the concierge can respond in a virtual space using natural movements and voice.
[1068] Step 15:
[1069] The user confirms the concierge's response by visually and audibly hearing the concierge in the virtual reality space reply to the user, "To relax, why not try some slow, deep breathing or a warm bath? We also recommend listening to your favorite music."
[1070] In this way, each step works in conjunction to realize a system that allows a user to ask questions in a virtual reality space and receive appropriate answers based on their emotions.
[1071] Example 2
[1072] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1073] Conventional virtual reality concierge systems can generate responses to questions entered by users, but they are unable to provide appropriate responses that take the user's emotions into consideration. This limits the user experience and prevents the user from receiving the personalized responses they desire.
[1074] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1075] In this invention, the server includes a means for allowing a user to wear a virtual reality device, a means for collecting voice or text input by the user, a built-in emotion engine for generating emotion data from the collected voice data, and a means for inputting the analyzed data into a generative artificial intelligence model to generate an appropriate response, thereby enabling a response that takes into account the user's emotions.
[1076] A "virtual reality device" is a device that allows a user to experience a virtual space through sight and sound.
[1077] "Voice or text collection means" refers to hardware and software components for collecting user-spoken voice or text information.
[1078] "Means for analyzing question data and performing natural language processing" refers to technology for analyzing the content of questions entered by users and processing them in an appropriate manner.
[1079] A "generative artificial intelligence model" is a set of artificial intelligence algorithms that perform natural language responses and other generative tasks based on given data.
[1080] "Means for three-dimensionally displaying response data in the form of a concierge in a virtual reality space" refers to a method that utilizes three-dimensional graphics technology to communicate the generated response to the user visually and audibly.
[1081] The "emotion engine" is an analysis device and a group of algorithms for extracting emotional data from a user's tone of voice and facial expressions.
[1082] The "means for generating a response that takes emotion into consideration" refers to a process and technology for generating a response that is appropriate to a user's emotion based on the user's emotion data.
[1083] This invention is a 3D concierge system that combines virtual reality technology and generative artificial intelligence, and further incorporates an emotion engine that recognizes the user's emotions, thereby providing responses that take the user's emotions into consideration. The system allows the user to wear a virtual reality device, input questions by voice or text, and displays the responses in 3D while simultaneously generating appropriate responses according to the user's emotions.
[1084] This system consists of three main components: the user, the terminal, and the server, each of which plays a specific role and works in conjunction with each other. The emotion engine has been newly incorporated, and the operation of each component is shown below.
[1085] User operations
[1086] First, the user puts on a virtual reality device (e.g., a "VR headset" as it is commonly called) and launches a dedicated concierge application. At this point, a three-dimensional concierge appears in the virtual space. The user asks questions through a voice recognition and text input interface. For example, when the user says, "How is the weather in Tokyo today?", the voice is sent to the device.
[1087] Terminal handling
[1088] The device records the user's voice input in real time and converts it into text data using voice recognition software (e.g., generically called a "voice recognition engine"). The device also has a built-in emotion engine (e.g., generically called an "emotion analysis module") that recognizes the user's emotions and analyzes voice tone and facial expression data to generate emotion data. For example, if a user says in a tired voice, "I'm really tired from work today. Please tell me how to relax," the emotion will be identified as "fatigue." The converted text question data and analyzed emotion data are packaged into a data packet and sent to the server.
[1089] Server Processing
[1090] The server receives the data packet sent from the device and analyzes its contents. It uses a natural language processing (NLP) module (e.g., a generic term "natural language processing engine") to analyze the question text and generate a prompt sentence, which is input to a generative artificial intelligence model (e.g., a generic term "generative AI model"). The generated prompt sentence also includes emotional data. For example, the following content is input to the generative AI model: "A user asks, 'I'm really tired from work today. Please tell me how to relax.' This user is feeling tired. Please suggest relaxation methods to the user that take their emotions into consideration." The generative AI model then generates an appropriate response to the question, such as, "Deep breathing and a warm bath are good ways to relax. I also recommend listening to your favorite music."
[1091] The generated response data is formatted and sent to the device. The device then renders the received response data as a 3D object. The concierge responds to the user in the virtual space using natural movements and voice. For example, if the user asks, "I'm really tired from work today. What are some ways to relax?" the concierge will tell the user, "Deep breathing and a warm bath are good ways to relax. I also recommend listening to your favorite music."
[1092] In this way, a system is realized that allows a user to ask questions in a virtual reality space and receive answers that take emotion into consideration.
[1093] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1094] Step 1:
[1095] The user wears a virtual reality device and launches a dedicated concierge application, which displays a 3D concierge in the virtual space. The input is the virtual reality device worn by the user and the application that operates it, and the output is the display of the 3D concierge.
[1096] Step 2:
[1097] The user inputs a question by voice or text. For example, they might say, "What's the weather like in Tokyo today?" The input is the user's voice or text, which is sent to the device. The output is the voice data sent to the device.
[1098] Step 3:
[1099] The device records the user's voice input in real time and converts it into text data using voice recognition software. The voice recognition software used is a general voice recognition engine. The input for text conversion is the recorded voice data, and the output is text data. Specifically, the voice signal is frequency analyzed and converted into text based on a language model.
[1100] Step 4:
[1101] The device's emotion engine analyzes the user's voice tone and facial expressions to generate emotion data. The emotion engine uses a general emotion analysis module to analyze voice intonation and facial muscle movements. The input is the user's voice tone and facial expression data, and the output is emotion data.
[1102] Step 5:
[1103] The terminal assembles the converted text question data and the analyzed emotion data into packets and sends them to the server. The input is text data and emotion data, and the output is the data packet sent to the server.
[1104] Step 6:
[1105] The server receives data packets sent from the terminal and analyzes their contents. The input is the received data packet, and the output is the analyzed question data and emotion data. The server checks the consistency of the data and requests a retransmission if there is an inconsistency.
[1106] Step 7:
[1107] The server uses a natural language processing engine to analyze the question text, generate a prompt sentence, and input it to the generative artificial intelligence model. The generative artificial intelligence model uses a general generative AI model. The input is question data and emotional data, and the output is response data that takes the emotion into consideration. For example, the prompt sentence input is, "The user asked, 'I'm really tired from work today. Please tell me how to relax.' This user is feeling tired. Please suggest a relaxation method for the user that takes their emotions into consideration."
[1108] Step 8:
[1109] The server formats the generated response data and sends it to the terminal. The input is the generated response data and the output is the formatted data sent to the terminal.
[1110] Step 9:
[1111] The device renders the received response data as a 3D object, using a 3D graphics engine. The input is the response data received from the server, and the output is a response made by the concierge in the virtual space using natural movements and voice. The concierge naturally tells the user, "To relax, deep breathing and a warm bath are good ways to do it. I also recommend listening to your favorite music."
[1112] (Application example 2)
[1113] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1114] Conventional systems using virtual reality technology have the ability to provide natural responses to user questions, but they are unable to provide individual responses based on the user's emotions. This limits the user's satisfaction and the quality of the experience. The objective of this invention is to provide a more personalized experience by recognizing the user's emotions and providing appropriate responses based on those emotions.
[1115] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1116] In this invention, the server includes a means for incorporating an emotion engine that recognizes and analyzes the user's emotional state, a means for inputting the data into a generative AI model to generate an appropriate response, and a means for displaying the generated response data in three dimensions in the form of a concierge in a virtual reality space, thereby enabling personalized responses according to the user's emotions.
[1117] A "virtual reality device" is a device worn by a user to immerse themselves in a virtual reality space, and refers to devices such as head-mounted displays and VR goggles.
[1118] "Voice or text collection means" refers to an input device such as a microphone or keyboard for capturing voice or text data input by a user in real time.
[1119] "Means for analyzing question data and performing natural language processing" refers to software or algorithms for analyzing collected user questions and semantically analyzing the text using natural language processing technology.
[1120] A "generative artificial intelligence model" is an artificial intelligence model that takes natural language processed data as input and generates appropriate responses, and refers to a model that generates responses based on learned patterns.
[1121] "Means for three-dimensional display" refers to rendering software and display devices for displaying the generated response in three dimensions within a virtual reality space.
[1122] An "emotion engine" refers to hardware or software that analyzes a user's tone of voice, facial expressions, etc., to identify the user's emotional state (e.g., joy, sadness, anger, etc.).
[1123] "Means for generating and displaying a response according to emotions" refers to a series of processes or devices that allow a generative artificial intelligence model to create a response based on the emotional state identified by the emotion engine and to display that response appropriately.
[1124] This invention is a 3D concierge system that combines virtual reality technology and generative artificial intelligence, incorporating an emotion engine that recognizes the user's emotions. With this system, the user puts on a virtual reality device and asks a question. The response to the question is displayed in 3D, and the user can receive an appropriate response according to their emotion.
[1125] The system mainly consists of three main components: users, terminals, and servers.
[1126] User operations
[1127] First, the user puts on a virtual reality device and launches a dedicated concierge application. This refers to a virtual reality device such as a head-mounted display or VR goggles. When the user asks a question by voice or text, the input is collected by the device. For example, the user might say, "Tell me about this product."
[1128] Terminal handling
[1129] The device converts the user's speech into text data using voice recognition software (e.g., Vosk). The device's built-in emotion engine analyzes the user's tone of voice and facial expressions to generate emotion data. The emotion engine uses OpenFace, for example. For example, if a user asks a question in a curious tone, the device identifies the user's curious emotional state.
[1130] Server Processing
[1131] The server receives the question data and emotion data sent from the device and performs natural language processing. The question data is input into a generative artificial intelligence model (such as OpenAI's GPT-3.5 Turbo) to generate a response based on the user's emotion. For example, if a user asks out of curiosity, "Tell me about this product," the server generates a response such as, "This product is new this season, made of silk, and is characterized by its extremely soft feel."
[1132] Terminal Processing (cont.)
[1133] The device receives the response data sent from the server and displays it in 3D as a concierge in the virtual reality space. This display is rendered using a VR library (e.g., VRRenderer). It is also possible to respond by voice using pyttsx3 or similar.
[1134] Prompt Sentence Examples
[1135] For example, the following prompt statement is generated:
[1136] The user was curious and asked: "Tell me about this product."
[1137] As described above, the present invention aims to enable users to ask questions in a virtual reality space and receive personalized responses to those questions based on their emotions, thereby providing a richer experience for the user.
[1138] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1139] Step 1: User interaction
[1140] The user puts on the virtual reality device and launches a dedicated concierge application. A concierge appears in the virtual space, and the user asks a question by voice or text. Input (voice or text) begins. For example, the user might say, "Tell me about this product."
[1141] Step 2: Collecting audio and converting it to text
[1142] The device records the user's speech with a microphone and converts it into text data using voice recognition software (e.g., Vosk). In the process of converting voice data (input) into text data (output), noise removal and voice analysis are performed. Specifically, the recorded voice data is converted into text data such as "Tell me about this product."
[1143] Step 3: Recognize emotions
[1144] The device's built-in emotion engine (e.g., OpenFace) captures the user's voice tone and facial expressions with a camera and analyzes the data. It then identifies the type of emotion and generates emotion data (e.g., curious). The emotion data (output) is packaged into a packet of question data (text data).
[1145] Step 4: Send data
[1146] The terminal assembles the converted question data and the analyzed emotion data into packets and sends them to the server. The input (question data and emotion data) is configured as a data packet to be sent to the server.
[1147] Step 5: Receive and analyze question and sentiment data
[1148] The server receives the question data and emotion data sent from the device and analyzes each data. It uses natural language processing (NLP) to analyze the question data and integrate the emotion data. It then outputs a prompt sentence generated by analyzing the data.
[1149] Step 6: Generate a response
[1150] The server inputs the prompt into a generative artificial intelligence model (e.g., GPT-3.5 Turbo) to generate an appropriate response. In this scenario, if the user's emotion of curiosity is identified, a detailed response tailored to that interest is generated. For example, a response such as "This product is new this season, made of silk, and is characterized by its extremely soft feel" may be output.
[1151] Step 7: Format and send response data
[1152] The server formats the generated response data and sends it as a data packet to the terminal. The response data (output) is now properly formatted and ready to be sent.
[1153] Step 8: Receiving response data and displaying it in 3D
[1154] The device receives the response data sent from the server and displays it in 3D as a concierge in the virtual space. Using VR rendering software (e.g., VRRenderer), the concierge responds with natural movements and voice. Specifically, the concierge explains, "This product is new this season, made of silk, and is extremely soft to the touch."
[1155] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1156] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1157] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1158] [Fourth embodiment]
[1159] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1160] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1161] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1162] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1163] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1164] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1165] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1166] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1167] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1168] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1169] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1170] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1171] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1172] The present invention provides a 3D concierge system that combines VR technology and generative AI. In this system, a user wears a virtual reality device, asks questions by voice or text, and an appropriate response is displayed in 3D. Specific embodiments for implementing the present invention are described below.
[1173] Program processing
[1174] The system consists of three main components: users, terminals, and servers, each of which plays a specific role and works in tandem.
[1175] User operations
[1176] 1. Booting the system
[1177] The user puts on the virtual reality device and launches a dedicated concierge application, at which point a concierge appears in the virtual space.
[1178] 2. Enter your question
[1179] The user asks a question by voice or text, for example, while cooking, "Tell me how to make tomato sauce."
[1180] Questions are entered into the system through the microphone or virtual keyboard of the virtual reality device.
[1181] Terminal handling
[1182] 3. Speech Recognition and Text Conversion
[1183] The user's voice input is recorded in real time by the terminal and converted into text data by voice recognition software.
[1184] The converted text data is packed into packets as question data and sent to the server.
[1185] 4. Receiving response data
[1186] Receives response data sent from the server, checks and converts the format.
[1187] The concierge then formats the responses appropriately so that they can be given to the user using natural movements and voice.
[1188] 5. Rendering 3D Objects
[1189] The device renders the received response data as a three-dimensional object.
[1190] A concierge in the virtual space presents the generated answer to the user visually and audibly.
[1191] Server Processing
[1192] 6. Receiving and analyzing query data
[1193] The server receives the query data sent from the terminal.
[1194] The received data is pre-processed for natural language processing, including tokenization and normalization.
[1195] 7. Answer generation using generative AI
[1196] The natural language processed data is input into a generative artificial intelligence model, which generates an appropriate answer.
[1197] For example, if you input "How to make tomato sauce," the generative AI will generate the answer, "Here's how to make tomato sauce: First, chop the tomatoes. Next, heat the olive oil in a frying pan, add the garlic and fry until fragrant. Then, adjust the flavor with the chopped tomatoes, salt and pepper, and simmer."
[1198] 8. Sending response data
[1199] The generated response data is appropriately formatted and transmitted to the terminal.
[1200] Specific examples
[1201] For example, the specific operations and processing flow when a user asks "Please tell me how to make curry" while cooking will be shown.
[1202] User
[1203] Put on the virtual reality device and say, "Tell me how to make curry."
[1204] Terminal
[1205] The user's speech is recorded and converted into text data using voice recognition software, such as "Please tell me how to make curry."
[1206] The converted text data is sent to the server.
[1207] server
[1208] It receives text data such as "Please tell me how to make curry" and performs natural language processing.
[1209] Using generative AI, the answer is generated: "Here's how to make curry: First, fry the onions. Next, add the meat, then the vegetables and simmer. Finally, add the curry roux and it's done."
[1210] The generated response data is transmitted to the terminal.
[1211] Terminal
[1212] Based on the received response data, it is rendered as a 3D object.
[1213] The concierge in the virtual space responds to the user using natural movements and voice, saying, "Here's how to make curry: First, fry the onions. Next, add the meat, then the vegetables and simmer. Finally, add the curry roux and it's done."
[1214] In this way, users can have a virtual reality experience that feels like they are having a conversation with a real person. The system responds immediately to detailed questions, enabling intuitive and efficient information acquisition.
[1215] The processing flow will be explained below.
[1216] Step 1:
[1217] The user puts on the virtual reality device and presses the start button on the device to launch the dedicated concierge application, which causes a concierge to appear in the virtual space.
[1218] Step 2:
[1219] The user enters a question. The user speaks into the VR device's microphone, such as "What's the weather like in Tokyo today?", or enters the question as text using the virtual keyboard.
[1220] Step 3:
[1221] The device collects voice data. The user's speech is recorded by the VR device's microphone and converted into text data in real time.
[1222] Step 4:
[1223] The device sends the text data to the server. The converted text data is packed into packets and sent to the server via the Internet.
[1224] Step 5:
[1225] The server receives the text data. The server receives and checks the data sent from the terminal.
[1226] Step 6:
[1227] The server analyzes the text data received, performs preprocessing for natural language processing, and tokenizes and normalizes the text data.
[1228] Step 7:
[1229] The server inputs the data into a generative AI model, which then inputs the analyzed text data into the model to generate an appropriate response.
[1230] Step 8:
[1231] The server formats the generated response data, and validates the generated response data and converts it into a format for display within the virtual reality space.
[1232] Step 9:
[1233] The server sends the formatted response data to the terminal, which then packages the response data into packets and sends them over the Internet to the terminal.
[1234] Step 10:
[1235] The terminal receives the response data. The terminal receives the response data sent from the server and checks its format.
[1236] Step 11:
[1237] The device renders the response data as a 3D object, and processes the received data so that the 3D concierge can respond to the user in a virtual space using natural movements and voice.
[1238] Step 12:
[1239] The user confirms the answer from the Moving Concierge. The concierge in the virtual reality space responds to the user, saying, "The weather in Tokyo is sunny today." The user confirms this visually and audibly.
[1240] In this way, each step works in cooperation to realize a system that allows a user to ask a question in a virtual reality space and receive an appropriate response immediately.
[1241] Example 1
[1242] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1243] In conventional virtual reality systems, it has been difficult to obtain natural responses in real time when a user asks a question. In particular, the processing required to provide detailed and appropriate responses to a user's question is complex, resulting in a non-intuitive user experience. Furthermore, conventional systems have a problem in that the quality of the visual and audio representation of the responses is low, resulting in a lack of realism.
[1244] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1245] In this invention, the server includes a means for a user to wear a virtual reality device, a means for collecting voice or text input by the user, a means for analyzing the collected user question data, a means for natural language processing the analyzed data, a means for inputting the processed natural language data into an artificial intelligence model for generating an appropriate response, a means for three-dimensionally displaying the generated response data in the form of an agent in a virtual reality space, and a means for presenting the generated response data in visual and audio formats, thereby enabling the user to have a more realistic interactive experience.
[1246] "User" refers to a person who uses the system.
[1247] "Virtual reality device" refers to a device used to immerse a user in a virtual reality environment, including, but not limited to, a headset and a handheld controller.
[1248] "Voice or text collection means" refers to devices or software that capture user-input voice or text data. Voice input is accomplished via a microphone, and text input is accomplished via a virtual keyboard or voice recognition software.
[1249] "Question data" refers to data including the question entered by the user.
[1250] "Analysis tools" refers to software and algorithms used to understand the collected question data and perform processes such as parsing, tokenization, and normalization.
[1251] "Natural language processing" refers to the technology that enables a machine to understand and appropriately process input human language. Specifically, it includes tokenization, grammatical analysis, and semantic analysis.
[1252] "Generative artificial intelligence model" refers to a machine learning algorithm or artificial intelligence technique that generates an appropriate response based on natural language processed data. Specific examples include generative AI models.
[1253] "Response data" refers to data containing answers generated by an artificial intelligence model.
[1254] "Means for three-dimensional display" refers to devices and software for displaying the generated response data in three-dimensional space, specifically including a 3D engine for visually displaying the response to an agent (concierge) in a virtual reality space.
[1255] "Means for visual and audio presentation" refers to devices and software for presenting the generated response data to the user in visual animation and audio, including, for example, speech synthesis software that converts text into speech and animation software that controls the behavior of an agent.
[1256] This invention relates to a 3D concierge system that combines virtual reality (VR) technology and generative artificial intelligence (AI). Specifically, the system provides a user with a virtual reality device, asks questions by voice or text, and displays appropriate responses in 3D. The system consists of three main components: a user, a terminal, and a server.
[1257] System configuration
[1258] User operations
[1259] The user wears a virtual reality device (such as Oculus Rift or HTC Vive). Through this device, the user can launch a dedicated concierge application and ask questions to a three-dimensional agent that appears in the virtual space. Questions can be asked by voice input or text input. Voice input uses a microphone, and text input uses a virtual keyboard.
[1260] Terminal handling
[1261] Speech recognition and text conversion
[1262] The device records the user's voice input in real time and converts it into text using speech recognition software such as the Google Speech-to-Text API or Microsoft Azure Speech Service. For example, if a user says, "Tell me how to make tomato sauce," this will be converted into text.
[1263] Sending text data
[1264] The converted text data is packaged into packets as question data and sent to the server via the HTTPS protocol.
[1265] Receiving and processing response data
[1266] Receives the response data sent from the server, checks and converts the format. Specifically, it converts text data into audio data and prepares it into data for setting the motion of a 3D agent.
[1267] Rendering 3D objects
[1268] The device uses a 3D engine such as Unity or Unreal Engine to display the virtual space, and the 3D agent uses natural movements (such as hand gestures and facial expressions) to present the generated responses visually and audibly.
[1269] Server Processing
[1270] Receiving and analyzing question data
[1271] The server receives the query data sent from the terminal. This data is received through a web server (such as Apache or Nginx). The received data is analyzed by a server-side application written in Python or Java. Natural language processing such as tokenization and normalization is performed in the initial analysis stage.
[1272] Answer generation using generative AI
[1273] The analyzed data is input into a generative AI model (e.g., OpenAI GPT-3). The generative AI model generates an appropriate response based on a prompt (e.g., "Please tell me how to make tomato sauce.") For example, it might generate an answer like, "Here's how to make tomato sauce: First, chop the tomatoes. Next, heat olive oil in a frying pan, add the garlic, and fry until fragrant. Next, add the chopped tomatoes, salt, and pepper to taste, and simmer."
[1274] Sending response data
[1275] The generated response data is formatted in a format such as JSON and sent to the terminal via the HTTPS protocol.
[1276] Specific examples
[1277] For example, the specific operations and processing flow when a user asks "Please tell me how to make curry" while cooking will be described.
[1278] 1. Users
[1279] Put on the VR device and say, "Please tell me how to make curry."
[1280] 2. Terminal
[1281] The user's speech is recorded and converted into text data using voice recognition software, such as "Please tell me how to make curry."
[1282] The converted text data is sent to the server.
[1283] 3. Server
[1284] It receives text data such as "Please tell me how to make curry" and performs natural language processing.
[1285] Using generative AI, the answer is generated: "Here's how to make curry: First, fry the onions. Next, add the meat, then the vegetables and simmer. Finally, add the curry roux and it's done."
[1286] The generated response data is transmitted to the terminal.
[1287] 4. Terminal
[1288] Based on the received response data, it is rendered as a 3D object.
[1289] The agent in the virtual space responds to the user using natural movements and voice, saying, "Here's how to make curry: First, fry the onions. Next, add the meat, then the vegetables and simmer. Finally, add the curry roux and it's done."
[1290] In this way, users can have a virtual reality experience that feels like they are having a conversation with a real person. The system responds immediately to detailed questions, enabling intuitive and efficient information acquisition.
[1291] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1292] Step 1:
[1293] The user puts on the virtual reality device and launches a dedicated concierge application. The input is the user putting on the virtual reality device and launching the application, and the output is the agent being ready to appear in the virtual space. Specifically, the user puts on the headset and controllers and clicks on the application icon to launch it.
[1294] Step 2:
[1295] The user asks a question by voice or text. The input is the user's speech or text input, and the output is the data of the user's question. Specifically, the user speaks "Tell me how to make tomato sauce" or enters similar text on a virtual keyboard.
[1296] Step 3:
[1297] The device records the user's voice input in real time and converts it into text data using the Google Speech-to-Text API or Microsoft Azure Speech Service. The input is the user's voice data and the output is text data. Specifically, the voice recognition software analyzes the voice data and converts it into text such as "Please tell me how to make tomato sauce."
[1298] Step 4:
[1299] The terminal assembles the converted text data into packets as question data and sends them to the server via the HTTPS protocol. The input is the text data converted from the voice, and the output is the question data sent to the server. Specifically, the terminal encodes and sends the data to the server.
[1300] Step 5:
[1301] The server receives the query data sent from the terminal. The input is the query data sent from the terminal, and the output is raw data for analysis. In concrete terms, the web server (e.g., Apache or Nginx) passes the received data to the application server.
[1302] Step 6:
[1303] The server analyzes the received question data using natural language processing (NLP). The input is the received question data, and the output is tokenized and normalized data. Specifically, an NLP engine implemented in Python, Java, or other languages tokenizes the data and performs grammatical and semantic analysis.
[1304] Step 7:
[1305] The server inputs the preprocessed data into a generative artificial intelligence model (e.g., OpenAI GPT-3) to generate an appropriate answer. The input is natural language processed data, and the output is the generated response data. Specifically, it inputs a prompt sentence (e.g., "Please tell me how to make tomato sauce.") into GPT-3 and obtains the generated response (e.g., "Here's how to make tomato sauce. First, chop the tomatoes...").
[1306] Step 8:
[1307] The server formats the generated response data in a format such as JSON and sends it to the terminal via the HTTPS protocol. The input is the generated response data, and the output is the response data sent to the terminal. Specifically, the server formats, encodes, and sends the data.
[1308] Step 9:
[1309] The terminal receives the response data sent from the server and checks and converts the format. The input is the response data sent from the server, and the output is renderable data. Specific operations include decoding the received data and generating speech synthesis and action generation data for the agent.
[1310] Step 10:
[1311] The device uses a 3D engine such as Unity or Unreal Engine to display the agent in a virtual space. The input is renderable data, and the output is the agent's response, presented visually and audibly. Specifically, the agent explains, with hand gestures and facial expressions, "Here's how to make tomato sauce. First, chop the tomatoes..."
[1312] (Application example 1)
[1313] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1314] In recent years, there have been many attempts to utilize technology to improve the customer experience in brick-and-mortar stores. However, current systems often make it difficult for users to quickly obtain detailed product information, especially when visiting a store for the first time or in an unfamiliar product category. Furthermore, when customers are hesitant to directly interact with store staff or when the store is crowded, quickly obtaining information becomes even more difficult. Therefore, there is a need for a system that allows users to easily obtain product information by voice and presents that information visually and audibly.
[1315] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1316] In this invention, the server includes: means for a user to wear a virtual reality device; means for collecting voice or text input by the user; means for analyzing the collected user question data and performing natural language processing; means for inputting the analyzed data into a generative artificial intelligence model and generating an appropriate response; means for three-dimensionally displaying the generated response data in the form of a concierge in a virtual reality space; means for forming a user interface for providing product information in a physical store; and means for generating answers to the user's voice questions to provide product information and presenting them to the user by voice, thereby enabling users to quickly and easily obtain product information in a physical store and receive that information visually and audibly.
[1317] A "user" is a person who uses the system to obtain information.
[1318] A "virtual reality device" is a device worn by a user to access a virtual reality space that includes vision and sound.
[1319] A "voice or text collection means" is any device or software that collects user input in the form of voice or text.
[1320] "Question data" is data including a question entered by a user.
[1321] "Natural language processing" is the process of analyzing collected question data and converting it into a format that a computer can understand.
[1322] A "generative artificial intelligence model" is an artificial intelligence model that inputs natural language processed data and generates an appropriate response.
[1323] "Response data" is data that includes an answer generated by a generative artificial intelligence model.
[1324] The "means for three-dimensional display" refers to a device or software for displaying the response data in the form of a three-dimensional concierge in a virtual reality space.
[1325] A "brick and mortar store" is a retail outlet that exists in a physical location.
[1326] "User interface" means an interface through which a user inputs information and interacts with the concierge system.
[1327] "Product information" is detailed information about the products being sold.
[1328] The "means for providing by voice" refers to a device or software for presenting the generated response data to the user in voice format.
[1329] The present invention provides a system that allows users to acquire product information in a physical store using a virtual reality device. The system displays appropriate responses to questions in the form of a three-dimensional concierge and presents them to the user via audio.
[1330] System configuration
[1331] The main components of the system are:
[1332] 1. Users
[1333] 2. Terminal
[1334] 3. Server
[1335] User operations
[1336] The user puts on the virtual reality device and asks a question about a product in the store by voice, for example, "What are the ingredients in this product?" This voice is transmitted to the terminal through the microphone of the virtual reality device.
[1337] Terminal handling
[1338] The device processes the user's voice input in the following steps:
[1339] 1. Use speech recognition software (e.g., SpeechRecognition library, Google Text-to-Speech) to convert the audio data into text data.
[1340] 2. The converted text data is packed into packets and sent to the server.
[1341] 3. Receive the response data from the server and format it appropriately.
[1342] 4. Render the generated answer as a 3D object and a virtual concierge will present the answer visually and audibly.
[1343] Server Processing
[1344] The server processes the query data in the following steps:
[1345] 1. Receive text data sent from the device.
[1346] 2. Perform natural language processing (NLP) and analyze the data.
[1347] 3. Use a generative AI model (e.g., GPT-3) to generate an appropriate answer. For example, if a user asks, "What are the ingredients in this product?", the generative AI will generate an answer such as, "This product's ingredients include water, glycerin, parabens, vitamin E, etc."
[1348] 4. The generated response data is sent to the terminal.
[1349] Hardware and software used
[1350] Hardware: Virtual reality devices (smart glasses, head-mounted displays, etc.), microphones, smartphones
[1351] Software: SpeechRecognition Library, Google Text-to-Speech (gTTS), OpenAI API
[1352] Specific examples
[1353] For example, if a user is browsing products in a physical store and asks through smart glasses, "What are the ingredients in this product?", the voice is sent to the device via a microphone. The device converts the voice into text, and this text data is sent to the server. The server uses a generative AI model to generate an answer, returning, for example, "The ingredients in this product include water, glycerin, parabens, vitamin E, etc." to the device. The device then provides this answer to the user as a 3D display and audio.
[1354] Prompt Sentence Examples
[1355] An example of a prompt sentence is, "A user asks a question about a product in a store. Question: 'What are the ingredients in this product?' Please provide an easy-to-understand answer to this question." By inputting this prompt sentence into a generative AI model, an appropriate answer will be generated.
[1356] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1357] Step 1:
[1358] The user wears a virtual reality device and asks a question about a product by voice. This voice is collected by the built-in microphone of the virtual reality device. The input is the user's voice, and the output is voice data.
[1359] Step 2:
[1360] The device uses speech recognition software (e.g., SpeechRecognition library, Google Text-to-Speech) to convert the collected voice data into text data. Here, speech is converted into text, such as "Please tell me the ingredients of this product." The input is voice data, and the output is text data.
[1361] Step 3:
[1362] The terminal packs the converted text data into packets and sends them to the server. The input is text data and the output is packet data.
[1363] Step 4:
[1364] The server receives packet data sent from the terminal and analyzes the data using natural language processing (NLP), where text data is tokenized and normalized. The input is packet data, and the output is analyzed data.
[1365] Step 5:
[1366] The server inputs the analyzed data into a generative artificial intelligence model (e.g., GPT-3) to generate an appropriate answer. For example, in response to the question, "What are the ingredients of this product?", an answer such as, "This product contains water, glycerin, parabens, vitamin E, etc." is generated. The input is the analyzed data, and the output is the generated answer data.
[1367] Step 6:
[1368] The server sends the generated response data to the terminal, where the input is the generated response data and the output is the reformatted response data.
[1369] Step 7:
[1370] The device then formats the answer data received from the server and renders it as a 3D object, where a concierge in a virtual reality space presents the generated answer visually and audibly. The input is the reformatted answer data, and the output is 3D visual and audio data.
[1371] Step 8:
[1372] The user receives visual and audio responses from the concierge in the virtual reality space and confirms detailed product information. The input is 3D visual and audio data, and the output is the user's understanding and acquisition of product information.
[1373] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1374] The present invention combines a 3D concierge system that combines VR technology and generative AI with an emotion engine that recognizes the user's emotions. This system allows the user to wear a virtual reality device, ask questions by voice or text, and display responses in 3D, while simultaneously generating appropriate responses based on the user's emotions. Specific embodiments for implementing the present invention are described below.
[1375] Program processing
[1376] The system consists of three main components: the user, the device, and the server, among which an emotion engine has been newly incorporated. Each plays a specific role and works in conjunction with each other.
[1377] User operations
[1378] 1. Booting the system
[1379] The user puts on the virtual reality device and launches a dedicated concierge application, at which point a concierge appears in the virtual space.
[1380] 2. Enter your question
[1381] The user asks a question by voice or text. For example, the user says, "What's the weather like in Tokyo today?"
[1382] Terminal handling
[1383] 3. Speech Recognition and Text Conversion
[1384] The user's voice input is recorded in real time by the terminal and converted into text data by voice recognition software.
[1385] 4. Emotion recognition
[1386] The device's built-in emotion engine analyzes the user's voice tone and facial expressions to generate emotion data, which identifies the user's emotional state (e.g., joy, sadness, anger, etc.).
[1387] 5. Sending Question Data and Emotion Data
[1388] The converted text question data and the analyzed emotion data are packaged into packets and sent to the server.
[1389] Server Processing
[1390] 6. Analysis of received data
[1391] The server receives and confirms the question data and emotion data sent from the terminal.
[1392] 7. Response Generation Using Generative AI
[1393] The question data is processed using natural language processing and input into a generative AI model. Emotional data is then incorporated to generate a response that reflects the user's emotions.
[1394] For example, if a user asks in a sad tone, "How's the weather in Tokyo today?", the generative AI will generate an answer that takes emotions into consideration, such as, "The weather in Tokyo is sunny today. It's nice, so why not go outside for a bit to change your mood?"
[1395] 8. Response Data Formatting and Transmission
[1396] The generated response data is formatted and sent to the terminal.
[1397] Terminal Processing (cont.)
[1398] 9. Receiving and displaying response data
[1399] The device receives the response data sent from the server and renders it as a three-dimensional object.
[1400] The concierge in the virtual space responds to the user with natural movements and voice, and also responds according to the user's emotions.
[1401] Specific examples
[1402] For example, the specific operations and processing flow when a user asks, "I'm really tired from work today. Please tell me how to relax" in a stressful daily life is shown below.
[1403] User
[1404] Put on the virtual reality device and say, "I'm really tired from work today. Please tell me how to relax."
[1405] Terminal
[1406] The user's speech is recorded and converted into text using speech recognition software.
[1407] The emotion engine detects fatigue and stress from the user's voice tone and generates emotion data.
[1408] The question data and emotion data are packaged into a packet and sent to the server.
[1409] server
[1410] Receives question data and emotion data and performs natural language processing.
[1411] Using generative AI, in response to the question, "I'm really tired from work today. Please tell me how to relax," the system generates a response that takes emotions into consideration, such as, "To relax, why not try taking some slow, deep breaths or taking a warm bath? I also recommend listening to your favorite music."
[1412] The generated response data is formatted and sent to the terminal.
[1413] Terminal
[1414] Based on the received response data, the concierge in the virtual space responds to the user with natural movements and voice.
[1415] The concierge tells the user, "To relax, why not try taking some slow, deep breaths or taking a warm bath? We also recommend listening to your favorite music."
[1416] In this way, by having each step work in conjunction with one another, a system is realized that allows a user to ask questions in a virtual reality space and receive answers that correspond to their emotions.
[1417] The processing flow will be explained below.
[1418] Step 1:
[1419] The user puts on the virtual reality device, turns on the device, and launches the dedicated concierge application, which displays the concierge in the virtual space.
[1420] Step 2:
[1421] The user types a question. The user can either speak aloud, such as "Work today is really tiring. What are some ways to relax?" or type the question as text using the virtual keyboard.
[1422] Step 3:
[1423] The device collects voice data: the user's speech is recorded by the VR device's microphone and sent to the voice recognition software.
[1424] Step 4:
[1425] The device converts the voice data into text data. Voice recognition software analyzes the voice data in real time and generates text data such as, "I'm really tired from work today. Please tell me how to relax."
[1426] Step 5:
[1427] The device collects the user's voice tone and facial expressions, and the emotion engine analyzes the user's voice tone and facial expressions to identify the user's emotional state.
[1428] Step 6:
[1429] The terminal generates emotion data, and based on the collected emotion information, generates emotion data indicating that the user is tired.
[1430] Step 7:
[1431] The device sends the question data and emotion data to the server, which then assembles the converted text data and generated emotion data into packets and sends them to the server via the Internet.
[1432] Step 8:
[1433] The server receives the question data and emotion data, analyzes the received data, and prepares it for processing.
[1434] Step 9:
[1435] The server performs natural language processing, tokenizing and normalizing the question data and converting it into a format that can be input to a generative AI model.
[1436] Step 10:
[1437] The server generates responses using a generative artificial intelligence model. It inputs question data and emotion data into the model and generates an appropriate response.
[1438] Step 11:
[1439] The server formats the response data, examining the generated response and formatting it in a format that can be displayed within the virtual reality space.
[1440] Step 12:
[1441] The server sends the formatted response data to the terminal, which then packages the response data into packets and sends them to the terminal over the Internet.
[1442] Step 13:
[1443] The terminal receives the response data. The terminal receives the response data sent from the server and checks its format.
[1444] Step 14:
[1445] The device renders the response data as a 3D object, and processes the received data so that the concierge can respond in a virtual space using natural movements and voice.
[1446] Step 15:
[1447] The user confirms the concierge's response by visually and audibly hearing the concierge in the virtual reality space reply to the user, "To relax, why not try some slow, deep breathing or a warm bath? We also recommend listening to your favorite music."
[1448] In this way, each step works in conjunction to realize a system that allows a user to ask questions in a virtual reality space and receive appropriate answers based on their emotions.
[1449] Example 2
[1450] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1451] Conventional virtual reality concierge systems can generate responses to questions entered by users, but they are unable to provide appropriate responses that take the user's emotions into consideration. This limits the user experience and prevents the user from receiving the personalized responses they desire.
[1452] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1453] In this invention, the server includes a means for allowing a user to wear a virtual reality device, a means for collecting voice or text input by the user, a built-in emotion engine for generating emotion data from the collected voice data, and a means for inputting the analyzed data into a generative artificial intelligence model to generate an appropriate response, thereby enabling a response that takes into account the user's emotions.
[1454] A "virtual reality device" is a device that allows a user to experience a virtual space through sight and sound.
[1455] "Voice or text collection means" refers to hardware and software components for collecting user-spoken voice or text information.
[1456] "Means for analyzing question data and performing natural language processing" refers to technology for analyzing the content of questions entered by users and processing them in an appropriate manner.
[1457] A "generative artificial intelligence model" is a set of artificial intelligence algorithms that perform natural language responses and other generative tasks based on given data.
[1458] "Means for three-dimensionally displaying response data in the form of a concierge in a virtual reality space" refers to a method that utilizes three-dimensional graphics technology to communicate the generated response to the user visually and audibly.
[1459] The "emotion engine" is an analysis device and a group of algorithms for extracting emotional data from a user's tone of voice and facial expressions.
[1460] The "means for generating a response that takes emotion into consideration" refers to a process and technology for generating a response that is appropriate to a user's emotion based on the user's emotion data.
[1461] This invention is a 3D concierge system that combines virtual reality technology and generative artificial intelligence, and further incorporates an emotion engine that recognizes the user's emotions, thereby providing responses that take the user's emotions into consideration. The system allows the user to wear a virtual reality device, input questions by voice or text, and displays the responses in 3D while simultaneously generating appropriate responses according to the user's emotions.
[1462] This system consists of three main components: the user, the terminal, and the server, each of which plays a specific role and works in conjunction with each other. The emotion engine has been newly incorporated, and the operation of each component is shown below.
[1463] User operations
[1464] First, the user puts on a virtual reality device (e.g., a "VR headset" as it is commonly called) and launches a dedicated concierge application. At this point, a three-dimensional concierge appears in the virtual space. The user asks questions through a voice recognition and text input interface. For example, when the user says, "How is the weather in Tokyo today?", the voice is sent to the device.
[1465] Terminal handling
[1466] The device records the user's voice input in real time and converts it into text data using voice recognition software (e.g., generically called a "voice recognition engine"). The device also has a built-in emotion engine (e.g., generically called an "emotion analysis module") that recognizes the user's emotions and analyzes voice tone and facial expression data to generate emotion data. For example, if a user says in a tired voice, "I'm really tired from work today. Please tell me how to relax," the emotion will be identified as "fatigue." The converted text question data and analyzed emotion data are packaged into a data packet and sent to the server.
[1467] Server Processing
[1468] The server receives the data packet sent from the device and analyzes its contents. It uses a natural language processing (NLP) module (e.g., a generic term "natural language processing engine") to analyze the question text and generate a prompt sentence, which is input to a generative artificial intelligence model (e.g., a generic term "generative AI model"). The generated prompt sentence also includes emotional data. For example, the following content is input to the generative AI model: "A user asks, 'I'm really tired from work today. Please tell me how to relax.' This user is feeling tired. Please suggest relaxation methods to the user that take their emotions into consideration." The generative AI model then generates an appropriate response to the question, such as, "Deep breathing and a warm bath are good ways to relax. I also recommend listening to your favorite music."
[1469] The generated response data is formatted and sent to the device. The device then renders the received response data as a 3D object. The concierge responds to the user in the virtual space using natural movements and voice. For example, if the user asks, "I'm really tired from work today. What are some ways to relax?" the concierge will tell the user, "Deep breathing and a warm bath are good ways to relax. I also recommend listening to your favorite music."
[1470] In this way, a system is realized that allows a user to ask questions in a virtual reality space and receive answers that take emotion into consideration.
[1471] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1472] Step 1:
[1473] The user wears a virtual reality device and launches a dedicated concierge application, which displays a 3D concierge in the virtual space. The input is the virtual reality device worn by the user and the application that operates it, and the output is the display of the 3D concierge.
[1474] Step 2:
[1475] The user inputs a question by voice or text. For example, they might say, "What's the weather like in Tokyo today?" The input is the user's voice or text, which is sent to the device. The output is the voice data sent to the device.
[1476] Step 3:
[1477] The device records the user's voice input in real time and converts it into text data using voice recognition software. The voice recognition software used is a general voice recognition engine. The input for text conversion is the recorded voice data, and the output is text data. Specifically, the voice signal is frequency analyzed and converted into text based on a language model.
[1478] Step 4:
[1479] The device's emotion engine analyzes the user's voice tone and facial expressions to generate emotion data. The emotion engine uses a general emotion analysis module to analyze voice intonation and facial muscle movements. The input is the user's voice tone and facial expression data, and the output is emotion data.
[1480] Step 5:
[1481] The terminal assembles the converted text question data and the analyzed emotion data into packets and sends them to the server. The input is text data and emotion data, and the output is the data packet sent to the server.
[1482] Step 6:
[1483] The server receives data packets sent from the terminal and analyzes their contents. The input is the received data packet, and the output is the analyzed question data and emotion data. The server checks the consistency of the data and requests a retransmission if there is an inconsistency.
[1484] Step 7:
[1485] The server uses a natural language processing engine to analyze the question text, generate a prompt sentence, and input it to the generative artificial intelligence model. The generative artificial intelligence model uses a general generative AI model. The input is question data and emotional data, and the output is response data that takes the emotion into consideration. For example, the prompt sentence input is, "The user asked, 'I'm really tired from work today. Please tell me how to relax.' This user is feeling tired. Please suggest a relaxation method for the user that takes their emotions into consideration."
[1486] Step 8:
[1487] The server formats the generated response data and sends it to the terminal. The input is the generated response data and the output is the formatted data sent to the terminal.
[1488] Step 9:
[1489] The device renders the received response data as a 3D object, using a 3D graphics engine. The input is the response data received from the server, and the output is a response made by the concierge in the virtual space using natural movements and voice. The concierge naturally tells the user, "To relax, deep breathing and a warm bath are good ways to do it. I also recommend listening to your favorite music."
[1490] (Application example 2)
[1491] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1492] Conventional systems using virtual reality technology have the ability to provide natural responses to user questions, but they are unable to provide individual responses based on the user's emotions. This limits the user's satisfaction and the quality of the experience. The objective of this invention is to provide a more personalized experience by recognizing the user's emotions and providing appropriate responses based on those emotions.
[1493] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1494] In this invention, the server includes a means for incorporating an emotion engine that recognizes and analyzes the user's emotional state, a means for inputting the data into a generative AI model to generate an appropriate response, and a means for displaying the generated response data in three dimensions in the form of a concierge in a virtual reality space, thereby enabling personalized responses according to the user's emotions.
[1495] A "virtual reality device" is a device worn by a user to immerse themselves in a virtual reality space, and refers to devices such as head-mounted displays and VR goggles.
[1496] "Voice or text collection means" refers to an input device such as a microphone or keyboard for capturing voice or text data input by a user in real time.
[1497] "Means for analyzing question data and performing natural language processing" refers to software or algorithms for analyzing collected user questions and semantically analyzing the text using natural language processing technology.
[1498] A "generative artificial intelligence model" is an artificial intelligence model that takes natural language processed data as input and generates appropriate responses, and refers to a model that generates responses based on learned patterns.
[1499] "Means for three-dimensional display" refers to rendering software and display devices for displaying the generated response in three dimensions within a virtual reality space.
[1500] An "emotion engine" refers to hardware or software that analyzes a user's tone of voice, facial expressions, etc., to identify the user's emotional state (e.g., joy, sadness, anger, etc.).
[1501] "Means for generating and displaying a response according to emotions" refers to a series of processes or devices that allow a generative artificial intelligence model to create a response based on the emotional state identified by the emotion engine and to display that response appropriately.
[1502] This invention is a 3D concierge system that combines virtual reality technology and generative artificial intelligence, incorporating an emotion engine that recognizes the user's emotions. With this system, the user puts on a virtual reality device and asks a question. The response to the question is displayed in 3D, and the user can receive an appropriate response according to their emotion.
[1503] The system mainly consists of three main components: users, terminals, and servers.
[1504] User operations
[1505] First, the user puts on a virtual reality device and launches a dedicated concierge application. This refers to a virtual reality device such as a head-mounted display or VR goggles. When the user asks a question by voice or text, the input is collected by the device. For example, the user might say, "Tell me about this product."
[1506] Terminal handling
[1507] The device converts the user's speech into text data using voice recognition software (e.g., Vosk). The device's built-in emotion engine analyzes the user's tone of voice and facial expressions to generate emotion data. The emotion engine uses OpenFace, for example. For example, if a user asks a question in a curious tone, the device identifies the user's curious emotional state.
[1508] Server Processing
[1509] The server receives the question data and emotion data sent from the device and performs natural language processing. The question data is input into a generative artificial intelligence model (such as OpenAI's GPT-3.5 Turbo) to generate a response based on the user's emotion. For example, if a user asks out of curiosity, "Tell me about this product," the server generates a response such as, "This product is new this season, made of silk, and is characterized by its extremely soft feel."
[1510] Terminal Processing (cont.)
[1511] The device receives the response data sent from the server and displays it in 3D as a concierge in the virtual reality space. This display is rendered using a VR library (e.g., VRRenderer). It is also possible to respond by voice using pyttsx3 or similar.
[1512] Prompt Sentence Examples
[1513] For example, the following prompt statement is generated:
[1514] The user was curious and asked: "Tell me about this product."
[1515] As described above, the present invention aims to enable users to ask questions in a virtual reality space and receive personalized responses to those questions based on their emotions, thereby providing a richer experience for the user.
[1516] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1517] Step 1: User interaction
[1518] The user puts on the virtual reality device and launches a dedicated concierge application. A concierge appears in the virtual space, and the user asks a question by voice or text. Input (voice or text) begins. For example, the user might say, "Tell me about this product."
[1519] Step 2: Collecting audio and converting it to text
[1520] The device records the user's speech with a microphone and converts it into text data using voice recognition software (e.g., Vosk). In the process of converting voice data (input) into text data (output), noise removal and voice analysis are performed. Specifically, the recorded voice data is converted into text data such as "Tell me about this product."
[1521] Step 3: Recognize emotions
[1522] The device's built-in emotion engine (e.g., OpenFace) captures the user's voice tone and facial expressions with a camera and analyzes the data. It then identifies the type of emotion and generates emotion data (e.g., curious). The emotion data (output) is packaged into a packet of question data (text data).
[1523] Step 4: Send data
[1524] The terminal assembles the converted question data and the analyzed emotion data into packets and sends them to the server. The input (question data and emotion data) is configured as a data packet to be sent to the server.
[1525] Step 5: Receive and analyze question and sentiment data
[1526] The server receives the question data and emotion data sent from the device and analyzes each data. It uses natural language processing (NLP) to analyze the question data and integrate the emotion data. It then outputs a prompt sentence generated by analyzing the data.
[1527] Step 6: Generate a response
[1528] The server inputs the prompt into a generative artificial intelligence model (e.g., GPT-3.5 Turbo) to generate an appropriate response. In this scenario, if the user's emotion of curiosity is identified, a detailed response tailored to that interest is generated. For example, a response such as "This product is new this season, made of silk, and is characterized by its extremely soft feel" may be output.
[1529] Step 7: Format and send response data
[1530] The server formats the generated response data and sends it as a data packet to the terminal. The response data (output) is now properly formatted and ready to be sent.
[1531] Step 8: Receiving response data and displaying it in 3D
[1532] The device receives the response data sent from the server and displays it in 3D as a concierge in the virtual space. Using VR rendering software (e.g., VRRenderer), the concierge responds with natural movements and voice. Specifically, the concierge explains, "This product is new this season, made of silk, and is extremely soft to the touch."
[1533] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1534] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1535] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1536] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1537] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1538] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1539] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1540] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1541] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1542] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1543] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1544] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1545] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1546] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1547] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1548] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1549] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1550] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1551] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1552] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1553] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1554] The following is further disclosed regarding the above embodiment.
[1555] (Claim 1)
[1556] means for a user to wear a virtual reality device;
[1557] means for collecting user-entered speech or text;
[1558] A means for analyzing the collected user question data and performing natural language processing;
[1559] means for inputting the analyzed data into a generative artificial intelligence model to generate an appropriate response;
[1560] a means for displaying the generated response data in three dimensions in the form of a concierge in a virtual reality space;
[1561] A system including:
[1562] (Claim 2)
[1563] 2. The system according to claim 1, wherein the three-dimensional representation of the concierge responds to the user with natural movements and voice.
[1564] (Claim 3)
[1565] 10. The system of claim 1, further comprising means for converting a user's voice input into text in real time and transmitting it to the generative artificial intelligence model.
[1566] "Example 1"
[1567] (Claim 1)
[1568] means for a user to wear a virtual reality device;
[1569] means for collecting user-entered speech or text;
[1570] A means for analyzing the collected user question data;
[1571] means for natural language processing the analyzed data;
[1572] means for inputting the natural language processed data into an artificial intelligence model for generating an appropriate response;
[1573] a means for three-dimensionally displaying the generated response data in the form of an agent in a virtual reality space;
[1574] means for presenting the generated response data in visual and audio form;
[1575] A system including:
[1576] (Claim 2)
[1577] 2. The system according to claim 1, wherein the three-dimensional representation of the agent responds to the user with natural movements and voices.
[1578] (Claim 3)
[1579] 10. The system of claim 1, further comprising means for converting user voice input into text in real time and transmitting the same to the artificial intelligence model for generation.
[1580] "Application Example 1"
[1581] (Claim 1)
[1582] means for a user to wear a virtual reality device;
[1583] means for collecting user-entered speech or text;
[1584] A means for analyzing the collected user question data and performing natural language processing;
[1585] means for inputting the analyzed data into a generative artificial intelligence model to generate an appropriate response;
[1586] a means for displaying the generated response data in three dimensions in the form of a concierge in a virtual reality space;
[1587] means for forming a user interface for providing product information in a physical store;
[1588] means for generating a response to a user's voice question to provide product information and presenting the response to the user by voice;
[1589] A system including:
[1590] (Claim 2)
[1591] 2. The system according to claim 1, wherein the three-dimensional representation of the concierge responds to the user with natural movements and voice.
[1592] (Claim 3)
[1593] 10. The system of claim 1, further comprising means for converting a user's voice input into text in real time and transmitting it to the generative artificial intelligence model.
[1594] "Example 2: Combining Emotion Engines"
[1595] (Claim 1)
[1596] means for a user to wear a virtual reality device;
[1597] means for collecting user-entered speech or text;
[1598] A means for analyzing the collected user question data and performing natural language processing;
[1599] means for inputting the analyzed data into a generative artificial intelligence model to generate an appropriate response;
[1600] a means for displaying the generated response data in three dimensions in the form of a concierge in a virtual reality space;
[1601] A means for incorporating an emotion engine that generates emotion data from collected voice data;
[1602] means for generating a response that takes into consideration the user's emotions based on the emotion data;
[1603] A system including:
[1604] (Claim 2)
[1605] 2. The system according to claim 1, wherein the three-dimensional representation of the concierge responds to the user with natural movements and voice.
[1606] (Claim 3)
[1607] 10. The system of claim 1, further comprising means for converting a user's voice input into text in real time and transmitting it to the generative artificial intelligence model.
[1608] "Application example 2 when combining emotion engines"
[1609] (Claim 1)
[1610] means for a user to wear a virtual reality device;
[1611] means for collecting user-entered speech or text;
[1612] A means for analyzing the collected user question data and performing natural language processing;
[1613] means for inputting the analyzed data into a generative artificial intelligence model to generate an appropriate response;
[1614] a means for displaying the generated response data in three dimensions in the form of a concierge in a virtual reality space;
[1615] means for incorporating an emotion engine for recognizing and analyzing the user's emotional state;
[1616] means for generating and displaying a response according to the user's emotion;
[1617] A system including:
[1618] (Claim 2)
[1619] 2. The system according to claim 1, wherein the three-dimensional representation of the concierge responds to the user with natural movements and voice.
[1620] (Claim 3)
[1621] 10. The system of claim 1, further comprising means for converting a user's voice input into text in real time and transmitting it to the generative artificial intelligence model. [Explanation of symbols]
[1622] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. means for a user to wear a virtual reality device; means for collecting user-entered speech or text; A means for analyzing the collected user question data and performing natural language processing; means for inputting the analyzed data into a generative artificial intelligence model to generate an appropriate response; a means for displaying the generated response data in three dimensions in the form of a concierge in a virtual reality space; A system including:
2. 2. The system according to claim 1, wherein the three-dimensional representation of the concierge responds to the user with natural movements and voices.
3. 10. The system of claim 1, further comprising means for converting user speech input into text in real time and transmitting it to the generative artificial intelligence model.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A