System
A system using sensors and cameras to predict and answer user questions addresses information overload by efficiently generating and providing relevant information in real-time.
Patent Information
- Application Number
- JP2024118065
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-23
- Publication Date
- 2026-02-04
AI Technical Summary
Users face challenges in quickly and accurately obtaining relevant information due to information overload, and lack effective mechanisms to express questions and receive appropriate answers, especially in new situations.
A system that includes sensors and cameras to detect user behavior and gaze, analyzing this data to predict questions, displaying them for selection, and generating appropriate answers using natural language processing.
Efficiently supports users in acquiring information by predicting and providing relevant answers, streamlining learning and daily tasks in information-overloaded environments.
Smart Images

Figure 2026017283000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] In modern society, there is an information overload, making it difficult for users to quickly and accurately obtain the information they need. It is also not easy to come up with the right questions at the right time in learning or daily life. In particular, when users are faced with new situations or information, they lack the means to effectively express their questions and obtain appropriate answers to those questions. Therefore, there is a need for a system that efficiently supports users' information acquisition and learning processes. [Means for solving the problem]
[0005] In order to solve the above problems, the present invention provides the following means.
[0006] We propose a system that includes a data collection means including a sensor and a camera for detecting user behavior, a data analysis means for analyzing the detected user behavior data and generating predicted questions, a display means for displaying the generated questions to the user and allowing the user to select one with their line of sight, and an answer generation means for generating and providing an appropriate answer to the question selected by the user. Furthermore, the sensor and camera include a means for tracking the user's line of sight, so that the system predicts questions based on the information the user is focusing on and provides appropriate answers using natural language processing technology.
[0007] A "sensor" is a device that detects changes or movements in the physical environment and a means of electronically recording and transmitting that data.
[0008] A "camera" is a device that acquires visual information and stores or transmits it as image or video data.
[0009] "Data collection means" refers to a series of mechanisms, including sensors and cameras, that detect the user's actions and gaze and record and transmit that information.
[0010] "Data analysis means" refers to the means for analyzing collected user behavior data and predicting and generating questions that users are likely to have.
[0011] The "display means" is a display device that visually presents the generated questions to the user and allows the user to select them with their line of sight.
[0012] The "answer generation means" is a means for generating an appropriate answer based on a question selected by a user and providing the answer using natural language processing technology or the like.
[0013] "Eye tracking" is a technology that tracks and records the direction of a user's gaze and point of attention in real time.
[0014] "Natural language processing technology" is a computer science technology for understanding and generating human language, and is used to generate appropriate answers. [Brief explanation of the drawings]
[0015] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14]FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0016] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0017] First, the terms used in the following description will be explained.
[0018] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0019] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0020] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0021] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0022] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0023] [First embodiment]
[0024] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0025] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0026] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0027] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0028] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0029] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0030] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0031] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0032] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0033] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0034] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0035] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0036] This invention provides a system that efficiently supports users in acquiring information by detecting user behavior, generating predicted questions, and providing appropriate answers to those questions. This system uses sensors and cameras to capture the user's behavior and gaze, analyzes the data using a cloud server, generates appropriate questions, displays them to the user, and generates and provides answers to the selected questions.
[0037] Program processing flow
[0038] This system operates based on the interaction between the "terminal" (smart glasses), the "server" on the cloud, and the "user."
[0039] Device: Using built-in sensors and cameras, the device captures the user's movements and gaze in real time. The device temporarily stores the collected data and sends it to a cloud server as needed.
[0040] Server: The cloud server receives and analyzes the user data sent from the device. The data analysis method generates predicted questions based on the user's current situation. The generated question list is then sent from the server to the device.
[0041] Terminal: The terminal displays the question list received from the server on the display. When the user selects a question candidate on the display with their gaze, the selected question is sent to the server.
[0042] Server: The server analyzes the question sent by the user and generates an appropriate answer using database search and natural language processing techniques. The generated answer is then sent to the device.
[0043] Terminal: The terminal provides the received answer to the user either visually on a display or audibly using speech synthesis technology.
[0044] Specific examples
[0045] Example 1: Use at a restaurant
[0046] User: Looking at a new menu at a restaurant.
[0047] Device: Captures menus with the built-in camera and tracks the user's gaze.
[0048] Server: Recognizes when the user is looking at the menu and generates relevant question suggestions (e.g., "What are the ingredients?", "How many calories?").
[0049] Device: The question candidates are displayed on the screen and the user selects one. The user selects a question such as "How many calories?"
[0050] Server: Retrieves the calorie information for the menu items from the database and generates the answer.
[0051] Terminal: Provides answers visually or audibly.
[0052] Example 2: Use during learning
[0053] User: Studying by reading a textbook.
[0054] Device: Built-in camera and sensors identify where you are reading and track your gaze.
[0055] Server: Generates relevant question candidates (e.g., "What is the background to this incident?", "Who are the main characters?") based on the content of the textbook.
[0056] Terminal: The question candidates are displayed on the screen and the user can select one. The user selects the question "What is the background of this incident?"
[0057] Server: Obtains background information about the incident from textbook content and related materials and generates answers.
[0058] Terminal: Provides answers visually or audibly.
[0059] The system of the present invention is a powerful support tool for users to quickly obtain appropriate information, and plays an important role in today's information-overloaded society. It streamlines the user's learning process and provides practical solutions to various questions in daily life.
[0060] The processing flow will be explained below.
[0061] Step 1:
[0062] Device: The device uses built-in sensors and cameras to capture the user's movements and gaze in real time. The sensors detect the user's movements (e.g., walking, standing, sitting, etc.), and the camera tracks the user's gaze direction. The captured data is temporarily stored on the device.
[0063] Step 2:
[0064] Terminal: Processes the collected data, extracts necessary information, and sends it to the cloud server. This transmitted data includes the user's current behavioral patterns and gaze information.
[0065] Step 3:
[0066] Server: The cloud server receives the transmitted data and uses data analysis methods to analyze the user's behavior and gaze data. Specifically, it predicts what information the user is paying attention to and what questions they are likely to have.
[0067] Step 4:
[0068] Server: Based on the analysis results, a list of predicted questions is generated. The list of questions includes questions that the user is likely to have in that situation (e.g., "What are the ingredients in this dish?"). The generated list of questions is then sent back to the device.
[0069] Step 5:
[0070] Terminal: The received question list is displayed on the display. The displayed questions are displayed so that they can be visually recognized within the user's field of vision.
[0071] Step 6:
[0072] User: Selects a question displayed on the display by gazing at it. For example, if a user fixates their gaze on the question "How many calories does it have?" for a certain period of time, that question is recognized as selected.
[0073] Step 7:
[0074] The terminal sends the selected question to the server. The sent data includes the selected question as well as the user's current situation and context information.
[0075] Step 8:
[0076] Server: The server analyzes the received question and generates an appropriate answer. It searches for relevant information from a database and uses natural language processing techniques to create an answer in an easy-to-read format.
[0077] Step 9:
[0078] Server: Sends the generated answer to the terminal, which may be presented visually or audibly.
[0079] Step 10:
[0080] Terminal: Provides the received answer to the user, either visually displaying the answer on a display or audibly using speech synthesis technology.
[0081] These are the specific processing steps of the "Chatty Glasses" system. By capturing the user's behavior and gaze in real time, and using that information to predict and generate appropriate questions and provide answers, the system helps users obtain information more efficiently.
[0082] Example 1
[0083] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0084] In modern society, it is important for users to quickly obtain appropriate information for efficient decision-making and learning. However, it is not easy to find the information they need from a vast amount of information. There is also a lack of mechanisms that can accurately identify when a user has a question and quickly provide an appropriate answer to that question. The present invention aims to solve these problems by providing a system that improves the efficiency of information acquisition and supports users in resolving their questions.
[0085] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0086] In this invention, the server includes a data collection means including a sensor and a camera for detecting user behavior, a communication means for transmitting the detected user behavior data to the cloud server, a data analysis means for analyzing the transmitted data and generating predicted questions, a display means for displaying the generated question list to the user and allowing the user to select one with their line of sight, and an answer generation means for generating and providing an appropriate answer to the question selected by the user, thereby enabling the user to efficiently obtain information and quickly resolve their question.
[0087] "Data collection means" refers to devices and systems that include sensors and cameras for detecting user behavior.
[0088] "Communication means" refers to the technology or mechanism for transmitting detected user behavior data to a cloud server.
[0089] "Data analysis means" refers to software or algorithms that analyze data sent to the cloud server and generate predicted questions.
[0090] The "display means" is a device or interface that displays the generated questions to the user and allows the user to select them with their line of sight.
[0091] The "answer generation means" is software or a device that generates and provides an appropriate answer to a question selected by a user.
[0092] A "sensor" is a device that detects a user's actions and gaze, and can sense changes in movement and gaze.
[0093] A "camera" is a photographing device for visually capturing a user's line of sight and actions.
[0094] A "cloud server" is a server that can be accessed remotely for data analysis and management.
[0095] "Eye tracking" is a technology that tracks the movement of a user's eyes in real time.
[0096] "Natural language processing technology" is a technology for analyzing text data and understanding and generating human language.
[0097] "Database searching" is the technique or method of searching through a database to retrieve specific information.
[0098] This invention provides a system that efficiently supports users in acquiring information by detecting user behavior, generating predicted questions, and providing appropriate answers to those questions. This system uses sensors and cameras to capture the user's behavior and gaze, analyzes the data using a cloud server, generates appropriate questions, displays them to the user, and generates and provides answers to the selected questions.
[0099] 1. Hardware
[0100] Terminal
[0101] Use smart glasses, which have built-in sensors and cameras to detect the user's gaze and behavior, such as devices like Tobii Pro Glasses 3.
[0102] server
[0103] Use a server on the cloud. The server runs on a cloud service such as an AWS EC2 instance.
[0104] 2. Software
[0105] Data collection
[0106] An application is installed on the device to collect data in real time and send it to a cloud server. This application is developed using Unity, OpenCV, etc.
[0107] Data analysis
[0108] The cloud server uses programs such as Python, TensorFlow, and PyTorch to analyze the received data, which then generates predicted questions based on the user's behavioral data.
[0109] Display
[0110] The device displays the generated question list on its screen, and AR technology is used as the interface for users to select questions with their gaze.
[0111] Answer generation
[0112] To generate answers to selected questions, generative AI models and database search techniques are used. Specific examples include the transformers library and Elasticsearch. The generated answers are sent to the device and presented visually or using speech synthesis techniques. The speech synthesis uses the Google Text-to-Speech API.
[0113] 3. Specific Examples
[0114] Example 1: Use at a restaurant
[0115] User: Looking at a new menu at a restaurant.
[0116] Device: Captures menus with the built-in camera and tracks the user's gaze.
[0117] Server: Recognizes when the user is looking at the menu and generates relevant question candidates (e.g., "What are the ingredients?", "How many calories?").
[0118] Device: Candidate questions are displayed on the screen, and the user selects the question "How many calories?" with their gaze.
[0119] Server: Retrieves the calorie information for the menu items from the database and generates the answer.
[0120] Terminal: Provides answers visually or audibly.
[0121] Example prompt sentence:
[0122] "If a user is looking at a new menu item at a restaurant, how does the system pose questions to the user and provide answers?"
[0123] Example 2: Use during learning
[0124] User: Studying by reading a textbook.
[0125] Device: Built-in camera and sensors identify where you are reading and track your gaze.
[0126] Server: Generates relevant question candidates (e.g., "What is the background of this incident?", "Who are the main characters?") based on the content of the textbook.
[0127] Device: Candidate questions are displayed on the screen, and the user selects a question such as "What is the background to this incident?" with their gaze.
[0128] Server: Obtains background information about the incident from textbook content and related materials and generates answers.
[0129] Terminal: Provides answers visually or audibly.
[0130] Example prompt sentence:
[0131] "When a user is studying a textbook, how does the system pose questions to the user and provide answers?"
[0132] The system of the present invention is a powerful support tool for users to quickly obtain appropriate information, and plays an important role in today's information-overloaded society. It streamlines the user's learning process and provides practical solutions to various questions in daily life.
[0133] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0134] Step 1:
[0135] Input: User behavior and gaze data
[0136] Processing: Data Collection
[0137] Output: User behavior and gaze data
[0138] The device uses sensors and cameras built into the smart glasses to capture the user's actions and gaze in real time. Specifically, the camera captures the user's gaze direction, and the sensor detects movement. For example, when a user picks up a book in a bookstore, their movement and gaze are captured. This allows detailed behavioral and gaze data to be collected.
[0139] Step 2:
[0140] Input: User behavior and gaze data
[0141] Action: Send data
[0142] Output: Data sent to the cloud server
[0143] The device temporarily stores the collected behavioral and gaze data in its internal memory and then transmits the data to a cloud server via Wi-Fi or Bluetooth, using the TLS encryption protocol to protect the data transmission. Specifically, the device generates collected data packets and sequentially transmits them to a designated endpoint on the cloud server.
[0144] Step 3:
[0145] Input: Data sent to the cloud server
[0146] Processing: Data analysis
[0147] Output: Generated question list
[0148] The server receives the data sent to the cloud server and analyzes it using programs such as Python and TensorFlow. Specifically, it analyzes the user's behavioral and gaze data to infer the user's intentions and interests. Based on this, it generates predicted questions. This analysis process uses natural language processing (NLP) technology and machine learning models. For example, if a user is interested in a particular book in a bookstore, it generates questions such as, "What are the reviews for this book?" and "Who is the author?"
[0149] Step 4:
[0150] Input: Generated question list
[0151] Action: Display Question
[0152] Output: The question displayed to the user
[0153] The server sends the generated question list to the device. The device displays the received question list on its display. The user selects a question using eye movements or a pointer on the interface. Specifically, the system uses AR technology to visually display the question list, allowing the user to visually select a question.
[0154] Step 5:
[0155] Input: User selected question
[0156] Process: Send selected data
[0157] Output: Selection data sent to the cloud server
[0158] The user selects a question displayed on the display by looking at it. The device detects this selection and sends the information to the cloud server. The selected question is identified by gaze detection, and the data is sent back to the cloud server.
[0159] Step 6:
[0160] Input: Selection data sent to the cloud server
[0161] Processing: Answer generation
[0162] Output: The generated answer
[0163] The server analyzes the question selected by the user and generates an appropriate answer. This is done using database search and generative AI models (e.g., the transformers library). Specifically, Elasticsearch is used to retrieve the necessary information from the database, and an NLP model generates an answer based on that information. For example, in response to the question, "What are the reviews of this book?", review information is retrieved from a book review database and a concise review is generated.
[0164] Step 7:
[0165] Input: Generated answer
[0166] Action: Provide a response
[0167] Output: The answer provided to the user
[0168] The server sends the generated answer to the device. The device then displays the received answer to the user visually on a display or provides it audibly using speech synthesis technology. Specifically, the device uses the Google Text-to-Speech API to synthesize speech and provide the answer to the user audibly. The answer is also displayed visually on the smart glasses display. For example, the answer might be, "This book has excellent reviews, especially its storytelling."
[0169] (Application example 1)
[0170] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0171] In today's brick-and-mortar stores, users are required to quickly and efficiently obtain the information they need from a large selection of products. In particular, if users are unable to easily obtain detailed product information and reviews when choosing a product, their motivation to purchase is likely to decrease. Furthermore, if the information provided by the store is insufficient, it can affect users' decision-making, resulting in a negative impact on sales. To solve these problems, a system is needed that allows users to obtain product information in real time.
[0172] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0173] In this invention, the server includes a means for capturing gazes using a camera and a sensor built into the smart glasses in a physical store and analyzing the gazes on the cloud server to provide information about products, a means for displaying a list of questions received from the cloud server on the display of the smart glasses and selecting an appropriate question based on the user's gaze, and a means for the smart glasses to retrieve information about the selected question from a database and provide it to the user using voice synthesis technology or a display, thereby enabling the user to quickly and efficiently obtain detailed information about products.
[0174] A "sensor" is a device that detects physical changes and converts them into electrical signals.
[0175] A "camera" is a device for capturing images or video.
[0176] The "data collection means" is the part of the system that includes sensors and cameras to detect user activity.
[0177] The "data analysis means" is the part of the system that analyzes collected user behavior data and generates predicted questions.
[0178] The "display means" is a device that displays the generated questions to the user and allows the user to select one by line of sight.
[0179] The "answer generator" is the part of the system that generates and provides appropriate answers to user-selected questions.
[0180] A "physical store" is a physical store that a user visits to purchase products.
[0181] "Smart glasses" are eyeglass-type devices that display information and have built-in cameras and sensors to track the user's gaze.
[0182] A "cloud server" is a server for processing and storing data remotely via the Internet.
[0183] "Gaze capture" means tracking the user's gaze in real time and acquiring the data.
[0184] A "question list" is a collection of multiple questions generated based on user behavior data.
[0185] A "database" is a system that systematically manages data and allows for quick retrieval of necessary information.
[0186] A "display" is a device for visually displaying information.
[0187] "Speech synthesis technology" is a technology that converts text data into voice data.
[0188] This invention is a system that allows users to quickly and efficiently obtain detailed product information in physical stores. The system uses smart glasses to track the user's gaze and provides relevant information via a cloud server. To implement this invention, smart glasses, a cloud server, a database, and voice synthesis technology are used in combination.
[0189] Hardware and software used
[0190] Hardware:
[0191] Smart glasses: Equipped with a built-in camera and sensors, they capture the user's gaze in real time and have communication capabilities to send gaze data to a cloud server.
[0192] Display: Built into the smart glasses, it visually displays questions and answers to the user.
[0193] software:
[0194] Cloud server: Receives the user's gaze data, analyzes the data, generates necessary questions, and retrieves appropriate information from the database to display to the user.
[0195] Data analysis method: Runs on a cloud server and is responsible for analyzing gaze data and generating questions.
[0196] Question list generator: Generates appropriate questions based on the user's gaze information and sends this question list to the smart glasses.
[0197] Answer generation method: Based on the question selected by the user, information is retrieved from the database and an answer is generated to provide to the user. Natural language processing techniques (e.g., OpenAI GPT model) are used.
[0198] Speech synthesis technology: Technology for providing answers to users via voice (e.g., Google TTS).
[0199] What the system does
[0200] Overview of the process flow:
[0201] 1. Smart Glasses:
[0202] It uses built-in cameras and sensors to capture the user's gaze in real time.
[0203] The acquired gaze data is sent to a cloud server.
[0204] 2. Cloud Server:
[0205] The received gaze data is analyzed to identify information about the product that the user is paying attention to.
[0206] A list of relevant questions is generated and sent to the smart glasses.
[0207] 3. Smart Glasses:
[0208] The received question list is displayed on a display, and the user selects a question with their line of sight.
[0209] The question selected by the user is sent to the cloud server.
[0210] 4. Cloud Server:
[0211] Retrieve answers to selected questions from a database.
[0212] Based on the acquired information, an appropriate answer is generated and sent to the smart glasses.
[0213] 5. Smart Glasses:
[0214] The received answer is either visually displayed on a display or provided aloud using speech synthesis technology.
[0215] Specific examples
[0216] Suppose a user is looking at a particular product (e.g., organic food) in a physical store (e.g., a supermarket). The user's gaze data is captured by the smart glasses and sent to a cloud server. The cloud server analyzes the gaze data to identify the product the user is looking at and generates a list of questions related to this product. Questions such as "What are the nutritional content of this product?" and "What is the best way to cook it?" are displayed on the display of the user's smart glasses. When the user selects a question with their gaze, the selected question is sent to the cloud server. The cloud server retrieves an appropriate answer to the selected question from a database and sends it to the user's smart glasses. Finally, the smart glasses either display the retrieved answer visually or provide the answer audibly using speech synthesis technology.
[0217] Prompt Sentence Examples
[0218] Describe the process flow for a smart glasses application that tracks a user's gaze in real time and generates and answers questions related to specific supermarket items. Include specific details about the hardware, software, and data processing methods used.
[0219] In this way, the present invention allows users to quickly and efficiently obtain detailed information about products in a physical store.
[0220] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0221] Step 1:
[0222] The smart glasses' cameras and sensors capture the user's gaze.
[0223] Input: User gaze data
[0224] How it works: The smart glasses' built-in cameras and sensors track the user's gaze in real time and capture data to identify the object they are looking at.
[0225] Output: Captured gaze data
[0226] Step 2:
[0227] The smart glasses transmit the gaze data to a cloud server.
[0228] Input: Captured gaze data
[0229] How it works: The smart glasses transmit the acquired gaze data to a cloud server via the internet.
[0230] Output: Gaze data received by the cloud server
[0231] Step 3:
[0232] The cloud server analyzes the received gaze data and identifies information about the product that the user is paying attention to.
[0233] Input: Gaze data received by the cloud server
[0234] Operation: The data analysis means on the cloud server analyzes the gaze data and performs calculations to identify the object (product) that the gaze is directed at.
[0235] Output: Identified product information (product ID, etc.)
[0236] Step 4:
[0237] The cloud server generates a list of related questions based on the identified products.
[0238] Input: Identified product information
[0239] How it works: A data analysis tool on a cloud server uses a generative AI model to generate questions related to the identified products. These questions are formulated based on product characteristics.
[0240] Output: Generated question list
[0241] Step 5:
[0242] The smart glasses display the received question list on the display.
[0243] Input: Generated question list
[0244] How it works: A list of questions sent from a cloud server is displayed on the smart glasses, allowing the user to select a question with their gaze.
[0245] Output: Question list displayed on the screen
[0246] Step 6:
[0247] The user selects a question with their gaze.
[0248] Input: Question list displayed on the screen
[0249] How it works: The user uses their gaze to select a question of interest. The smart glasses sense the selected question.
[0250] Output: Selected Question
[0251] Step 7:
[0252] The smart glasses send the selected question to a cloud server.
[0253] Input: Selected Question
[0254] How it works: The smart glasses send the selected question over the internet to a cloud server.
[0255] Output: Questions received by the cloud server
[0256] Step 8:
[0257] The cloud server retrieves answers to the received questions from the database.
[0258] Input: A question received by the cloud server
[0259] Operation: The answer generation means on the cloud server searches the database for relevant information and processes the data to generate an answer.
[0260] Output: Retrieved answer information
[0261] Step 9:
[0262] The cloud server transmits the acquired answer information to the smart glasses.
[0263] Input: Retrieved answer information
[0264] How it works: The cloud server sends the generated answer information to the smart glasses via the internet.
[0265] Output: Answer information received by smart glasses
[0266] Step 10:
[0267] The smart glasses provide the received answer information to the user.
[0268] Input: Answer information received by smart glasses
[0269] How it works: The smart glasses will either visually display the answer information on a display or provide it aloud using speech synthesis technology.
[0270] Output: Answer information provided to the user
[0271] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0272] This invention provides a system that efficiently supports users in acquiring information by detecting a user's behavior, gaze, and even emotions, generating predicted questions, and providing appropriate answers to those questions. This system uses sensors and a camera to capture the user's behavior and gaze, recognizes the user's emotions using an emotion engine, analyzes the data using a cloud server, generates appropriate questions, displays them to the user, and generates and provides answers to the selected questions.
[0273] Program processing flow
[0274] This system operates based on the interaction between the "terminal" (smart glasses), the "server" on the cloud, and the "user." In addition, an emotion engine is built in to recognize the user's emotions.
[0275] Device: The device uses built-in sensors and cameras to capture the user's movements and gaze in real time. The sensors detect the user's movements (e.g., walking, standing, sitting, etc.), and the camera tracks the user's gaze direction and facial expressions. The captured data is temporarily stored on the device, and the emotion engine analyzes the user's facial expressions and voice to recognize emotions.
[0276] Server: The cloud server receives the data sent from the device and analyzes it in multiple ways. In particular, it uses data analysis tools to analyze the user's current behavior, gaze, and emotions to predict questions the user is likely to have in that situation. It then generates a list of predicted questions. The generated list of questions is sent from the cloud server to the device.
[0277] Terminal: The terminal displays the question list received from the server on a display, and the displayed questions are displayed so that they can be visually recognized within the user's field of vision.
[0278] User: The user selects a question displayed on the display by gazing at it. For example, if the user fixates their gaze on the question "How many calories does it have?" for a certain period of time, that question is recognized as selected.
[0279] Device: Sends the selected question to the server. The sent data includes the selected question, as well as the user's current situation and context information (including emotions).
[0280] Server: The server analyzes the received question and generates an appropriate answer. It searches for relevant information from a database and uses natural language processing techniques to create an answer in an easy-to-read format. It also adjusts the answer based on the emotions recognized by the emotion engine.
[0281] Terminal: Provides the received answer to the user. The answer is displayed visually on the display or provided aloud using speech synthesis technology. The answer is provided in an appropriate expression according to the user's emotion.
[0282] Specific examples
[0283] Example 1: Use at a restaurant
[0284] User: Looking at a new menu at a restaurant.
[0285] Device: Captures menus with the built-in camera and tracks the user's gaze and facial expressions.
[0286] Server: Recognizes when the user is looking at the menu and generates relevant question candidates (e.g., "What are the ingredients?", "How many calories?"). If the user's facial expression indicates interest, it prioritizes displaying relevant information.
[0287] Device: Candidate questions are displayed on the screen, and the user selects a question such as "How many calories?" with their gaze.
[0288] Server: Retrieves the calorie information of the menu item from the database and generates an answer. If the user looks surprised, the answer is provided in a softer tone.
[0289] Terminal: Provides answers visually or audibly.
[0290] Example 2: Use during learning
[0291] User: Studying by reading a textbook.
[0292] On the device: Built-in cameras and sensors identify where the user is reading, track their gaze and facial expressions, and an emotion engine recognizes the user's emotions (e.g., confusion, interest, etc.).
[0293] Server: Generates relevant question candidates (e.g., "What is the background of this incident?", "Who are the main characters?") based on the textbook content. If it determines that the user is confused, it prioritizes question candidates that include explanations.
[0294] Device: Candidate questions are displayed on the screen, and the user selects a question such as "What is the background to this incident?" with their gaze.
[0295] Server: Retrieves background information from textbook content and related materials and generates an answer. If the user is dissatisfied, the answer is expanded.
[0296] Terminal: Provides answers visually or audibly.
[0297] The system of this invention optimizes the provision of information by capturing the user's behavior, gaze, and emotions from multiple angles, and supports the user in resolving questions in their studies and daily life. By providing information according to the user's emotions, more effective learning and information acquisition are realized.
[0298] The processing flow will be explained below.
[0299] Step 1:
[0300] Device: The device uses built-in sensors and cameras to capture the user's movements, gaze, and facial expressions in real time. The sensors detect the user's movements (e.g., walking, standing, sitting, etc.), and the camera tracks the user's gaze direction and facial expressions. The captured data is temporarily stored on the device.
[0301] Step 2:
[0302] Device: Processes the collected behavior, gaze, and facial expression data, extracts necessary information, and sends it to the cloud server. This transmitted data includes the user's current behavioral patterns, gaze information, and facial expression data.
[0303] Step 3:
[0304] Server: The cloud server receives the transmitted data and analyzes it using data analysis methods and an emotion engine. In particular, it determines what information the user is paying attention to and what emotion (interest, confusion, delight, etc.) the user is feeling.
[0305] Step 4:
[0306] Server: Based on the analysis results, predicts the questions the user is likely to have in that situation. Generates a list of predicted questions that also reflects the user's emotional state. The generated list of questions is sent from the cloud server to the device.
[0307] Step 5:
[0308] Terminal: The received question list is displayed on the display. The displayed questions are displayed so that they can be visually recognized within the user's field of vision.
[0309] Step 6:
[0310] User: Selects a question displayed on the display by gazing at it. For example, if a user fixates their gaze on the question "How many calories does it have?" for a certain period of time, that question is recognized as selected.
[0311] Step 7:
[0312] The terminal sends the selected question to the server. The sent data includes the selected question, as well as the user's current situation and emotional information.
[0313] Step 8:
[0314] Server: The server analyzes the received question and generates an appropriate answer. It searches for relevant information from a database and creates an answer using natural language processing technology. It also adjusts the expression and level of detail of the answer based on the emotions recognized by the emotion engine.
[0315] Step 9:
[0316] Server: Sends the generated answer to the terminal, which may be presented as a visual display or audio.
[0317] Step 10:
[0318] Terminal: Provides the received answer to the user. The answer is displayed visually on the display or provided aloud using speech synthesis technology. The answer is provided in an appropriate expression according to the user's emotions.
[0319] Step 11:
[0320] Server: User responses and selection patterns are logged and used for later analysis, forming a feedback loop to improve the accuracy of future predictions.
[0321] These are the specific processing steps of this system. By comprehensively capturing the user's behavior, gaze, and emotions, and providing predicted questions and answers based on this information in real time, the system streamlines the user's learning process and information acquisition.
[0322] Example 2
[0323] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0324] Conventional information acquisition systems have had difficulty efficiently providing users with the information they seek. In particular, they have been unable to sense data such as the user's gaze and emotions in real time, pose appropriate questions, and provide answers tailored to the user. As a result, the user's information acquisition process has been subject to significant time and effort.
[0325] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes a data collection means including a sensor and a camera for detecting the user's movements, gaze data, and emotions, a data analysis means for analyzing the detected user's movement data, gaze data, and emotion data and generating predicted questions, a display means for displaying a question list generated based on the analysis means to the user and allowing the user to select a question with their gaze, an answer generation means for generating and providing an appropriate answer to the question selected by the user, and an emotion analysis means for adjusting the generated answer based on the user's emotion data. This makes it possible to efficiently obtain the information the user desires.
[0326] A "data collection means" is a device that includes sensors and cameras to detect the user's movements, gaze, and emotions.
[0327] The "data analysis means" is a device or software that analyzes the detected user's motion data, gaze data, and emotion data and generates predicted questions.
[0328] The "display means" is a device or software that displays the question list generated based on the analysis means to the user and allows the user to select a question by line of sight.
[0329] The "answer generation means" is a device or software that generates and provides an appropriate answer to a question selected by a user.
[0330] The "emotion analysis means" is a device or software that adjusts the response generated based on the user's emotional data.
[0331] This invention is a system that efficiently supports users in acquiring information by detecting a user's movements, gaze, and emotions, generating predicted questions, and providing appropriate answers to those questions. This system uses sensors and a camera to capture the user's movements and gaze, recognizes the user's emotions using an emotion engine, analyzes the data using a cloud server, generates appropriate questions and displays them to the user, and generates and provides answers to the selected questions.
[0332] The following components are primarily involved in the operation of the system:
[0333] Device: Uses a device (such as smart glasses) equipped with sensors and a camera to capture the user's movements, gaze, and emotions in real time. The sensor tracks the user's movements, and the camera tracks the direction of gaze and facial expressions, and this data is temporarily stored on the device. The device is also equipped with an emotion engine that analyzes the user's facial expressions and voice to recognize emotions.
[0334] Server: The cloud server receives the data sent from the device and uses data analysis methods to analyze the user's current behavior, gaze, and emotions. This analysis predicts the questions the user is likely to have in that situation and generates a list of questions. This list of questions is then sent from the cloud server to the device.
[0335] Display method: The device displays the question list received from the server on the display. The display format is an interactive format using pop-ups or eye tracking. For example, when a user is looking at a new menu, questions such as "How many calories are in this dish?" and "What are the ingredients?" are displayed.
[0336] User: The user selects a question by gazing at it. By fixing their gaze for a certain period of time, the system recognizes the user's selection and determines that the question has been selected.
[0337] Answer generation method: The selected question and context information (user's facial expression and current situation information) are sent from the device to the server. The server analyzes the question, searches for relevant information from a database, and generates an answer using natural language processing technology. It also adjusts the answer based on an emotion engine, for example, providing a softer tone if the user is surprised.
[0338] Providing an answer: The device provides the received answer to the user either visually on the display or audibly using speech synthesis technology. For example, it may visually display or audibly guide the user by saying, "This dish is 500 kcal."
[0339] Specific examples
[0340] Example 1: Use at a restaurant
[0341] 1. User: Looking at a new menu at a restaurant.
[0342] 2. Device: Capture menus with the built-in camera and track the user's gaze and facial expressions.
[0343] 3. Server: Recognizes that the user is looking at the menu and generates relevant question candidates ("What are the ingredients?", "How many calories?").
[0344] 4. Device: Question candidates are displayed on the screen, and the user selects a question such as "How many calories?" with their gaze.
[0345] 5. Server: Retrieves the calorie information of the menu item from the database and generates an answer. If the user looks surprised, the answer is provided in a softer tone.
[0346] 6. Terminal: Provides answers visually or audibly.
[0347] Example 2: Use during learning
[0348] 1. User: Reading and studying a textbook.
[0349] 2. Device: The built-in camera and sensors identify where the user is reading, track their gaze and facial expressions, and the emotion engine recognizes the user's emotions (e.g., confusion, interest, etc.).
[0350] 3. Server: Generates candidate questions ("What is the background of this incident?", "Who are the main characters?") based on the content of the textbook. If it determines that the user is confused, it prioritizes candidate questions that include explanations.
[0351] 4. Terminal: Candidate questions are displayed on the screen, and the user selects a question such as "What is the background to this incident?" with their gaze.
[0352] 5. Server: Retrieves background information from textbook content and related materials and generates an answer. If the user is dissatisfied, the answer is expanded.
[0353] 6. Terminal: Provides answers visually or audibly.
[0354] Prompt Sentence Examples
[0355] "A user is looking at a new menu item at a restaurant. If the user looks interested, suggest possible questions about the menu."
[0356] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0357] Step 1: Capture data with your device
[0358] Device: Activates built-in sensors and cameras to capture the user's movements, gaze, and facial expressions in real time. Inputs include the user's physical movements, gaze, and facial expressions, which are detected by the sensors and camera and temporarily stored as digital data. Specific movements include detecting walking and sitting, tracking gaze direction, and analyzing facial expressions. Outputs include movement data, gaze data, and facial expression data.
[0359] Step 2: Send data to the server
[0360] Terminal: Sends the captured data to the cloud server. The inputs are the movement data, gaze data, and facial expression data generated in step 1, and are sent to the server in real time. Specific operations include packaging and sending the data. The output is the data sent to the server.
[0361] Step 3: Data analysis and question generation by the server
[0362] Server: Analyzes the received data to determine the user's current behavior, gaze, and emotions. Inputs include motion data, gaze data, and facial expression data sent from the device, which are analyzed using machine learning algorithms and an emotion analysis engine. Specific operations include analyzing behavioral patterns, identifying gaze focus, and recognizing emotions. The output is a list of questions the user is likely to have in the situation.
[0363] Step 4: Sending the Question List from the Server to the Device
[0364] Server: Sends the question list generated by the analysis to the terminal. The input is the question list generated in step 3, which is sent to the terminal. Specifically, the server packages and sends the question list. The output is the question list sent to the terminal.
[0365] Step 5: Ask a question on the device
[0366] Terminal: The terminal displays the question list received from the server. The input is the question list sent from the server, which is displayed in a format that is easy for the user to view. Specific operations include displaying questions in a pop-up or interactive format. The output is the questions displayed on the screen.
[0367] Step 6: User selects question
[0368] User: Selects a question candidate displayed on the display with their gaze. The input is a list of questions displayed on the device, and the user fixates their gaze on a specific question for a certain period of time. Specific operations include processing the gaze tracking data and selecting a question. The output is the selected question.
[0369] Step 7: Sending the selected question to the server
[0370] Terminal: Sends the user's current situation and context information to the server along with the selected question. The inputs are the user's selected question and related context information (motion data, gaze data, and facial expression data), which are then sent to the server. Specific operations include packaging and sending the data. The output is the question and context information sent to the server.
[0371] Step 8: Server Generates Answer
[0372] Server: Analyzes the received question, retrieves relevant information from a database, and generates an answer. The input is the question and context information sent from the device, which is analyzed using natural language processing technology. Specific operations include database search, natural language generation, and response adjustment using an emotion engine. The output is the generated answer.
[0373] Step 9: Provide answers on your device
[0374] Terminal: Provides the received answer to the user. The input is the answer sent from the server, which is displayed on a screen or provided as voice using speech synthesis technology. Specific operations include visual display and speech synthesis. The output is the answer provided to the user.
[0375] (Application example 2)
[0376] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0377] Currently, when a customer has a question about a product in a physical store, it often takes time and effort to find a store employee to resolve the question. It can also be difficult for store staff to respond to all questions quickly and accurately. Furthermore, there is no system that can predict what customers are wondering about and provide information to answer those questions. This can lead to a poor customer experience and affect store sales.
[0378] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0379] In this invention, the server includes a data collection means including a sensor and a camera for detecting user behavior, a data analysis means for analyzing the detected user behavior data and emotion data and generating predicted questions, a display means for displaying the generated questions to the user and allowing the user to select one with their line of sight, and an answer generation means for generating and providing an appropriate answer to the question selected by the user. This makes it possible to provide quick and appropriate answers to questions that users have in physical stores.
[0380] "User" refers to an individual who uses the system.
[0381] An "action" is when a user performs a specific physical action or visual attention.
[0382] A "sensor" is a device that senses physical movements and environmental information and acquires data based on that information.
[0383] A "camera" is an optical device for capturing images or videos.
[0384] The "data collection means" is a means for collecting user behavioral data and emotional data using sensors and cameras.
[0385] "Emotion data" refers to emotional information sensed from the user's facial expressions, voice, etc.
[0386] "Data analysis means" refers to means for analyzing collected behavioral and emotional data and generating predicted questions.
[0387] The "display means" is a device that displays the generated questions to the user and allows the user to select one by line of sight.
[0388] The "answer generation means" is a means for generating and providing an appropriate answer to a question selected by a user.
[0389] "Natural language processing technology" is a technology that enables computers to understand and generate natural language.
[0390] A "generative AI model" is an artificial intelligence model that generates new information based on previously learned data.
[0391] The present invention provides a system that detects a user's behavior, gaze, and emotions, generates predicted questions based on the detected behavior, and provides appropriate answers to the questions. This system includes a data collection means, a data analysis means, a display means, and an answer generation means.
[0392] The system is implemented using the following hardware and software.
[0393] Hardware:
[0394] Smart glasses: Equipped with a built-in camera and display, they capture the user's gaze and facial expressions in real time.
[0395] Sensor: A device that detects user movements, such as an accelerometer or gyro sensor.
[0396] software:
[0397] OpenCV: A library for capturing and processing video data from cameras.
[0398] EmotionEngine: A system for analyzing emotions from a user's facial expressions and voice.
[0399] Cloud server: A server for data analysis and answer generation, utilizing natural language processing technology and generative AI models.
[0400] Natural language processing techniques: For example, using advanced language models such as GPT-3.
[0401] Process flow:
[0402] 1. Data collection methods:
[0403] The smart glasses' built-in cameras and sensors collect user behavioral and emotional data.
[0404] 2. Data analysis methods:
[0405] The collected data is analyzed on the cloud server to understand the user's current behavior and emotions and generate relevant questions.
[0406] 3. Display means:
[0407] The generated question list is displayed on the smart glasses display, and the user can select a question by eye gaze.
[0408] 4. Answer generation means:
[0409] Appropriate answers to questions selected by the user are generated on a cloud server and provided to the smart glasses either displayed or via voice.
[0410] Examples:
[0411] Example 1: Supermarket use:
[0412] User: Looking for food in the supermarket.
[0413] Smart glasses: Use a camera to track the user's gaze and analyze emotions.
[0414] Cloud server: Generates food-related question candidates (e.g., "What is the allergy information for this ingredient?").
[0415] Smart glasses: A list of questions is displayed and the user selects one by looking at the screen.
[0416] Cloud server: Generates answers to selected questions and displays them to the user.
[0417] Example prompt sentence:
[0418] "Please provide information about the ingredients of foods that may be of interest to users."
[0419] User gaze information: Fresh vegetable section
[0420] User sentiment: Showing interest
[0421] Questions generated: "What is the nutritional value of this vegetable?", "What dishes can it be used in?", "What is the allergy information?"
[0422] In this way, by using the system of the present invention, users can smoothly resolve their questions in physical stores, thereby improving the customer experience.
[0423] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0424] Step 1:
[0425] The device uses a built-in camera and sensors to capture the user's actions and gaze. The camera captures the user's gaze direction and facial expressions in real time, and the sensor detects the user's actions (e.g., reaching for a product). This captured data is collected by a data collection means. The input is the user's physical actions and gaze information, and the output is the captured action and gaze data.
[0426] Step 2:
[0427] The device sends the collected behavioral data and gaze data to a cloud server. The cloud server receives this data and analyzes it using data analysis means. The data analysis means analyzes the user's emotions from their gaze direction, movements, and facial expressions, and predicts questions that are likely to interest the user. For example, if the user is looking at a particular food item, it generates a question related to that food (e.g., "What are the ingredients in this food?"). The input is behavioral data and gaze data, and the output is a list of predicted questions.
[0428] Step 3:
[0429] The server generates questions based on the analysis results and sends them to the device in the form of a list. The device displays the received question list on the smart glasses display. The user visually checks the displayed question list and fixes their gaze on questions that interest them. The input is the predicted question list, and the output is the displayed question list.
[0430] Step 4:
[0431] When a user fixates their gaze on a question, the device detects the gaze information and determines that the user has selected the question. The selected question is then sent back to the cloud server, which then generates an appropriate answer for the question. The input is the question selected by the user, and the output is the generated answer.
[0432] Step 5:
[0433] The cloud server uses natural language processing technology and generative AI models (e.g., GPT-3) to generate answers to questions selected by the user. This answer is generated by searching for relevant information from a database and formatting it in an appropriate linguistic format. Furthermore, it takes into account the user's emotional information and generates an answer in a tone appropriate to the emotion. The input is the question selected by the user and emotional data, and the output is the generated answer.
[0434] Step 6:
[0435] The device receives the answer from the cloud server and provides it to the user either visually or audibly using speech synthesis technology. This allows the user to quickly obtain an answer to their question. The input is the generated answer, and the output is the answer provided to the user.
[0436] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0437] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0438] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0439] [Second embodiment]
[0440] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0441] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0442] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0443] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0444] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0445] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0446] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0447] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0448] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0449] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0450] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0451] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0452] This invention provides a system that efficiently supports users in acquiring information by detecting user behavior, generating predicted questions, and providing appropriate answers to those questions. This system uses sensors and cameras to capture the user's behavior and gaze, analyzes the data using a cloud server, generates appropriate questions, displays them to the user, and generates and provides answers to the selected questions.
[0453] Program processing flow
[0454] This system operates based on the interaction between the "terminal" (smart glasses), the "server" on the cloud, and the "user."
[0455] Device: Using built-in sensors and cameras, the device captures the user's movements and gaze in real time. The device temporarily stores the collected data and sends it to a cloud server as needed.
[0456] Server: The cloud server receives and analyzes the user data sent from the device. The data analysis method generates predicted questions based on the user's current situation. The generated question list is then sent from the server to the device.
[0457] Terminal: The terminal displays the question list received from the server on the display. When the user selects a question candidate on the display with their gaze, the selected question is sent to the server.
[0458] Server: The server analyzes the question sent by the user and generates an appropriate answer using database search and natural language processing techniques. The generated answer is then sent to the device.
[0459] Terminal: The terminal provides the received answer to the user either visually on a display or audibly using speech synthesis technology.
[0460] Specific examples
[0461] Example 1: Use at a restaurant
[0462] User: Looking at a new menu at a restaurant.
[0463] Device: Captures menus with the built-in camera and tracks the user's gaze.
[0464] Server: Recognizes when the user is looking at the menu and generates relevant question suggestions (e.g., "What are the ingredients?", "How many calories?").
[0465] Device: The question candidates are displayed on the screen and the user selects one. The user selects a question such as "How many calories?"
[0466] Server: Retrieves the calorie information for the menu items from the database and generates the answer.
[0467] Terminal: Provides answers visually or audibly.
[0468] Example 2: Use during learning
[0469] User: Studying by reading a textbook.
[0470] Device: Built-in camera and sensors identify where you are reading and track your gaze.
[0471] Server: Generates relevant question candidates (e.g., "What is the background to this incident?", "Who are the main characters?") based on the content of the textbook.
[0472] Terminal: The question candidates are displayed on the screen and the user can select one. The user selects the question "What is the background of this incident?"
[0473] Server: Obtains background information about the incident from textbook content and related materials and generates answers.
[0474] Terminal: Provides answers visually or audibly.
[0475] The system of the present invention is a powerful support tool for users to quickly obtain appropriate information, and plays an important role in today's information-overloaded society. It streamlines the user's learning process and provides practical solutions to various questions in daily life.
[0476] The processing flow will be explained below.
[0477] Step 1:
[0478] Device: The device uses built-in sensors and cameras to capture the user's movements and gaze in real time. The sensors detect the user's movements (e.g., walking, standing, sitting, etc.), and the camera tracks the user's gaze direction. The captured data is temporarily stored on the device.
[0479] Step 2:
[0480] Terminal: Processes the collected data, extracts necessary information, and sends it to the cloud server. This transmitted data includes the user's current behavioral patterns and gaze information.
[0481] Step 3:
[0482] Server: The cloud server receives the transmitted data and uses data analysis methods to analyze the user's behavior and gaze data. Specifically, it predicts what information the user is paying attention to and what questions they are likely to have.
[0483] Step 4:
[0484] Server: Based on the analysis results, a list of predicted questions is generated. The list of questions includes questions that the user is likely to have in that situation (e.g., "What are the ingredients in this dish?"). The generated list of questions is then sent back to the device.
[0485] Step 5:
[0486] Terminal: The received question list is displayed on the display. The displayed questions are displayed so that they can be visually recognized within the user's field of vision.
[0487] Step 6:
[0488] User: Selects a question displayed on the display by gazing at it. For example, if a user fixates their gaze on the question "How many calories does it have?" for a certain period of time, that question is recognized as selected.
[0489] Step 7:
[0490] The terminal sends the selected question to the server. The sent data includes the selected question as well as the user's current situation and context information.
[0491] Step 8:
[0492] Server: The server analyzes the received question and generates an appropriate answer. It searches for relevant information from a database and uses natural language processing techniques to create an answer in an easy-to-read format.
[0493] Step 9:
[0494] Server: Sends the generated answer to the terminal, which may be presented visually or audibly.
[0495] Step 10:
[0496] Terminal: Provides the received answer to the user, either visually displaying the answer on a display or audibly using speech synthesis technology.
[0497] These are the specific processing steps of the "Chatty Glasses" system. By capturing the user's behavior and gaze in real time, and using that information to predict and generate appropriate questions and provide answers, the system helps users obtain information more efficiently.
[0498] Example 1
[0499] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0500] In modern society, it is important for users to quickly obtain appropriate information for efficient decision-making and learning. However, it is not easy to find the information they need from a vast amount of information. There is also a lack of mechanisms that can accurately identify when a user has a question and quickly provide an appropriate answer to that question. The present invention aims to solve these problems by providing a system that improves the efficiency of information acquisition and supports users in resolving their questions.
[0501] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0502] In this invention, the server includes a data collection means including a sensor and a camera for detecting user behavior, a communication means for transmitting the detected user behavior data to the cloud server, a data analysis means for analyzing the transmitted data and generating predicted questions, a display means for displaying the generated question list to the user and allowing the user to select one with their line of sight, and an answer generation means for generating and providing an appropriate answer to the question selected by the user, thereby enabling the user to efficiently obtain information and quickly resolve their question.
[0503] "Data collection means" refers to devices and systems that include sensors and cameras for detecting user behavior.
[0504] "Communication means" refers to the technology or mechanism for transmitting detected user behavior data to a cloud server.
[0505] "Data analysis means" refers to software or algorithms that analyze data sent to the cloud server and generate predicted questions.
[0506] The "display means" is a device or interface that displays the generated questions to the user and allows the user to select them with their line of sight.
[0507] The "answer generation means" is software or a device that generates and provides an appropriate answer to a question selected by a user.
[0508] A "sensor" is a device that detects a user's actions and gaze, and can sense changes in movement and gaze.
[0509] A "camera" is a photographing device for visually capturing a user's line of sight and actions.
[0510] A "cloud server" is a server that can be accessed remotely for data analysis and management.
[0511] "Eye tracking" is a technology that tracks the movement of a user's eyes in real time.
[0512] "Natural language processing technology" is a technology for analyzing text data and understanding and generating human language.
[0513] "Database searching" is the technique or method of searching through a database to retrieve specific information.
[0514] This invention provides a system that efficiently supports users in acquiring information by detecting user behavior, generating predicted questions, and providing appropriate answers to those questions. This system uses sensors and cameras to capture the user's behavior and gaze, analyzes the data using a cloud server, generates appropriate questions, displays them to the user, and generates and provides answers to the selected questions.
[0515] 1. Hardware
[0516] Terminal
[0517] Use smart glasses, which have built-in sensors and cameras to detect the user's gaze and behavior, such as devices like Tobii Pro Glasses 3.
[0518] server
[0519] Use a server on the cloud. The server runs on a cloud service such as an AWS EC2 instance.
[0520] 2. Software
[0521] Data collection
[0522] An application is installed on the device to collect data in real time and send it to a cloud server. This application is developed using Unity, OpenCV, etc.
[0523] Data analysis
[0524] The cloud server uses programs such as Python, TensorFlow, and PyTorch to analyze the received data, which then generates predicted questions based on the user's behavioral data.
[0525] Display
[0526] The device displays the generated question list on its screen, and AR technology is used as the interface for users to select questions with their gaze.
[0527] Answer generation
[0528] To generate answers to selected questions, generative AI models and database search techniques are used. Specific examples include the transformers library and Elasticsearch. The generated answers are sent to the device and presented visually or using speech synthesis techniques. The speech synthesis uses the Google Text-to-Speech API.
[0529] 3. Specific Examples
[0530] Example 1: Use at a restaurant
[0531] User: Looking at a new menu at a restaurant.
[0532] Device: Captures menus with the built-in camera and tracks the user's gaze.
[0533] Server: Recognizes when the user is looking at the menu and generates relevant question candidates (e.g., "What are the ingredients?", "How many calories?").
[0534] Device: Candidate questions are displayed on the screen, and the user selects the question "How many calories?" with their gaze.
[0535] Server: Retrieves the calorie information for the menu items from the database and generates the answer.
[0536] Terminal: Provides answers visually or audibly.
[0537] Example prompt sentence:
[0538] "If a user is looking at a new menu item at a restaurant, how does the system pose questions to the user and provide answers?"
[0539] Example 2: Use during learning
[0540] User: Studying by reading a textbook.
[0541] Device: Built-in camera and sensors identify where you are reading and track your gaze.
[0542] Server: Generates relevant question candidates (e.g., "What is the background of this incident?", "Who are the main characters?") based on the content of the textbook.
[0543] Device: Candidate questions are displayed on the screen, and the user selects a question such as "What is the background to this incident?" with their gaze.
[0544] Server: Obtains background information about the incident from textbook content and related materials and generates answers.
[0545] Terminal: Provides answers visually or audibly.
[0546] Example prompt sentence:
[0547] "When a user is studying a textbook, how does the system pose questions to the user and provide answers?"
[0548] The system of the present invention is a powerful support tool for users to quickly obtain appropriate information, and plays an important role in today's information-overloaded society. It streamlines the user's learning process and provides practical solutions to various questions in daily life.
[0549] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0550] Step 1:
[0551] Input: User behavior and gaze data
[0552] Processing: Data Collection
[0553] Output: User behavior and gaze data
[0554] The device uses sensors and cameras built into the smart glasses to capture the user's actions and gaze in real time. Specifically, the camera captures the user's gaze direction, and the sensor detects movement. For example, when a user picks up a book in a bookstore, their movement and gaze are captured. This allows detailed behavioral and gaze data to be collected.
[0555] Step 2:
[0556] Input: User behavior and gaze data
[0557] Action: Send data
[0558] Output: Data sent to the cloud server
[0559] The device temporarily stores the collected behavioral and gaze data in its internal memory and then transmits the data to a cloud server via Wi-Fi or Bluetooth, using the TLS encryption protocol to protect the data transmission. Specifically, the device generates collected data packets and sequentially transmits them to a designated endpoint on the cloud server.
[0560] Step 3:
[0561] Input: Data sent to the cloud server
[0562] Processing: Data analysis
[0563] Output: Generated question list
[0564] The server receives the data sent to the cloud server and analyzes it using programs such as Python and TensorFlow. Specifically, it analyzes the user's behavioral and gaze data to infer the user's intentions and interests. Based on this, it generates predicted questions. This analysis process uses natural language processing (NLP) technology and machine learning models. For example, if a user is interested in a particular book in a bookstore, it generates questions such as, "What are the reviews for this book?" and "Who is the author?"
[0565] Step 4:
[0566] Input: Generated question list
[0567] Action: Display Question
[0568] Output: The question displayed to the user
[0569] The server sends the generated question list to the device. The device displays the received question list on its display. The user selects a question using eye movements or a pointer on the interface. Specifically, the system uses AR technology to visually display the question list, allowing the user to visually select a question.
[0570] Step 5:
[0571] Input: User selected question
[0572] Process: Send selected data
[0573] Output: Selection data sent to the cloud server
[0574] The user selects a question displayed on the display by looking at it. The device detects this selection and sends the information to the cloud server. The selected question is identified by gaze detection, and the data is sent back to the cloud server.
[0575] Step 6:
[0576] Input: Selection data sent to the cloud server
[0577] Processing: Answer generation
[0578] Output: The generated answer
[0579] The server analyzes the question selected by the user and generates an appropriate answer. This is done using database search and generative AI models (e.g., the transformers library). Specifically, Elasticsearch is used to retrieve the necessary information from the database, and an NLP model generates an answer based on that information. For example, in response to the question, "What are the reviews of this book?", review information is retrieved from a book review database and a concise review is generated.
[0580] Step 7:
[0581] Input: Generated answer
[0582] Action: Provide a response
[0583] Output: The answer provided to the user
[0584] The server sends the generated answer to the device. The device then displays the received answer to the user visually on a display or provides it audibly using speech synthesis technology. Specifically, the device uses the Google Text-to-Speech API to synthesize speech and provide the answer to the user audibly. The answer is also displayed visually on the smart glasses display. For example, the answer might be, "This book has excellent reviews, especially its storytelling."
[0585] (Application example 1)
[0586] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0587] In today's brick-and-mortar stores, users are required to quickly and efficiently obtain the information they need from a large selection of products. In particular, if users are unable to easily obtain detailed product information and reviews when choosing a product, their motivation to purchase is likely to decrease. Furthermore, if the information provided by the store is insufficient, it can affect users' decision-making, resulting in a negative impact on sales. To solve these problems, a system is needed that allows users to obtain product information in real time.
[0588] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0589] In this invention, the server includes a means for capturing gazes using a camera and a sensor built into the smart glasses in a physical store and analyzing the gazes on the cloud server to provide information about products, a means for displaying a list of questions received from the cloud server on the display of the smart glasses and selecting an appropriate question based on the user's gaze, and a means for the smart glasses to retrieve information about the selected question from a database and provide it to the user using voice synthesis technology or a display, thereby enabling the user to quickly and efficiently obtain detailed information about products.
[0590] A "sensor" is a device that detects physical changes and converts them into electrical signals.
[0591] A "camera" is a device for capturing images or video.
[0592] The "data collection means" is the part of the system that includes sensors and cameras to detect user activity.
[0593] The "data analysis means" is the part of the system that analyzes collected user behavior data and generates predicted questions.
[0594] The "display means" is a device that displays the generated questions to the user and allows the user to select one by line of sight.
[0595] The "answer generator" is the part of the system that generates and provides appropriate answers to user-selected questions.
[0596] A "physical store" is a physical store that a user visits to purchase products.
[0597] "Smart glasses" are eyeglass-type devices that display information and have built-in cameras and sensors to track the user's gaze.
[0598] A "cloud server" is a server for processing and storing data remotely via the Internet.
[0599] "Gaze capture" means tracking the user's gaze in real time and acquiring the data.
[0600] A "question list" is a collection of multiple questions generated based on user behavior data.
[0601] A "database" is a system that systematically manages data and allows for quick retrieval of necessary information.
[0602] A "display" is a device for visually displaying information.
[0603] "Speech synthesis technology" is a technology that converts text data into voice data.
[0604] This invention is a system that allows users to quickly and efficiently obtain detailed product information in physical stores. The system uses smart glasses to track the user's gaze and provides relevant information via a cloud server. To implement this invention, smart glasses, a cloud server, a database, and voice synthesis technology are used in combination.
[0605] Hardware and software used
[0606] Hardware:
[0607] Smart glasses: Equipped with a built-in camera and sensors, they capture the user's gaze in real time and have communication capabilities to send gaze data to a cloud server.
[0608] Display: Built into the smart glasses, it visually displays questions and answers to the user.
[0609] software:
[0610] Cloud server: Receives the user's gaze data, analyzes the data, generates necessary questions, and retrieves appropriate information from the database to display to the user.
[0611] Data analysis method: Runs on a cloud server and is responsible for analyzing gaze data and generating questions.
[0612] Question list generator: Generates appropriate questions based on the user's gaze information and sends this question list to the smart glasses.
[0613] Answer generation method: Based on the question selected by the user, information is retrieved from the database and an answer is generated to provide to the user. Natural language processing techniques (e.g., OpenAI GPT model) are used.
[0614] Speech synthesis technology: Technology for providing answers to users via voice (e.g., Google TTS).
[0615] What the system does
[0616] Overview of the process flow:
[0617] 1. Smart Glasses:
[0618] It uses built-in cameras and sensors to capture the user's gaze in real time.
[0619] The acquired gaze data is sent to a cloud server.
[0620] 2. Cloud Server:
[0621] The received gaze data is analyzed to identify information about the product that the user is paying attention to.
[0622] A list of relevant questions is generated and sent to the smart glasses.
[0623] 3. Smart Glasses:
[0624] The received question list is displayed on a display, and the user selects a question with their line of sight.
[0625] The question selected by the user is sent to the cloud server.
[0626] 4. Cloud Server:
[0627] Retrieve answers to selected questions from a database.
[0628] Based on the acquired information, an appropriate answer is generated and sent to the smart glasses.
[0629] 5. Smart Glasses:
[0630] The received answer is either visually displayed on a display or provided aloud using speech synthesis technology.
[0631] Specific examples
[0632] Suppose a user is looking at a particular product (e.g., organic food) in a physical store (e.g., a supermarket). The user's gaze data is captured by the smart glasses and sent to a cloud server. The cloud server analyzes the gaze data to identify the product the user is looking at and generates a list of questions related to this product. Questions such as "What are the nutritional content of this product?" and "What is the best way to cook it?" are displayed on the display of the user's smart glasses. When the user selects a question with their gaze, the selected question is sent to the cloud server. The cloud server retrieves an appropriate answer to the selected question from a database and sends it to the user's smart glasses. Finally, the smart glasses either display the retrieved answer visually or provide the answer audibly using speech synthesis technology.
[0633] Prompt Sentence Examples
[0634] Describe the process flow for a smart glasses application that tracks a user's gaze in real time and generates and answers questions related to specific supermarket items. Include specific details about the hardware, software, and data processing methods used.
[0635] In this way, the present invention allows users to quickly and efficiently obtain detailed information about products in a physical store.
[0636] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0637] Step 1:
[0638] The smart glasses' cameras and sensors capture the user's gaze.
[0639] Input: User gaze data
[0640] How it works: The smart glasses' built-in cameras and sensors track the user's gaze in real time and capture data to identify the object they are looking at.
[0641] Output: Captured gaze data
[0642] Step 2:
[0643] The smart glasses transmit the gaze data to a cloud server.
[0644] Input: Captured gaze data
[0645] How it works: The smart glasses transmit the acquired gaze data to a cloud server via the internet.
[0646] Output: Gaze data received by the cloud server
[0647] Step 3:
[0648] The cloud server analyzes the received gaze data and identifies information about the product that the user is paying attention to.
[0649] Input: Gaze data received by the cloud server
[0650] Operation: The data analysis means on the cloud server analyzes the gaze data and performs calculations to identify the object (product) that the gaze is directed at.
[0651] Output: Identified product information (product ID, etc.)
[0652] Step 4:
[0653] The cloud server generates a list of related questions based on the identified products.
[0654] Input: Identified product information
[0655] How it works: A data analysis tool on a cloud server uses a generative AI model to generate questions related to the identified products. These questions are formulated based on product characteristics.
[0656] Output: Generated question list
[0657] Step 5:
[0658] The smart glasses display the received question list on the display.
[0659] Input: Generated question list
[0660] How it works: A list of questions sent from a cloud server is displayed on the smart glasses, allowing the user to select a question with their gaze.
[0661] Output: Question list displayed on the screen
[0662] Step 6:
[0663] The user selects a question with their gaze.
[0664] Input: Question list displayed on the screen
[0665] How it works: The user uses their gaze to select a question of interest. The smart glasses sense the selected question.
[0666] Output: Selected Question
[0667] Step 7:
[0668] The smart glasses send the selected question to a cloud server.
[0669] Input: Selected Question
[0670] How it works: The smart glasses send the selected question over the internet to a cloud server.
[0671] Output: Questions received by the cloud server
[0672] Step 8:
[0673] The cloud server retrieves answers to the received questions from the database.
[0674] Input: A question received by the cloud server
[0675] Operation: The answer generation means on the cloud server searches the database for relevant information and processes the data to generate an answer.
[0676] Output: Retrieved answer information
[0677] Step 9:
[0678] The cloud server transmits the acquired answer information to the smart glasses.
[0679] Input: Retrieved answer information
[0680] How it works: The cloud server sends the generated answer information to the smart glasses via the internet.
[0681] Output: Answer information received by smart glasses
[0682] Step 10:
[0683] The smart glasses provide the received answer information to the user.
[0684] Input: Answer information received by smart glasses
[0685] How it works: The smart glasses will either visually display the answer information on a display or provide it aloud using speech synthesis technology.
[0686] Output: Answer information provided to the user
[0687] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0688] This invention provides a system that efficiently supports users in acquiring information by detecting a user's behavior, gaze, and even emotions, generating predicted questions, and providing appropriate answers to those questions. This system uses sensors and a camera to capture the user's behavior and gaze, recognizes the user's emotions using an emotion engine, analyzes the data using a cloud server, generates appropriate questions, displays them to the user, and generates and provides answers to the selected questions.
[0689] Program processing flow
[0690] This system operates based on the interaction between the "terminal" (smart glasses), the "server" on the cloud, and the "user." In addition, an emotion engine is built in to recognize the user's emotions.
[0691] Device: The device uses built-in sensors and cameras to capture the user's movements and gaze in real time. The sensors detect the user's movements (e.g., walking, standing, sitting, etc.), and the camera tracks the user's gaze direction and facial expressions. The captured data is temporarily stored on the device, and the emotion engine analyzes the user's facial expressions and voice to recognize emotions.
[0692] Server: The cloud server receives the data sent from the device and analyzes it in multiple ways. In particular, it uses data analysis tools to analyze the user's current behavior, gaze, and emotions to predict questions the user is likely to have in that situation. It then generates a list of predicted questions. The generated list of questions is sent from the cloud server to the device.
[0693] Terminal: The terminal displays the question list received from the server on a display, and the displayed questions are displayed so that they can be visually recognized within the user's field of vision.
[0694] User: The user selects a question displayed on the display by gazing at it. For example, if the user fixates their gaze on the question "How many calories does it have?" for a certain period of time, that question is recognized as selected.
[0695] Device: Sends the selected question to the server. The sent data includes the selected question, as well as the user's current situation and context information (including emotions).
[0696] Server: The server analyzes the received question and generates an appropriate answer. It searches for relevant information from a database and uses natural language processing techniques to create an answer in an easy-to-read format. It also adjusts the answer based on the emotions recognized by the emotion engine.
[0697] Terminal: Provides the received answer to the user. The answer is displayed visually on the display or provided aloud using speech synthesis technology. The answer is provided in an appropriate expression according to the user's emotion.
[0698] Specific examples
[0699] Example 1: Use at a restaurant
[0700] User: Looking at a new menu at a restaurant.
[0701] Device: Captures menus with the built-in camera and tracks the user's gaze and facial expressions.
[0702] Server: Recognizes when the user is looking at the menu and generates relevant question candidates (e.g., "What are the ingredients?", "How many calories?"). If the user's facial expression indicates interest, it prioritizes displaying relevant information.
[0703] Device: Candidate questions are displayed on the screen, and the user selects a question such as "How many calories?" with their gaze.
[0704] Server: Retrieves the calorie information of the menu item from the database and generates an answer. If the user looks surprised, the answer is provided in a softer tone.
[0705] Terminal: Provides answers visually or audibly.
[0706] Example 2: Use during learning
[0707] User: Studying by reading a textbook.
[0708] On the device: Built-in cameras and sensors identify where the user is reading, track their gaze and facial expressions, and an emotion engine recognizes the user's emotions (e.g., confusion, interest, etc.).
[0709] Server: Generates relevant question candidates (e.g., "What is the background of this incident?", "Who are the main characters?") based on the textbook content. If it determines that the user is confused, it prioritizes question candidates that include explanations.
[0710] Device: Candidate questions are displayed on the screen, and the user selects a question such as "What is the background to this incident?" with their gaze.
[0711] Server: Retrieves background information from textbook content and related materials and generates an answer. If the user is dissatisfied, the answer is expanded.
[0712] Terminal: Provides answers visually or audibly.
[0713] The system of this invention optimizes the provision of information by capturing the user's behavior, gaze, and emotions from multiple angles, and supports the user in resolving questions in their studies and daily life. By providing information according to the user's emotions, more effective learning and information acquisition are realized.
[0714] The processing flow will be explained below.
[0715] Step 1:
[0716] Device: The device uses built-in sensors and cameras to capture the user's movements, gaze, and facial expressions in real time. The sensors detect the user's movements (e.g., walking, standing, sitting, etc.), and the camera tracks the user's gaze direction and facial expressions. The captured data is temporarily stored on the device.
[0717] Step 2:
[0718] Device: Processes the collected behavior, gaze, and facial expression data, extracts necessary information, and sends it to the cloud server. This transmitted data includes the user's current behavioral patterns, gaze information, and facial expression data.
[0719] Step 3:
[0720] Server: The cloud server receives the transmitted data and analyzes it using data analysis methods and an emotion engine. In particular, it determines what information the user is paying attention to and what emotion (interest, confusion, delight, etc.) the user is feeling.
[0721] Step 4:
[0722] Server: Based on the analysis results, predicts the questions the user is likely to have in that situation. Generates a list of predicted questions that also reflects the user's emotional state. The generated list of questions is sent from the cloud server to the device.
[0723] Step 5:
[0724] Terminal: The received question list is displayed on the display. The displayed questions are displayed so that they can be visually recognized within the user's field of vision.
[0725] Step 6:
[0726] User: Selects a question displayed on the display by gazing at it. For example, if a user fixates their gaze on the question "How many calories does it have?" for a certain period of time, that question is recognized as selected.
[0727] Step 7:
[0728] The terminal sends the selected question to the server. The sent data includes the selected question, as well as the user's current situation and emotional information.
[0729] Step 8:
[0730] Server: The server analyzes the received question and generates an appropriate answer. It searches for relevant information from a database and creates an answer using natural language processing technology. It also adjusts the expression and level of detail of the answer based on the emotions recognized by the emotion engine.
[0731] Step 9:
[0732] Server: Sends the generated answer to the terminal, which may be presented as a visual display or audio.
[0733] Step 10:
[0734] Terminal: Provides the received answer to the user. The answer is displayed visually on the display or provided aloud using speech synthesis technology. The answer is provided in an appropriate expression according to the user's emotions.
[0735] Step 11:
[0736] Server: User responses and selection patterns are logged and used for later analysis, forming a feedback loop to improve the accuracy of future predictions.
[0737] These are the specific processing steps of this system. By comprehensively capturing the user's behavior, gaze, and emotions, and providing predicted questions and answers based on this information in real time, the system streamlines the user's learning process and information acquisition.
[0738] Example 2
[0739] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0740] Conventional information acquisition systems have had difficulty efficiently providing users with the information they seek. In particular, they have been unable to sense data such as the user's gaze and emotions in real time, pose appropriate questions, and provide answers tailored to the user. As a result, the user's information acquisition process has been subject to significant time and effort.
[0741] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes a data collection means including a sensor and a camera for detecting the user's movements, gaze data, and emotions, a data analysis means for analyzing the detected user's movement data, gaze data, and emotion data and generating predicted questions, a display means for displaying a question list generated based on the analysis means to the user and allowing the user to select a question with their gaze, an answer generation means for generating and providing an appropriate answer to the question selected by the user, and an emotion analysis means for adjusting the generated answer based on the user's emotion data. This makes it possible to efficiently obtain the information the user desires.
[0742] A "data collection means" is a device that includes sensors and cameras to detect the user's movements, gaze, and emotions.
[0743] The "data analysis means" is a device or software that analyzes the detected user's motion data, gaze data, and emotion data and generates predicted questions.
[0744] The "display means" is a device or software that displays the question list generated based on the analysis means to the user and allows the user to select a question by line of sight.
[0745] The "answer generation means" is a device or software that generates and provides an appropriate answer to a question selected by a user.
[0746] The "emotion analysis means" is a device or software that adjusts the response generated based on the user's emotional data.
[0747] This invention is a system that efficiently supports users in acquiring information by detecting a user's movements, gaze, and emotions, generating predicted questions, and providing appropriate answers to those questions. This system uses sensors and a camera to capture the user's movements and gaze, recognizes the user's emotions using an emotion engine, analyzes the data using a cloud server, generates appropriate questions and displays them to the user, and generates and provides answers to the selected questions.
[0748] The following components are primarily involved in the operation of the system:
[0749] Device: Uses a device (such as smart glasses) equipped with sensors and a camera to capture the user's movements, gaze, and emotions in real time. The sensor tracks the user's movements, and the camera tracks the direction of gaze and facial expressions, and this data is temporarily stored on the device. The device is also equipped with an emotion engine that analyzes the user's facial expressions and voice to recognize emotions.
[0750] Server: The cloud server receives the data sent from the device and uses data analysis methods to analyze the user's current behavior, gaze, and emotions. This analysis predicts the questions the user is likely to have in that situation and generates a list of questions. This list of questions is then sent from the cloud server to the device.
[0751] Display method: The device displays the question list received from the server on the display. The display format is an interactive format using pop-ups or eye tracking. For example, when a user is looking at a new menu, questions such as "How many calories are in this dish?" and "What are the ingredients?" are displayed.
[0752] User: The user selects a question by gazing at it. By fixing their gaze for a certain period of time, the system recognizes the user's selection and determines that the question has been selected.
[0753] Answer generation method: The selected question and context information (user's facial expression and current situation information) are sent from the device to the server. The server analyzes the question, searches for relevant information from a database, and generates an answer using natural language processing technology. It also adjusts the answer based on an emotion engine, for example, providing a softer tone if the user is surprised.
[0754] Providing an answer: The device provides the received answer to the user either visually on the display or audibly using speech synthesis technology. For example, it may visually display or audibly guide the user by saying, "This dish is 500 kcal."
[0755] Specific examples
[0756] Example 1: Use at a restaurant
[0757] 1. User: Looking at a new menu at a restaurant.
[0758] 2. Device: Capture menus with the built-in camera and track the user's gaze and facial expressions.
[0759] 3. Server: Recognizes that the user is looking at the menu and generates relevant question candidates ("What are the ingredients?", "How many calories?").
[0760] 4. Device: Question candidates are displayed on the screen, and the user selects a question such as "How many calories?" with their gaze.
[0761] 5. Server: Retrieves the calorie information of the menu item from the database and generates an answer. If the user looks surprised, the answer is provided in a softer tone.
[0762] 6. Terminal: Provides answers visually or audibly.
[0763] Example 2: Use during learning
[0764] 1. User: Reading and studying a textbook.
[0765] 2. Device: The built-in camera and sensors identify where the user is reading, track their gaze and facial expressions, and the emotion engine recognizes the user's emotions (e.g., confusion, interest, etc.).
[0766] 3. Server: Generates candidate questions ("What is the background of this incident?", "Who are the main characters?") based on the content of the textbook. If it determines that the user is confused, it prioritizes candidate questions that include explanations.
[0767] 4. Terminal: Candidate questions are displayed on the screen, and the user selects a question such as "What is the background to this incident?" with their gaze.
[0768] 5. Server: Retrieves background information from textbook content and related materials and generates an answer. If the user is dissatisfied, the answer is expanded.
[0769] 6. Terminal: Provides answers visually or audibly.
[0770] Prompt Sentence Examples
[0771] "A user is looking at a new menu item at a restaurant. If the user looks interested, suggest possible questions about the menu."
[0772] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0773] Step 1: Capture data with your device
[0774] Device: Activates built-in sensors and cameras to capture the user's movements, gaze, and facial expressions in real time. Inputs include the user's physical movements, gaze, and facial expressions, which are detected by the sensors and camera and temporarily stored as digital data. Specific movements include detecting walking and sitting, tracking gaze direction, and analyzing facial expressions. Outputs include movement data, gaze data, and facial expression data.
[0775] Step 2: Send data to the server
[0776] Terminal: Sends the captured data to the cloud server. The inputs are the movement data, gaze data, and facial expression data generated in step 1, and are sent to the server in real time. Specific operations include packaging and sending the data. The output is the data sent to the server.
[0777] Step 3: Data analysis and question generation by the server
[0778] Server: Analyzes the received data to determine the user's current behavior, gaze, and emotions. Inputs include motion data, gaze data, and facial expression data sent from the device, which are analyzed using machine learning algorithms and an emotion analysis engine. Specific operations include analyzing behavioral patterns, identifying gaze focus, and recognizing emotions. The output is a list of questions the user is likely to have in the situation.
[0779] Step 4: Sending the Question List from the Server to the Device
[0780] Server: Sends the question list generated by the analysis to the terminal. The input is the question list generated in step 3, which is sent to the terminal. Specifically, the server packages and sends the question list. The output is the question list sent to the terminal.
[0781] Step 5: Ask a question on the device
[0782] Terminal: The terminal displays the question list received from the server. The input is the question list sent from the server, which is displayed in a format that is easy for the user to view. Specific operations include displaying questions in a pop-up or interactive format. The output is the questions displayed on the screen.
[0783] Step 6: User selects question
[0784] User: Selects a question candidate displayed on the display with their gaze. The input is a list of questions displayed on the device, and the user fixates their gaze on a specific question for a certain period of time. Specific operations include processing the gaze tracking data and selecting a question. The output is the selected question.
[0785] Step 7: Sending the selected question to the server
[0786] Terminal: Sends the user's current situation and context information to the server along with the selected question. The inputs are the user's selected question and related context information (motion data, gaze data, and facial expression data), which are then sent to the server. Specific operations include packaging and sending the data. The output is the question and context information sent to the server.
[0787] Step 8: Server Generates Answer
[0788] Server: Analyzes the received question, retrieves relevant information from a database, and generates an answer. The input is the question and context information sent from the device, which is analyzed using natural language processing technology. Specific operations include database search, natural language generation, and response adjustment using an emotion engine. The output is the generated answer.
[0789] Step 9: Provide answers on your device
[0790] Terminal: Provides the received answer to the user. The input is the answer sent from the server, which is displayed on a screen or provided as voice using speech synthesis technology. Specific operations include visual display and speech synthesis. The output is the answer provided to the user.
[0791] (Application example 2)
[0792] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0793] Currently, when a customer has a question about a product in a physical store, it often takes time and effort to find a store employee to resolve the question. It can also be difficult for store staff to respond to all questions quickly and accurately. Furthermore, there is no system that can predict what customers are wondering about and provide information to answer those questions. This can lead to a poor customer experience and affect store sales.
[0794] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0795] In this invention, the server includes a data collection means including a sensor and a camera for detecting user behavior, a data analysis means for analyzing the detected user behavior data and emotion data and generating predicted questions, a display means for displaying the generated questions to the user and allowing the user to select one with their line of sight, and an answer generation means for generating and providing an appropriate answer to the question selected by the user. This makes it possible to provide quick and appropriate answers to questions that users have in physical stores.
[0796] "User" refers to an individual who uses the system.
[0797] An "action" is when a user performs a specific physical action or visual attention.
[0798] A "sensor" is a device that senses physical movements and environmental information and acquires data based on that information.
[0799] A "camera" is an optical device for capturing images or videos.
[0800] The "data collection means" is a means for collecting user behavioral data and emotional data using sensors and cameras.
[0801] "Emotion data" refers to emotional information sensed from the user's facial expressions, voice, etc.
[0802] "Data analysis means" refers to means for analyzing collected behavioral and emotional data and generating predicted questions.
[0803] The "display means" is a device that displays the generated questions to the user and allows the user to select one by line of sight.
[0804] The "answer generation means" is a means for generating and providing an appropriate answer to a question selected by a user.
[0805] "Natural language processing technology" is a technology that enables computers to understand and generate natural language.
[0806] A "generative AI model" is an artificial intelligence model that generates new information based on previously learned data.
[0807] The present invention provides a system that detects a user's behavior, gaze, and emotions, generates predicted questions based on the detected behavior, and provides appropriate answers to the questions. This system includes a data collection means, a data analysis means, a display means, and an answer generation means.
[0808] The system is implemented using the following hardware and software.
[0809] Hardware:
[0810] Smart glasses: Equipped with a built-in camera and display, they capture the user's gaze and facial expressions in real time.
[0811] Sensor: A device that detects user movements, such as an accelerometer or gyro sensor.
[0812] software:
[0813] OpenCV: A library for capturing and processing video data from cameras.
[0814] EmotionEngine: A system for analyzing emotions from a user's facial expressions and voice.
[0815] Cloud server: A server for data analysis and answer generation, utilizing natural language processing technology and generative AI models.
[0816] Natural language processing techniques: For example, using advanced language models such as GPT-3.
[0817] Process flow:
[0818] 1. Data collection methods:
[0819] The smart glasses' built-in cameras and sensors collect user behavioral and emotional data.
[0820] 2. Data analysis methods:
[0821] The collected data is analyzed on the cloud server to understand the user's current behavior and emotions and generate relevant questions.
[0822] 3. Display means:
[0823] The generated question list is displayed on the smart glasses display, and the user can select a question by eye gaze.
[0824] 4. Answer generation means:
[0825] Appropriate answers to questions selected by the user are generated on a cloud server and provided to the smart glasses either displayed or via voice.
[0826] Examples:
[0827] Example 1: Supermarket use:
[0828] User: Looking for food in the supermarket.
[0829] Smart glasses: Use a camera to track the user's gaze and analyze emotions.
[0830] Cloud server: Generates food-related question candidates (e.g., "What is the allergy information for this ingredient?").
[0831] Smart glasses: A list of questions is displayed and the user selects one by looking at the screen.
[0832] Cloud server: Generates answers to selected questions and displays them to the user.
[0833] Example prompt sentence:
[0834] "Please provide information about the ingredients of foods that may be of interest to users."
[0835] User gaze information: Fresh vegetable section
[0836] User sentiment: Showing interest
[0837] Questions generated: "What is the nutritional value of this vegetable?", "What dishes can it be used in?", "What is the allergy information?"
[0838] In this way, by using the system of the present invention, users can smoothly resolve their questions in physical stores, thereby improving the customer experience.
[0839] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0840] Step 1:
[0841] The device uses a built-in camera and sensors to capture the user's actions and gaze. The camera captures the user's gaze direction and facial expressions in real time, and the sensor detects the user's actions (e.g., reaching for a product). This captured data is collected by a data collection means. The input is the user's physical actions and gaze information, and the output is the captured action and gaze data.
[0842] Step 2:
[0843] The device sends the collected behavioral data and gaze data to a cloud server. The cloud server receives this data and analyzes it using data analysis means. The data analysis means analyzes the user's emotions from their gaze direction, movements, and facial expressions, and predicts questions that are likely to interest the user. For example, if the user is looking at a particular food item, it generates a question related to that food (e.g., "What are the ingredients in this food?"). The input is behavioral data and gaze data, and the output is a list of predicted questions.
[0844] Step 3:
[0845] The server generates questions based on the analysis results and sends them to the device in the form of a list. The device displays the received question list on the smart glasses display. The user visually checks the displayed question list and fixes their gaze on questions that interest them. The input is the predicted question list, and the output is the displayed question list.
[0846] Step 4:
[0847] When a user fixates their gaze on a question, the device detects the gaze information and determines that the user has selected the question. The selected question is then sent back to the cloud server, which then generates an appropriate answer for the question. The input is the question selected by the user, and the output is the generated answer.
[0848] Step 5:
[0849] The cloud server uses natural language processing technology and generative AI models (e.g., GPT-3) to generate answers to questions selected by the user. This answer is generated by searching for relevant information from a database and formatting it in an appropriate linguistic format. Furthermore, it takes into account the user's emotional information and generates an answer in a tone appropriate to the emotion. The input is the question selected by the user and emotional data, and the output is the generated answer.
[0850] Step 6:
[0851] The device receives the answer from the cloud server and provides it to the user either visually or audibly using speech synthesis technology. This allows the user to quickly obtain an answer to their question. The input is the generated answer, and the output is the answer provided to the user.
[0852] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0853] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0854] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0855] [Third embodiment]
[0856] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0857] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0858] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0859] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0860] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0861] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0862] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0863] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0864] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0865] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0866] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0867] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0868] This invention provides a system that efficiently supports users in acquiring information by detecting user behavior, generating predicted questions, and providing appropriate answers to those questions. This system uses sensors and cameras to capture the user's behavior and gaze, analyzes the data using a cloud server, generates appropriate questions, displays them to the user, and generates and provides answers to the selected questions.
[0869] Program processing flow
[0870] This system operates based on the interaction between the "terminal" (smart glasses), the "server" on the cloud, and the "user."
[0871] Device: Using built-in sensors and cameras, the device captures the user's movements and gaze in real time. The device temporarily stores the collected data and sends it to a cloud server as needed.
[0872] Server: The cloud server receives and analyzes the user data sent from the device. The data analysis method generates predicted questions based on the user's current situation. The generated question list is then sent from the server to the device.
[0873] Terminal: The terminal displays the question list received from the server on the display. When the user selects a question candidate on the display with their gaze, the selected question is sent to the server.
[0874] Server: The server analyzes the question sent by the user and generates an appropriate answer using database search and natural language processing techniques. The generated answer is then sent to the device.
[0875] Terminal: The terminal provides the received answer to the user either visually on a display or audibly using speech synthesis technology.
[0876] Specific examples
[0877] Example 1: Use at a restaurant
[0878] User: Looking at a new menu at a restaurant.
[0879] Device: Captures menus with the built-in camera and tracks the user's gaze.
[0880] Server: Recognizes when the user is looking at the menu and generates relevant question suggestions (e.g., "What are the ingredients?", "How many calories?").
[0881] Device: The question candidates are displayed on the screen and the user selects one. The user selects a question such as "How many calories?"
[0882] Server: Retrieves the calorie information for the menu items from the database and generates the answer.
[0883] Terminal: Provides answers visually or audibly.
[0884] Example 2: Use during learning
[0885] User: Studying by reading a textbook.
[0886] Device: Built-in camera and sensors identify where you are reading and track your gaze.
[0887] Server: Generates relevant question candidates (e.g., "What is the background to this incident?", "Who are the main characters?") based on the content of the textbook.
[0888] Terminal: The question candidates are displayed on the screen and the user can select one. The user selects the question "What is the background of this incident?"
[0889] Server: Obtains background information about the incident from textbook content and related materials and generates answers.
[0890] Terminal: Provides answers visually or audibly.
[0891] The system of the present invention is a powerful support tool for users to quickly obtain appropriate information, and plays an important role in today's information-overloaded society. It streamlines the user's learning process and provides practical solutions to various questions in daily life.
[0892] The processing flow will be explained below.
[0893] Step 1:
[0894] Device: The device uses built-in sensors and cameras to capture the user's movements and gaze in real time. The sensors detect the user's movements (e.g., walking, standing, sitting, etc.), and the camera tracks the user's gaze direction. The captured data is temporarily stored on the device.
[0895] Step 2:
[0896] Terminal: Processes the collected data, extracts necessary information, and sends it to the cloud server. This transmitted data includes the user's current behavioral patterns and gaze information.
[0897] Step 3:
[0898] Server: The cloud server receives the transmitted data and uses data analysis methods to analyze the user's behavior and gaze data. Specifically, it predicts what information the user is paying attention to and what questions they are likely to have.
[0899] Step 4:
[0900] Server: Based on the analysis results, a list of predicted questions is generated. The list of questions includes questions that the user is likely to have in that situation (e.g., "What are the ingredients in this dish?"). The generated list of questions is then sent back to the device.
[0901] Step 5:
[0902] Terminal: The received question list is displayed on the display. The displayed questions are displayed so that they can be visually recognized within the user's field of vision.
[0903] Step 6:
[0904] User: Selects a question displayed on the display by gazing at it. For example, if a user fixates their gaze on the question "How many calories does it have?" for a certain period of time, that question is recognized as selected.
[0905] Step 7:
[0906] The terminal sends the selected question to the server. The sent data includes the selected question as well as the user's current situation and context information.
[0907] Step 8:
[0908] Server: The server analyzes the received question and generates an appropriate answer. It searches for relevant information from a database and uses natural language processing techniques to create an answer in an easy-to-read format.
[0909] Step 9:
[0910] Server: Sends the generated answer to the terminal, which may be presented visually or audibly.
[0911] Step 10:
[0912] Terminal: Provides the received answer to the user, either visually displaying the answer on a display or audibly using speech synthesis technology.
[0913] These are the specific processing steps of the "Chatty Glasses" system. By capturing the user's behavior and gaze in real time, and using that information to predict and generate appropriate questions and provide answers, the system helps users obtain information more efficiently.
[0914] Example 1
[0915] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0916] In modern society, it is important for users to quickly obtain appropriate information for efficient decision-making and learning. However, it is not easy to find the information they need from a vast amount of information. There is also a lack of mechanisms that can accurately identify when a user has a question and quickly provide an appropriate answer to that question. The present invention aims to solve these problems by providing a system that improves the efficiency of information acquisition and supports users in resolving their questions.
[0917] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0918] In this invention, the server includes a data collection means including a sensor and a camera for detecting user behavior, a communication means for transmitting the detected user behavior data to the cloud server, a data analysis means for analyzing the transmitted data and generating predicted questions, a display means for displaying the generated question list to the user and allowing the user to select one with their line of sight, and an answer generation means for generating and providing an appropriate answer to the question selected by the user, thereby enabling the user to efficiently obtain information and quickly resolve their question.
[0919] "Data collection means" refers to devices and systems that include sensors and cameras for detecting user behavior.
[0920] "Communication means" refers to the technology or mechanism for transmitting detected user behavior data to a cloud server.
[0921] "Data analysis means" refers to software or algorithms that analyze data sent to the cloud server and generate predicted questions.
[0922] The "display means" is a device or interface that displays the generated questions to the user and allows the user to select them with their line of sight.
[0923] The "answer generation means" is software or a device that generates and provides an appropriate answer to a question selected by a user.
[0924] A "sensor" is a device that detects a user's actions and gaze, and can sense changes in movement and gaze.
[0925] A "camera" is a photographing device for visually capturing a user's line of sight and actions.
[0926] A "cloud server" is a server that can be accessed remotely for data analysis and management.
[0927] "Eye tracking" is a technology that tracks the movement of a user's eyes in real time.
[0928] "Natural language processing technology" is a technology for analyzing text data and understanding and generating human language.
[0929] "Database searching" is the technique or method of searching through a database to retrieve specific information.
[0930] This invention provides a system that efficiently supports users in acquiring information by detecting user behavior, generating predicted questions, and providing appropriate answers to those questions. This system uses sensors and cameras to capture the user's behavior and gaze, analyzes the data using a cloud server, generates appropriate questions, displays them to the user, and generates and provides answers to the selected questions.
[0931] 1. Hardware
[0932] Terminal
[0933] Use smart glasses, which have built-in sensors and cameras to detect the user's gaze and behavior, such as devices like Tobii Pro Glasses 3.
[0934] server
[0935] Use a server on the cloud. The server runs on a cloud service such as an AWS EC2 instance.
[0936] 2. Software
[0937] Data collection
[0938] An application is installed on the device to collect data in real time and send it to a cloud server. This application is developed using Unity, OpenCV, etc.
[0939] Data analysis
[0940] The cloud server uses programs such as Python, TensorFlow, and PyTorch to analyze the received data, which then generates predicted questions based on the user's behavioral data.
[0941] Display
[0942] The device displays the generated question list on its screen, and AR technology is used as the interface for users to select questions with their gaze.
[0943] Answer generation
[0944] To generate answers to selected questions, generative AI models and database search techniques are used. Specific examples include the transformers library and Elasticsearch. The generated answers are sent to the device and presented visually or using speech synthesis techniques. The speech synthesis uses the Google Text-to-Speech API.
[0945] 3. Specific Examples
[0946] Example 1: Use at a restaurant
[0947] User: Looking at a new menu at a restaurant.
[0948] Device: Captures menus with the built-in camera and tracks the user's gaze.
[0949] Server: Recognizes when the user is looking at the menu and generates relevant question candidates (e.g., "What are the ingredients?", "How many calories?").
[0950] Device: Candidate questions are displayed on the screen, and the user selects the question "How many calories?" with their gaze.
[0951] Server: Retrieves the calorie information for the menu items from the database and generates the answer.
[0952] Terminal: Provides answers visually or audibly.
[0953] Example prompt sentence:
[0954] "If a user is looking at a new menu item at a restaurant, how does the system pose questions to the user and provide answers?"
[0955] Example 2: Use during learning
[0956] User: Studying by reading a textbook.
[0957] Device: Built-in camera and sensors identify where you are reading and track your gaze.
[0958] Server: Generates relevant question candidates (e.g., "What is the background of this incident?", "Who are the main characters?") based on the content of the textbook.
[0959] Device: Candidate questions are displayed on the screen, and the user selects a question such as "What is the background to this incident?" with their gaze.
[0960] Server: Obtains background information about the incident from textbook content and related materials and generates answers.
[0961] Terminal: Provides answers visually or audibly.
[0962] Example prompt sentence:
[0963] "When a user is studying a textbook, how does the system pose questions to the user and provide answers?"
[0964] The system of the present invention is a powerful support tool for users to quickly obtain appropriate information, and plays an important role in today's information-overloaded society. It streamlines the user's learning process and provides practical solutions to various questions in daily life.
[0965] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0966] Step 1:
[0967] Input: User behavior and gaze data
[0968] Processing: Data Collection
[0969] Output: User behavior and gaze data
[0970] The device uses sensors and cameras built into the smart glasses to capture the user's actions and gaze in real time. Specifically, the camera captures the user's gaze direction, and the sensor detects movement. For example, when a user picks up a book in a bookstore, their movement and gaze are captured. This allows detailed behavioral and gaze data to be collected.
[0971] Step 2:
[0972] Input: User behavior and gaze data
[0973] Action: Send data
[0974] Output: Data sent to the cloud server
[0975] The device temporarily stores the collected behavioral and gaze data in its internal memory and then transmits the data to a cloud server via Wi-Fi or Bluetooth, using the TLS encryption protocol to protect the data transmission. Specifically, the device generates collected data packets and sequentially transmits them to a designated endpoint on the cloud server.
[0976] Step 3:
[0977] Input: Data sent to the cloud server
[0978] Processing: Data analysis
[0979] Output: Generated question list
[0980] The server receives the data sent to the cloud server and analyzes it using programs such as Python and TensorFlow. Specifically, it analyzes the user's behavioral and gaze data to infer the user's intentions and interests. Based on this, it generates predicted questions. This analysis process uses natural language processing (NLP) technology and machine learning models. For example, if a user is interested in a particular book in a bookstore, it generates questions such as, "What are the reviews for this book?" and "Who is the author?"
[0981] Step 4:
[0982] Input: Generated question list
[0983] Action: Display Question
[0984] Output: The question displayed to the user
[0985] The server sends the generated question list to the device. The device displays the received question list on its display. The user selects a question using eye movements or a pointer on the interface. Specifically, the system uses AR technology to visually display the question list, allowing the user to visually select a question.
[0986] Step 5:
[0987] Input: User selected question
[0988] Process: Send selected data
[0989] Output: Selection data sent to the cloud server
[0990] The user selects a question displayed on the display by looking at it. The device detects this selection and sends the information to the cloud server. The selected question is identified by gaze detection, and the data is sent back to the cloud server.
[0991] Step 6:
[0992] Input: Selection data sent to the cloud server
[0993] Processing: Answer generation
[0994] Output: The generated answer
[0995] The server analyzes the question selected by the user and generates an appropriate answer. This is done using database search and generative AI models (e.g., the transformers library). Specifically, Elasticsearch is used to retrieve the necessary information from the database, and an NLP model generates an answer based on that information. For example, in response to the question, "What are the reviews of this book?", review information is retrieved from a book review database and a concise review is generated.
[0996] Step 7:
[0997] Input: Generated answer
[0998] Action: Provide a response
[0999] Output: The answer provided to the user
[1000] The server sends the generated answer to the device. The device then displays the received answer to the user visually on a display or provides it audibly using speech synthesis technology. Specifically, the device uses the Google Text-to-Speech API to synthesize speech and provide the answer to the user audibly. The answer is also displayed visually on the smart glasses display. For example, the answer might be, "This book has excellent reviews, especially its storytelling."
[1001] (Application example 1)
[1002] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1003] In today's brick-and-mortar stores, users are required to quickly and efficiently obtain the information they need from a large selection of products. In particular, if users are unable to easily obtain detailed product information and reviews when choosing a product, their motivation to purchase is likely to decrease. Furthermore, if the information provided by the store is insufficient, it can affect users' decision-making, resulting in a negative impact on sales. To solve these problems, a system is needed that allows users to obtain product information in real time.
[1004] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1005] In this invention, the server includes a means for capturing gazes using a camera and a sensor built into the smart glasses in a physical store and analyzing the gazes on the cloud server to provide information about products, a means for displaying a list of questions received from the cloud server on the display of the smart glasses and selecting an appropriate question based on the user's gaze, and a means for the smart glasses to retrieve information about the selected question from a database and provide it to the user using voice synthesis technology or a display, thereby enabling the user to quickly and efficiently obtain detailed information about products.
[1006] A "sensor" is a device that detects physical changes and converts them into electrical signals.
[1007] A "camera" is a device for capturing images or video.
[1008] The "data collection means" is the part of the system that includes sensors and cameras to detect user activity.
[1009] The "data analysis means" is the part of the system that analyzes collected user behavior data and generates predicted questions.
[1010] The "display means" is a device that displays the generated questions to the user and allows the user to select one by line of sight.
[1011] The "answer generator" is the part of the system that generates and provides appropriate answers to user-selected questions.
[1012] A "physical store" is a physical store that a user visits to purchase products.
[1013] "Smart glasses" are eyeglass-type devices that display information and have built-in cameras and sensors to track the user's gaze.
[1014] A "cloud server" is a server for processing and storing data remotely via the Internet.
[1015] "Gaze capture" means tracking the user's gaze in real time and acquiring the data.
[1016] A "question list" is a collection of multiple questions generated based on user behavior data.
[1017] A "database" is a system that systematically manages data and allows for quick retrieval of necessary information.
[1018] A "display" is a device for visually displaying information.
[1019] "Speech synthesis technology" is a technology that converts text data into voice data.
[1020] This invention is a system that allows users to quickly and efficiently obtain detailed product information in physical stores. The system uses smart glasses to track the user's gaze and provides relevant information via a cloud server. To implement this invention, smart glasses, a cloud server, a database, and voice synthesis technology are used in combination.
[1021] Hardware and software used
[1022] Hardware:
[1023] Smart glasses: Equipped with a built-in camera and sensors, they capture the user's gaze in real time and have communication capabilities to send gaze data to a cloud server.
[1024] Display: Built into the smart glasses, it visually displays questions and answers to the user.
[1025] software:
[1026] Cloud server: Receives the user's gaze data, analyzes the data, generates necessary questions, and retrieves appropriate information from the database to display to the user.
[1027] Data analysis method: Runs on a cloud server and is responsible for analyzing gaze data and generating questions.
[1028] Question list generator: Generates appropriate questions based on the user's gaze information and sends this question list to the smart glasses.
[1029] Answer generation method: Based on the question selected by the user, information is retrieved from the database and an answer is generated to provide to the user. Natural language processing techniques (e.g., OpenAI GPT model) are used.
[1030] Speech synthesis technology: Technology for providing answers to users via voice (e.g., Google TTS).
[1031] What the system does
[1032] Overview of the process flow:
[1033] 1. Smart Glasses:
[1034] It uses built-in cameras and sensors to capture the user's gaze in real time.
[1035] The acquired gaze data is sent to a cloud server.
[1036] 2. Cloud Server:
[1037] The received gaze data is analyzed to identify information about the product that the user is paying attention to.
[1038] A list of relevant questions is generated and sent to the smart glasses.
[1039] 3. Smart Glasses:
[1040] The received question list is displayed on a display, and the user selects a question with their line of sight.
[1041] The question selected by the user is sent to the cloud server.
[1042] 4. Cloud Server:
[1043] Retrieve answers to selected questions from a database.
[1044] Based on the acquired information, an appropriate answer is generated and sent to the smart glasses.
[1045] 5. Smart Glasses:
[1046] The received answer is either visually displayed on a display or provided aloud using speech synthesis technology.
[1047] Specific examples
[1048] Suppose a user is looking at a particular product (e.g., organic food) in a physical store (e.g., a supermarket). The user's gaze data is captured by the smart glasses and sent to a cloud server. The cloud server analyzes the gaze data to identify the product the user is looking at and generates a list of questions related to this product. Questions such as "What are the nutritional content of this product?" and "What is the best way to cook it?" are displayed on the display of the user's smart glasses. When the user selects a question with their gaze, the selected question is sent to the cloud server. The cloud server retrieves an appropriate answer to the selected question from a database and sends it to the user's smart glasses. Finally, the smart glasses either display the retrieved answer visually or provide the answer audibly using speech synthesis technology.
[1049] Prompt Sentence Examples
[1050] Describe the process flow for a smart glasses application that tracks a user's gaze in real time and generates and answers questions related to specific supermarket items. Include specific details about the hardware, software, and data processing methods used.
[1051] In this way, the present invention allows users to quickly and efficiently obtain detailed information about products in a physical store.
[1052] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1053] Step 1:
[1054] The smart glasses' cameras and sensors capture the user's gaze.
[1055] Input: User gaze data
[1056] How it works: The smart glasses' built-in cameras and sensors track the user's gaze in real time and capture data to identify the object they are looking at.
[1057] Output: Captured gaze data
[1058] Step 2:
[1059] The smart glasses transmit the gaze data to a cloud server.
[1060] Input: Captured gaze data
[1061] How it works: The smart glasses transmit the acquired gaze data to a cloud server via the internet.
[1062] Output: Gaze data received by the cloud server
[1063] Step 3:
[1064] The cloud server analyzes the received gaze data and identifies information about the product that the user is paying attention to.
[1065] Input: Gaze data received by the cloud server
[1066] Operation: The data analysis means on the cloud server analyzes the gaze data and performs calculations to identify the object (product) that the gaze is directed at.
[1067] Output: Identified product information (product ID, etc.)
[1068] Step 4:
[1069] The cloud server generates a list of related questions based on the identified products.
[1070] Input: Identified product information
[1071] How it works: A data analysis tool on a cloud server uses a generative AI model to generate questions related to the identified products. These questions are formulated based on product characteristics.
[1072] Output: Generated question list
[1073] Step 5:
[1074] The smart glasses display the received question list on the display.
[1075] Input: Generated question list
[1076] How it works: A list of questions sent from a cloud server is displayed on the smart glasses, allowing the user to select a question with their gaze.
[1077] Output: Question list displayed on the screen
[1078] Step 6:
[1079] The user selects a question with their gaze.
[1080] Input: Question list displayed on the screen
[1081] How it works: The user uses their gaze to select a question of interest. The smart glasses sense the selected question.
[1082] Output: Selected Question
[1083] Step 7:
[1084] The smart glasses send the selected question to a cloud server.
[1085] Input: Selected Question
[1086] How it works: The smart glasses send the selected question over the internet to a cloud server.
[1087] Output: Questions received by the cloud server
[1088] Step 8:
[1089] The cloud server retrieves answers to the received questions from the database.
[1090] Input: A question received by the cloud server
[1091] Operation: The answer generation means on the cloud server searches the database for relevant information and processes the data to generate an answer.
[1092] Output: Retrieved answer information
[1093] Step 9:
[1094] The cloud server transmits the acquired answer information to the smart glasses.
[1095] Input: Retrieved answer information
[1096] How it works: The cloud server sends the generated answer information to the smart glasses via the internet.
[1097] Output: Answer information received by smart glasses
[1098] Step 10:
[1099] The smart glasses provide the received answer information to the user.
[1100] Input: Answer information received by smart glasses
[1101] How it works: The smart glasses will either visually display the answer information on a display or provide it aloud using speech synthesis technology.
[1102] Output: Answer information provided to the user
[1103] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1104] This invention provides a system that efficiently supports users in acquiring information by detecting a user's behavior, gaze, and even emotions, generating predicted questions, and providing appropriate answers to those questions. This system uses sensors and a camera to capture the user's behavior and gaze, recognizes the user's emotions using an emotion engine, analyzes the data using a cloud server, generates appropriate questions, displays them to the user, and generates and provides answers to the selected questions.
[1105] Program processing flow
[1106] This system operates based on the interaction between the "terminal" (smart glasses), the "server" on the cloud, and the "user." In addition, an emotion engine is built in to recognize the user's emotions.
[1107] Device: The device uses built-in sensors and cameras to capture the user's movements and gaze in real time. The sensors detect the user's movements (e.g., walking, standing, sitting, etc.), and the camera tracks the user's gaze direction and facial expressions. The captured data is temporarily stored on the device, and the emotion engine analyzes the user's facial expressions and voice to recognize emotions.
[1108] Server: The cloud server receives the data sent from the device and analyzes it in multiple ways. In particular, it uses data analysis tools to analyze the user's current behavior, gaze, and emotions to predict questions the user is likely to have in that situation. It then generates a list of predicted questions. The generated list of questions is sent from the cloud server to the device.
[1109] Terminal: The terminal displays the question list received from the server on a display, and the displayed questions are displayed so that they can be visually recognized within the user's field of vision.
[1110] User: The user selects a question displayed on the display by gazing at it. For example, if the user fixates their gaze on the question "How many calories does it have?" for a certain period of time, that question is recognized as selected.
[1111] Device: Sends the selected question to the server. The sent data includes the selected question, as well as the user's current situation and context information (including emotions).
[1112] Server: The server analyzes the received question and generates an appropriate answer. It searches for relevant information from a database and uses natural language processing techniques to create an answer in an easy-to-read format. It also adjusts the answer based on the emotions recognized by the emotion engine.
[1113] Terminal: Provides the received answer to the user. The answer is displayed visually on the display or provided aloud using speech synthesis technology. The answer is provided in an appropriate expression according to the user's emotion.
[1114] Specific examples
[1115] Example 1: Use at a restaurant
[1116] User: Looking at a new menu at a restaurant.
[1117] Device: Captures menus with the built-in camera and tracks the user's gaze and facial expressions.
[1118] Server: Recognizes when the user is looking at the menu and generates relevant question candidates (e.g., "What are the ingredients?", "How many calories?"). If the user's facial expression indicates interest, it prioritizes displaying relevant information.
[1119] Device: Candidate questions are displayed on the screen, and the user selects a question such as "How many calories?" with their gaze.
[1120] Server: Retrieves the calorie information of the menu item from the database and generates an answer. If the user looks surprised, the answer is provided in a softer tone.
[1121] Terminal: Provides answers visually or audibly.
[1122] Example 2: Use during learning
[1123] User: Studying by reading a textbook.
[1124] On the device: Built-in cameras and sensors identify where the user is reading, track their gaze and facial expressions, and an emotion engine recognizes the user's emotions (e.g., confusion, interest, etc.).
[1125] Server: Generates relevant question candidates (e.g., "What is the background of this incident?", "Who are the main characters?") based on the textbook content. If it determines that the user is confused, it prioritizes question candidates that include explanations.
[1126] Device: Candidate questions are displayed on the screen, and the user selects a question such as "What is the background to this incident?" with their gaze.
[1127] Server: Retrieves background information from textbook content and related materials and generates an answer. If the user is dissatisfied, the answer is expanded.
[1128] Terminal: Provides answers visually or audibly.
[1129] The system of this invention optimizes the provision of information by capturing the user's behavior, gaze, and emotions from multiple angles, and supports the user in resolving questions in their studies and daily life. By providing information according to the user's emotions, more effective learning and information acquisition are realized.
[1130] The processing flow will be explained below.
[1131] Step 1:
[1132] Device: The device uses built-in sensors and cameras to capture the user's movements, gaze, and facial expressions in real time. The sensors detect the user's movements (e.g., walking, standing, sitting, etc.), and the camera tracks the user's gaze direction and facial expressions. The captured data is temporarily stored on the device.
[1133] Step 2:
[1134] Device: Processes the collected behavior, gaze, and facial expression data, extracts necessary information, and sends it to the cloud server. This transmitted data includes the user's current behavioral patterns, gaze information, and facial expression data.
[1135] Step 3:
[1136] Server: The cloud server receives the transmitted data and analyzes it using data analysis methods and an emotion engine. In particular, it determines what information the user is paying attention to and what emotion (interest, confusion, delight, etc.) the user is feeling.
[1137] Step 4:
[1138] Server: Based on the analysis results, predicts the questions the user is likely to have in that situation. Generates a list of predicted questions that also reflects the user's emotional state. The generated list of questions is sent from the cloud server to the device.
[1139] Step 5:
[1140] Terminal: The received question list is displayed on the display. The displayed questions are displayed so that they can be visually recognized within the user's field of vision.
[1141] Step 6:
[1142] User: Selects a question displayed on the display by gazing at it. For example, if a user fixates their gaze on the question "How many calories does it have?" for a certain period of time, that question is recognized as selected.
[1143] Step 7:
[1144] The terminal sends the selected question to the server. The sent data includes the selected question, as well as the user's current situation and emotional information.
[1145] Step 8:
[1146] Server: The server analyzes the received question and generates an appropriate answer. It searches for relevant information from a database and creates an answer using natural language processing technology. It also adjusts the expression and level of detail of the answer based on the emotions recognized by the emotion engine.
[1147] Step 9:
[1148] Server: Sends the generated answer to the terminal, which may be presented as a visual display or audio.
[1149] Step 10:
[1150] Terminal: Provides the received answer to the user. The answer is displayed visually on the display or provided aloud using speech synthesis technology. The answer is provided in an appropriate expression according to the user's emotions.
[1151] Step 11:
[1152] Server: User responses and selection patterns are logged and used for later analysis, forming a feedback loop to improve the accuracy of future predictions.
[1153] These are the specific processing steps of this system. By comprehensively capturing the user's behavior, gaze, and emotions, and providing predicted questions and answers based on this information in real time, the system streamlines the user's learning process and information acquisition.
[1154] Example 2
[1155] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1156] Conventional information acquisition systems have had difficulty efficiently providing users with the information they seek. In particular, they have been unable to sense data such as the user's gaze and emotions in real time, pose appropriate questions, and provide answers tailored to the user. As a result, the user's information acquisition process has been subject to significant time and effort.
[1157] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes a data collection means including a sensor and a camera for detecting the user's movements, gaze data, and emotions, a data analysis means for analyzing the detected user's movement data, gaze data, and emotion data and generating predicted questions, a display means for displaying a question list generated based on the analysis means to the user and allowing the user to select a question with their gaze, an answer generation means for generating and providing an appropriate answer to the question selected by the user, and an emotion analysis means for adjusting the generated answer based on the user's emotion data. This makes it possible to efficiently obtain the information the user desires.
[1158] A "data collection means" is a device that includes sensors and cameras to detect the user's movements, gaze, and emotions.
[1159] The "data analysis means" is a device or software that analyzes the detected user's motion data, gaze data, and emotion data and generates predicted questions.
[1160] The "display means" is a device or software that displays the question list generated based on the analysis means to the user and allows the user to select a question by line of sight.
[1161] The "answer generation means" is a device or software that generates and provides an appropriate answer to a question selected by a user.
[1162] The "emotion analysis means" is a device or software that adjusts the response generated based on the user's emotional data.
[1163] This invention is a system that efficiently supports users in acquiring information by detecting a user's movements, gaze, and emotions, generating predicted questions, and providing appropriate answers to those questions. This system uses sensors and a camera to capture the user's movements and gaze, recognizes the user's emotions using an emotion engine, analyzes the data using a cloud server, generates appropriate questions and displays them to the user, and generates and provides answers to the selected questions.
[1164] The following components are primarily involved in the operation of the system:
[1165] Device: Uses a device (such as smart glasses) equipped with sensors and a camera to capture the user's movements, gaze, and emotions in real time. The sensor tracks the user's movements, and the camera tracks the direction of gaze and facial expressions, and this data is temporarily stored on the device. The device is also equipped with an emotion engine that analyzes the user's facial expressions and voice to recognize emotions.
[1166] Server: The cloud server receives the data sent from the device and uses data analysis methods to analyze the user's current behavior, gaze, and emotions. This analysis predicts the questions the user is likely to have in that situation and generates a list of questions. This list of questions is then sent from the cloud server to the device.
[1167] Display method: The device displays the question list received from the server on the display. The display format is an interactive format using pop-ups or eye tracking. For example, when a user is looking at a new menu, questions such as "How many calories are in this dish?" and "What are the ingredients?" are displayed.
[1168] User: The user selects a question by gazing at it. By fixing their gaze for a certain period of time, the system recognizes the user's selection and determines that the question has been selected.
[1169] Answer generation method: The selected question and context information (user's facial expression and current situation information) are sent from the device to the server. The server analyzes the question, searches for relevant information from a database, and generates an answer using natural language processing technology. It also adjusts the answer based on an emotion engine, for example, providing a softer tone if the user is surprised.
[1170] Providing an answer: The device provides the received answer to the user either visually on the display or audibly using speech synthesis technology. For example, it may visually display or audibly guide the user by saying, "This dish is 500 kcal."
[1171] Specific examples
[1172] Example 1: Use at a restaurant
[1173] 1. User: Looking at a new menu at a restaurant.
[1174] 2. Device: Capture menus with the built-in camera and track the user's gaze and facial expressions.
[1175] 3. Server: Recognizes that the user is looking at the menu and generates relevant question candidates ("What are the ingredients?", "How many calories?").
[1176] 4. Device: Question candidates are displayed on the screen, and the user selects a question such as "How many calories?" with their gaze.
[1177] 5. Server: Retrieves the calorie information of the menu item from the database and generates an answer. If the user looks surprised, the answer is provided in a softer tone.
[1178] 6. Terminal: Provides answers visually or audibly.
[1179] Example 2: Use during learning
[1180] 1. User: Reading and studying a textbook.
[1181] 2. Device: The built-in camera and sensors identify where the user is reading, track their gaze and facial expressions, and the emotion engine recognizes the user's emotions (e.g., confusion, interest, etc.).
[1182] 3. Server: Generates candidate questions ("What is the background of this incident?", "Who are the main characters?") based on the content of the textbook. If it determines that the user is confused, it prioritizes candidate questions that include explanations.
[1183] 4. Terminal: Candidate questions are displayed on the screen, and the user selects a question such as "What is the background to this incident?" with their gaze.
[1184] 5. Server: Retrieves background information from textbook content and related materials and generates an answer. If the user is dissatisfied, the answer is expanded.
[1185] 6. Terminal: Provides answers visually or audibly.
[1186] Prompt Sentence Examples
[1187] "A user is looking at a new menu item at a restaurant. If the user looks interested, suggest possible questions about the menu."
[1188] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1189] Step 1: Capture data with your device
[1190] Device: Activates built-in sensors and cameras to capture the user's movements, gaze, and facial expressions in real time. Inputs include the user's physical movements, gaze, and facial expressions, which are detected by the sensors and camera and temporarily stored as digital data. Specific movements include detecting walking and sitting, tracking gaze direction, and analyzing facial expressions. Outputs include movement data, gaze data, and facial expression data.
[1191] Step 2: Send data to the server
[1192] Terminal: Sends the captured data to the cloud server. The inputs are the movement data, gaze data, and facial expression data generated in step 1, and are sent to the server in real time. Specific operations include packaging and sending the data. The output is the data sent to the server.
[1193] Step 3: Data analysis and question generation by the server
[1194] Server: Analyzes the received data to determine the user's current behavior, gaze, and emotions. Inputs include motion data, gaze data, and facial expression data sent from the device, which are analyzed using machine learning algorithms and an emotion analysis engine. Specific operations include analyzing behavioral patterns, identifying gaze focus, and recognizing emotions. The output is a list of questions the user is likely to have in the situation.
[1195] Step 4: Sending the Question List from the Server to the Device
[1196] Server: Sends the question list generated by the analysis to the terminal. The input is the question list generated in step 3, which is sent to the terminal. Specifically, the server packages and sends the question list. The output is the question list sent to the terminal.
[1197] Step 5: Ask a question on the device
[1198] Terminal: The terminal displays the question list received from the server. The input is the question list sent from the server, which is displayed in a format that is easy for the user to view. Specific operations include displaying questions in a pop-up or interactive format. The output is the questions displayed on the screen.
[1199] Step 6: User selects question
[1200] User: Selects a question candidate displayed on the display with their gaze. The input is a list of questions displayed on the device, and the user fixates their gaze on a specific question for a certain period of time. Specific operations include processing the gaze tracking data and selecting a question. The output is the selected question.
[1201] Step 7: Sending the selected question to the server
[1202] Terminal: Sends the user's current situation and context information to the server along with the selected question. The inputs are the user's selected question and related context information (motion data, gaze data, and facial expression data), which are then sent to the server. Specific operations include packaging and sending the data. The output is the question and context information sent to the server.
[1203] Step 8: Server Generates Answer
[1204] Server: Analyzes the received question, retrieves relevant information from a database, and generates an answer. The input is the question and context information sent from the device, which is analyzed using natural language processing technology. Specific operations include database search, natural language generation, and response adjustment using an emotion engine. The output is the generated answer.
[1205] Step 9: Provide answers on your device
[1206] Terminal: Provides the received answer to the user. The input is the answer sent from the server, which is displayed on a screen or provided as voice using speech synthesis technology. Specific operations include visual display and speech synthesis. The output is the answer provided to the user.
[1207] (Application example 2)
[1208] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1209] Currently, when a customer has a question about a product in a physical store, it often takes time and effort to find a store employee to resolve the question. It can also be difficult for store staff to respond to all questions quickly and accurately. Furthermore, there is no system that can predict what customers are wondering about and provide information to answer those questions. This can lead to a poor customer experience and affect store sales.
[1210] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1211] In this invention, the server includes a data collection means including a sensor and a camera for detecting user behavior, a data analysis means for analyzing the detected user behavior data and emotion data and generating predicted questions, a display means for displaying the generated questions to the user and allowing the user to select one with their line of sight, and an answer generation means for generating and providing an appropriate answer to the question selected by the user. This makes it possible to provide quick and appropriate answers to questions that users have in physical stores.
[1212] "User" refers to an individual who uses the system.
[1213] An "action" is when a user performs a specific physical action or visual attention.
[1214] A "sensor" is a device that senses physical movements and environmental information and acquires data based on that information.
[1215] A "camera" is an optical device for capturing images or videos.
[1216] The "data collection means" is a means for collecting user behavioral data and emotional data using sensors and cameras.
[1217] "Emotion data" refers to emotional information sensed from the user's facial expressions, voice, etc.
[1218] "Data analysis means" refers to means for analyzing collected behavioral and emotional data and generating predicted questions.
[1219] The "display means" is a device that displays the generated questions to the user and allows the user to select one by line of sight.
[1220] The "answer generation means" is a means for generating and providing an appropriate answer to a question selected by a user.
[1221] "Natural language processing technology" is a technology that enables computers to understand and generate natural language.
[1222] A "generative AI model" is an artificial intelligence model that generates new information based on previously learned data.
[1223] The present invention provides a system that detects a user's behavior, gaze, and emotions, generates predicted questions based on the detected behavior, and provides appropriate answers to the questions. This system includes a data collection means, a data analysis means, a display means, and an answer generation means.
[1224] The system is implemented using the following hardware and software.
[1225] Hardware:
[1226] Smart glasses: Equipped with a built-in camera and display, they capture the user's gaze and facial expressions in real time.
[1227] Sensor: A device that detects user movements, such as an accelerometer or gyro sensor.
[1228] software:
[1229] OpenCV: A library for capturing and processing video data from cameras.
[1230] EmotionEngine: A system for analyzing emotions from a user's facial expressions and voice.
[1231] Cloud server: A server for data analysis and answer generation, utilizing natural language processing technology and generative AI models.
[1232] Natural language processing techniques: For example, using advanced language models such as GPT-3.
[1233] Process flow:
[1234] 1. Data collection methods:
[1235] The smart glasses' built-in cameras and sensors collect user behavioral and emotional data.
[1236] 2. Data analysis methods:
[1237] The collected data is analyzed on the cloud server to understand the user's current behavior and emotions and generate relevant questions.
[1238] 3. Display means:
[1239] The generated question list is displayed on the smart glasses display, and the user can select a question by eye gaze.
[1240] 4. Answer generation means:
[1241] Appropriate answers to questions selected by the user are generated on a cloud server and provided to the smart glasses either displayed or via voice.
[1242] Examples:
[1243] Example 1: Supermarket use:
[1244] User: Looking for food in the supermarket.
[1245] Smart glasses: Use a camera to track the user's gaze and analyze emotions.
[1246] Cloud server: Generates food-related question candidates (e.g., "What is the allergy information for this ingredient?").
[1247] Smart glasses: A list of questions is displayed and the user selects one by looking at the screen.
[1248] Cloud server: Generates answers to selected questions and displays them to the user.
[1249] Example prompt sentence:
[1250] "Please provide information about the ingredients of foods that may be of interest to users."
[1251] User gaze information: Fresh vegetable section
[1252] User sentiment: Showing interest
[1253] Questions generated: "What is the nutritional value of this vegetable?", "What dishes can it be used in?", "What is the allergy information?"
[1254] In this way, by using the system of the present invention, users can smoothly resolve their questions in physical stores, thereby improving the customer experience.
[1255] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1256] Step 1:
[1257] The device uses a built-in camera and sensors to capture the user's actions and gaze. The camera captures the user's gaze direction and facial expressions in real time, and the sensor detects the user's actions (e.g., reaching for a product). This captured data is collected by a data collection means. The input is the user's physical actions and gaze information, and the output is the captured action and gaze data.
[1258] Step 2:
[1259] The device sends the collected behavioral data and gaze data to a cloud server. The cloud server receives this data and analyzes it using data analysis means. The data analysis means analyzes the user's emotions from their gaze direction, movements, and facial expressions, and predicts questions that are likely to interest the user. For example, if the user is looking at a particular food item, it generates a question related to that food (e.g., "What are the ingredients in this food?"). The input is behavioral data and gaze data, and the output is a list of predicted questions.
[1260] Step 3:
[1261] The server generates questions based on the analysis results and sends them to the device in the form of a list. The device displays the received question list on the smart glasses display. The user visually checks the displayed question list and fixes their gaze on questions that interest them. The input is the predicted question list, and the output is the displayed question list.
[1262] Step 4:
[1263] When a user fixates their gaze on a question, the device detects the gaze information and determines that the user has selected the question. The selected question is then sent back to the cloud server, which then generates an appropriate answer for the question. The input is the question selected by the user, and the output is the generated answer.
[1264] Step 5:
[1265] The cloud server uses natural language processing technology and generative AI models (e.g., GPT-3) to generate answers to questions selected by the user. This answer is generated by searching for relevant information from a database and formatting it in an appropriate linguistic format. Furthermore, it takes into account the user's emotional information and generates an answer in a tone appropriate to the emotion. The input is the question selected by the user and emotional data, and the output is the generated answer.
[1266] Step 6:
[1267] The device receives the answer from the cloud server and provides it to the user either visually or audibly using speech synthesis technology. This allows the user to quickly obtain an answer to their question. The input is the generated answer, and the output is the answer provided to the user.
[1268] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1269] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1270] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1271] [Fourth embodiment]
[1272] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1273] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1274] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1275] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1276] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1277] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1278] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1279] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1280] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1281] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1282] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1283] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1284] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1285] This invention provides a system that efficiently supports users in acquiring information by detecting user behavior, generating predicted questions, and providing appropriate answers to those questions. This system uses sensors and cameras to capture the user's behavior and gaze, analyzes the data using a cloud server, generates appropriate questions, displays them to the user, and generates and provides answers to the selected questions.
[1286] Program processing flow
[1287] This system operates based on the interaction between the "terminal" (smart glasses), the "server" on the cloud, and the "user."
[1288] Device: Using built-in sensors and cameras, the device captures the user's movements and gaze in real time. The device temporarily stores the collected data and sends it to a cloud server as needed.
[1289] Server: The cloud server receives and analyzes the user data sent from the device. The data analysis method generates predicted questions based on the user's current situation. The generated question list is then sent from the server to the device.
[1290] Terminal: The terminal displays the question list received from the server on the display. When the user selects a question candidate on the display with their gaze, the selected question is sent to the server.
[1291] Server: The server analyzes the question sent by the user and generates an appropriate answer using database search and natural language processing techniques. The generated answer is then sent to the device.
[1292] Terminal: The terminal provides the received answer to the user either visually on a display or audibly using speech synthesis technology.
[1293] Specific examples
[1294] Example 1: Use at a restaurant
[1295] User: Looking at a new menu at a restaurant.
[1296] Device: Captures menus with the built-in camera and tracks the user's gaze.
[1297] Server: Recognizes when the user is looking at the menu and generates relevant question suggestions (e.g., "What are the ingredients?", "How many calories?").
[1298] Device: The question candidates are displayed on the screen and the user selects one. The user selects a question such as "How many calories?"
[1299] Server: Retrieves the calorie information for the menu items from the database and generates the answer.
[1300] Terminal: Provides answers visually or audibly.
[1301] Example 2: Use during learning
[1302] User: Studying by reading a textbook.
[1303] Device: Built-in camera and sensors identify where you are reading and track your gaze.
[1304] Server: Generates relevant question candidates (e.g., "What is the background to this incident?", "Who are the main characters?") based on the content of the textbook.
[1305] Terminal: The question candidates are displayed on the screen and the user can select one. The user selects the question "What is the background of this incident?"
[1306] Server: Obtains background information about the incident from textbook content and related materials and generates answers.
[1307] Terminal: Provides answers visually or audibly.
[1308] The system of the present invention is a powerful support tool for users to quickly obtain appropriate information, and plays an important role in today's information-overloaded society. It streamlines the user's learning process and provides practical solutions to various questions in daily life.
[1309] The processing flow will be explained below.
[1310] Step 1:
[1311] Device: The device uses built-in sensors and cameras to capture the user's movements and gaze in real time. The sensors detect the user's movements (e.g., walking, standing, sitting, etc.), and the camera tracks the user's gaze direction. The captured data is temporarily stored on the device.
[1312] Step 2:
[1313] Terminal: Processes the collected data, extracts necessary information, and sends it to the cloud server. This transmitted data includes the user's current behavioral patterns and gaze information.
[1314] Step 3:
[1315] Server: The cloud server receives the transmitted data and uses data analysis methods to analyze the user's behavior and gaze data. Specifically, it predicts what information the user is paying attention to and what questions they are likely to have.
[1316] Step 4:
[1317] Server: Generates a list of predicted questions based on the analysis results. The list of questions includes questions that the user is likely to have in that situation (e.g., "What are the ingredients in this dish?"). The generated list of questions is then sent back to the device.
[1318] Step 5:
[1319] Terminal: The received question list is displayed on the display. The displayed questions are displayed so that they can be visually recognized within the user's field of vision.
[1320] Step 6:
[1321] User: Selects a question displayed on the display by gazing at it. For example, if a user fixates their gaze on the question "How many calories does it have?" for a certain period of time, that question is recognized as selected.
[1322] Step 7:
[1323] The terminal sends the selected question to the server. The sent data includes the selected question as well as the user's current situation and context information.
[1324] Step 8:
[1325] Server: The server analyzes the received question and generates an appropriate answer, searching for relevant information from a database and using natural language processing techniques to create an answer in an easy-to-read format.
[1326] Step 9:
[1327] Server: Sends the generated answer to the terminal, which may be presented visually or audibly.
[1328] Step 10:
[1329] Terminal: Provides the received answer to the user, either visually displaying the answer on a display or audibly using speech synthesis technology.
[1330] These are the specific processing steps of the "Chatty Glasses" system. By capturing the user's behavior and gaze in real time, and using that information to predict and generate appropriate questions and provide answers, the system helps users obtain information more efficiently.
[1331] Example 1
[1332] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1333] In modern society, it is important for users to quickly obtain appropriate information for efficient decision-making and learning. However, it is not easy to find the information they need from a vast amount of information. There is also a lack of mechanisms that can accurately identify when a user has a question and quickly provide an appropriate answer to that question. The present invention aims to solve these problems by providing a system that improves the efficiency of information acquisition and supports users in resolving their questions.
[1334] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1335] In this invention, the server includes a data collection means including a sensor and a camera for detecting user behavior, a communication means for transmitting the detected user behavior data to the cloud server, a data analysis means for analyzing the transmitted data and generating predicted questions, a display means for displaying the generated question list to the user and allowing the user to select one with their line of sight, and an answer generation means for generating and providing an appropriate answer to the question selected by the user, thereby enabling the user to efficiently obtain information and quickly resolve their question.
[1336] "Data collection means" refers to devices and systems that include sensors and cameras for detecting user behavior.
[1337] "Communication means" refers to the technology or mechanism for transmitting detected user behavior data to a cloud server.
[1338] "Data analysis means" refers to software or algorithms that analyze data sent to the cloud server and generate predicted questions.
[1339] The "display means" is a device or interface that displays the generated questions to the user and allows the user to select them with their line of sight.
[1340] The "answer generation means" is software or a device that generates and provides an appropriate answer to a question selected by a user.
[1341] A "sensor" is a device that detects a user's actions and gaze, and can sense changes in movement and gaze.
[1342] A "camera" is a photographing device for visually capturing a user's line of sight and actions.
[1343] A "cloud server" is a server that can be accessed remotely for data analysis and management.
[1344] "Eye tracking" is a technology that tracks the movement of a user's eyes in real time.
[1345] "Natural language processing technology" is a technology for analyzing text data and understanding and generating human language.
[1346] "Database searching" is the technique or method of searching through a database to retrieve specific information.
[1347] This invention provides a system that efficiently supports users in acquiring information by detecting user behavior, generating predicted questions, and providing appropriate answers to those questions. This system uses sensors and cameras to capture the user's behavior and gaze, analyzes the data using a cloud server, generates appropriate questions, displays them to the user, and generates and provides answers to the selected questions.
[1348] 1. Hardware
[1349] Terminal
[1350] Use smart glasses, which have built-in sensors and cameras to detect the user's gaze and behavior, such as devices like Tobii Pro Glasses 3.
[1351] server
[1352] Use a server on the cloud. The server runs on a cloud service such as an AWS EC2 instance.
[1353] 2. Software
[1354] Data collection
[1355] An application is installed on the device to collect data in real time and send it to a cloud server. This application is developed using Unity, OpenCV, etc.
[1356] Data analysis
[1357] The cloud server uses programs such as Python, TensorFlow, and PyTorch to analyze the received data, which then generates predicted questions based on the user's behavioral data.
[1358] Display
[1359] The device displays the generated question list on its screen, and AR technology is used as the interface for users to select questions with their gaze.
[1360] Answer generation
[1361] To generate answers to selected questions, generative AI models and database search techniques are used. Specific examples include the transformers library and Elasticsearch. The generated answers are sent to the device and presented visually or using speech synthesis techniques. The speech synthesis uses the Google Text-to-Speech API.
[1362] 3. Specific Examples
[1363] Example 1: Use at a restaurant
[1364] User: Looking at a new menu at a restaurant.
[1365] Device: Captures menus with the built-in camera and tracks the user's gaze.
[1366] Server: Recognizes when the user is looking at the menu and generates relevant question candidates (e.g., "What are the ingredients?", "How many calories?").
[1367] Device: Candidate questions are displayed on the screen, and the user selects the question "How many calories?" with their gaze.
[1368] Server: Retrieves the calorie information for the menu items from the database and generates the answer.
[1369] Terminal: Provides answers visually or audibly.
[1370] Example prompt sentence:
[1371] "If a user is looking at a new menu item at a restaurant, how does the system pose questions to the user and provide answers?"
[1372] Example 2: Use during learning
[1373] User: Studying by reading a textbook.
[1374] Device: Built-in camera and sensors identify where you are reading and track your gaze.
[1375] Server: Generates relevant question candidates (e.g., "What is the background of this incident?", "Who are the main characters?") based on the content of the textbook.
[1376] Device: Candidate questions are displayed on the screen, and the user selects a question such as "What is the background to this incident?" with their gaze.
[1377] Server: Obtains background information about the incident from textbook content and related materials and generates answers.
[1378] Terminal: Provides answers visually or audibly.
[1379] Example prompt sentence:
[1380] "When a user is studying a textbook, how does the system pose questions to the user and provide answers?"
[1381] The system of the present invention is a powerful support tool for users to quickly obtain appropriate information, and plays an important role in today's information-overloaded society. It streamlines the user's learning process and provides practical solutions to various questions in daily life.
[1382] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1383] Step 1:
[1384] Input: User behavior and gaze data
[1385] Processing: Data Collection
[1386] Output: User behavior and gaze data
[1387] The device uses sensors and cameras built into the smart glasses to capture the user's actions and gaze in real time. Specifically, the camera captures the user's gaze direction, and the sensor detects movement. For example, when a user picks up a book in a bookstore, their movement and gaze are captured. This allows detailed behavioral and gaze data to be collected.
[1388] Step 2:
[1389] Input: User behavior and gaze data
[1390] Action: Send data
[1391] Output: Data sent to the cloud server
[1392] The device temporarily stores the collected behavioral and gaze data in its internal memory and then transmits the data to a cloud server via Wi-Fi or Bluetooth, using the TLS encryption protocol to protect the data transmission. Specifically, the device generates collected data packets and sequentially transmits them to a designated endpoint on the cloud server.
[1393] Step 3:
[1394] Input: Data sent to the cloud server
[1395] Processing: Data analysis
[1396] Output: Generated question list
[1397] The server receives the data sent to the cloud server and analyzes it using programs such as Python and TensorFlow. Specifically, it analyzes the user's behavioral and gaze data to infer the user's intentions and interests. Based on this, it generates predicted questions. This analysis process uses natural language processing (NLP) technology and machine learning models. For example, if a user is interested in a particular book in a bookstore, it generates questions such as, "What are the reviews for this book?" and "Who is the author?"
[1398] Step 4:
[1399] Input: Generated question list
[1400] Action: Display Question
[1401] Output: The question displayed to the user
[1402] The server sends the generated question list to the device. The device displays the received question list on its display. The user selects a question using eye movements or a pointer on the interface. Specifically, the system uses AR technology to visually display the question list, allowing the user to visually select a question.
[1403] Step 5:
[1404] Input: User selected question
[1405] Process: Send selected data
[1406] Output: Selection data sent to the cloud server
[1407] The user selects a question displayed on the display by looking at it. The device detects this selection and sends the information to the cloud server. The selected question is identified by gaze detection, and the data is sent back to the cloud server.
[1408] Step 6:
[1409] Input: Selection data sent to the cloud server
[1410] Processing: Answer generation
[1411] Output: The generated answer
[1412] The server analyzes the question selected by the user and generates an appropriate answer. This is done using database search and generative AI models (e.g., the transformers library). Specifically, Elasticsearch is used to retrieve the necessary information from the database, and an NLP model generates an answer based on that information. For example, in response to the question, "What are the reviews of this book?", review information is retrieved from a book review database and a concise review is generated.
[1413] Step 7:
[1414] Input: Generated answer
[1415] Action: Provide a response
[1416] Output: The answer provided to the user
[1417] The server sends the generated answer to the device. The device then displays the received answer to the user visually on a display or provides it audibly using speech synthesis technology. Specifically, the device uses the Google Text-to-Speech API to synthesize speech and provide the answer to the user audibly. The answer is also displayed visually on the smart glasses display. For example, the answer might be, "This book has excellent reviews, especially its storytelling."
[1418] (Application example 1)
[1419] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1420] In today's brick-and-mortar stores, users are required to quickly and efficiently obtain the information they need from a large selection of products. In particular, if users are unable to easily obtain detailed product information and reviews when choosing a product, their motivation to purchase is likely to decrease. Furthermore, if the information provided by the store is insufficient, it can affect users' decision-making, resulting in a negative impact on sales. To solve these problems, a system is needed that allows users to obtain product information in real time.
[1421] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1422] In this invention, the server includes a means for capturing gazes using a camera and a sensor built into the smart glasses in a physical store and analyzing the gazes on the cloud server to provide information about products, a means for displaying a list of questions received from the cloud server on the display of the smart glasses and selecting an appropriate question based on the user's gaze, and a means for the smart glasses to retrieve information about the selected question from a database and provide it to the user using voice synthesis technology or a display, thereby enabling the user to quickly and efficiently obtain detailed information about products.
[1423] A "sensor" is a device that detects physical changes and converts them into electrical signals.
[1424] A "camera" is a device for capturing images or video.
[1425] The "data collection means" is the part of the system that includes sensors and cameras to detect user activity.
[1426] The "data analysis means" is the part of the system that analyzes collected user behavior data and generates predicted questions.
[1427] The "display means" is a device that displays the generated questions to the user and allows the user to select one by line of sight.
[1428] The "answer generator" is the part of the system that generates and provides appropriate answers to user-selected questions.
[1429] A "physical store" is a physical store that a user visits to purchase products.
[1430] "Smart glasses" are eyeglass-type devices that display information and have built-in cameras and sensors to track the user's gaze.
[1431] A "cloud server" is a server for processing and storing data remotely via the Internet.
[1432] "Gaze capture" means tracking the user's gaze in real time and acquiring the data.
[1433] A "question list" is a collection of multiple questions generated based on user behavior data.
[1434] A "database" is a system that systematically manages data and allows for quick retrieval of necessary information.
[1435] A "display" is a device for visually displaying information.
[1436] "Speech synthesis technology" is a technology that converts text data into voice data.
[1437] This invention is a system that allows users to quickly and efficiently obtain detailed product information in physical stores. The system uses smart glasses to track the user's gaze and provides relevant information via a cloud server. To implement this invention, smart glasses, a cloud server, a database, and voice synthesis technology are used in combination.
[1438] Hardware and software used
[1439] Hardware:
[1440] Smart glasses: Equipped with a built-in camera and sensors, they capture the user's gaze in real time and have communication capabilities to send gaze data to a cloud server.
[1441] Display: Built into the smart glasses, it visually displays questions and answers to the user.
[1442] software:
[1443] Cloud server: Receives the user's gaze data, analyzes the data, generates necessary questions, and retrieves appropriate information from the database to display to the user.
[1444] Data analysis method: Runs on a cloud server and is responsible for analyzing gaze data and generating questions.
[1445] Question list generator: Generates appropriate questions based on the user's gaze information and sends this question list to the smart glasses.
[1446] Answer generation method: Based on the question selected by the user, information is retrieved from the database and an answer is generated to provide to the user. Natural language processing techniques (e.g., OpenAI GPT model) are used.
[1447] Speech synthesis technology: Technology for providing answers to users via voice (e.g., Google TTS).
[1448] What the system does
[1449] Overview of the process flow:
[1450] 1. Smart Glasses:
[1451] It uses built-in cameras and sensors to capture the user's gaze in real time.
[1452] The acquired gaze data is sent to a cloud server.
[1453] 2. Cloud Server:
[1454] The received gaze data is analyzed to identify information about the product that the user is paying attention to.
[1455] A list of relevant questions is generated and sent to the smart glasses.
[1456] 3. Smart Glasses:
[1457] The received question list is displayed on a display, and the user selects a question with their line of sight.
[1458] The question selected by the user is sent to the cloud server.
[1459] 4. Cloud Server:
[1460] Retrieve answers to selected questions from a database.
[1461] Based on the acquired information, an appropriate answer is generated and sent to the smart glasses.
[1462] 5. Smart Glasses:
[1463] The received answer is either visually displayed on a display or provided aloud using speech synthesis technology.
[1464] Specific examples
[1465] Suppose a user is looking at a particular product (e.g., organic food) in a physical store (e.g., a supermarket). The user's gaze data is captured by the smart glasses and sent to a cloud server. The cloud server analyzes the gaze data to identify the product the user is looking at and generates a list of questions related to this product. Questions such as "What are the nutritional content of this product?" and "What is the best way to cook it?" are displayed on the display of the user's smart glasses. When the user selects a question with their gaze, the selected question is sent to the cloud server. The cloud server retrieves an appropriate answer to the selected question from a database and sends it to the user's smart glasses. Finally, the smart glasses either display the retrieved answer visually or provide the answer audibly using speech synthesis technology.
[1466] Prompt Sentence Examples
[1467] Describe the process flow for a smart glasses application that tracks a user's gaze in real time and generates and answers questions related to specific supermarket items. Include specific details about the hardware, software, and data processing methods used.
[1468] In this way, the present invention allows users to quickly and efficiently obtain detailed information about products in a physical store.
[1469] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1470] Step 1:
[1471] The smart glasses' cameras and sensors capture the user's gaze.
[1472] Input: User gaze data
[1473] How it works: The smart glasses' built-in cameras and sensors track the user's gaze in real time and capture data to identify the object they are looking at.
[1474] Output: Captured gaze data
[1475] Step 2:
[1476] The smart glasses transmit the gaze data to a cloud server.
[1477] Input: Captured gaze data
[1478] How it works: The smart glasses transmit the acquired gaze data to a cloud server via the internet.
[1479] Output: Gaze data received by the cloud server
[1480] Step 3:
[1481] The cloud server analyzes the received gaze data and identifies information about the product that the user is paying attention to.
[1482] Input: Gaze data received by the cloud server
[1483] Operation: The data analysis means on the cloud server analyzes the gaze data and performs calculations to identify the object (product) that the gaze is directed at.
[1484] Output: Identified product information (product ID, etc.)
[1485] Step 4:
[1486] The cloud server generates a list of related questions based on the identified products.
[1487] Input: Identified product information
[1488] How it works: A data analysis tool on a cloud server uses a generative AI model to generate questions related to the identified products. These questions are formulated based on product characteristics.
[1489] Output: Generated question list
[1490] Step 5:
[1491] The smart glasses display the received question list on the display.
[1492] Input: Generated question list
[1493] How it works: A list of questions sent from a cloud server is displayed on the smart glasses, allowing the user to select a question with their gaze.
[1494] Output: Question list displayed on the screen
[1495] Step 6:
[1496] The user selects a question with their gaze.
[1497] Input: Question list displayed on the screen
[1498] How it works: The user uses their gaze to select a question of interest. The smart glasses sense the selected question.
[1499] Output: Selected Question
[1500] Step 7:
[1501] The smart glasses send the selected question to a cloud server.
[1502] Input: Selected Question
[1503] How it works: The smart glasses send the selected question over the internet to a cloud server.
[1504] Output: Questions received by the cloud server
[1505] Step 8:
[1506] The cloud server retrieves answers to the received questions from the database.
[1507] Input: A question received by the cloud server
[1508] Operation: The answer generation means on the cloud server searches the database for relevant information and processes the data to generate an answer.
[1509] Output: Retrieved answer information
[1510] Step 9:
[1511] The cloud server transmits the acquired answer information to the smart glasses.
[1512] Input: Retrieved answer information
[1513] How it works: The cloud server sends the generated answer information to the smart glasses via the internet.
[1514] Output: Answer information received by smart glasses
[1515] Step 10:
[1516] The smart glasses provide the received answer information to the user.
[1517] Input: Answer information received by smart glasses
[1518] How it works: The smart glasses will either visually display the answer information on a display or provide it aloud using speech synthesis technology.
[1519] Output: Answer information provided to the user
[1520] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1521] This invention provides a system that efficiently supports users in acquiring information by detecting a user's behavior, gaze, and even emotions, generating predicted questions, and providing appropriate answers to those questions. This system uses sensors and a camera to capture the user's behavior and gaze, recognizes the user's emotions using an emotion engine, analyzes the data using a cloud server, generates appropriate questions, displays them to the user, and generates and provides answers to the selected questions.
[1522] Program processing flow
[1523] This system operates based on the interaction between the "terminal" (smart glasses), the "server" on the cloud, and the "user." In addition, an emotion engine is built in to recognize the user's emotions.
[1524] Device: The device uses built-in sensors and cameras to capture the user's movements and gaze in real time. The sensors detect the user's movements (e.g., walking, standing, sitting, etc.), and the camera tracks the user's gaze direction and facial expressions. The captured data is temporarily stored on the device, and the emotion engine analyzes the user's facial expressions and voice to recognize emotions.
[1525] Server: The cloud server receives the data sent from the device and analyzes it in multiple ways. In particular, it uses data analysis tools to analyze the user's current behavior, gaze, and emotions to predict questions the user is likely to have in that situation. It then generates a list of predicted questions. The generated list of questions is sent from the cloud server to the device.
[1526] Terminal: The terminal displays the question list received from the server on a display, and the displayed questions are displayed so that they can be visually recognized within the user's field of vision.
[1527] User: The user selects a question displayed on the display by gazing at it. For example, if the user fixates their gaze on the question "How many calories does it have?" for a certain period of time, that question is recognized as selected.
[1528] Device: Sends the selected question to the server. The sent data includes the selected question, as well as the user's current situation and context information (including emotions).
[1529] Server: The server analyzes the received question and generates an appropriate answer. It searches for relevant information from a database and uses natural language processing techniques to create an answer in an easy-to-read format. It also adjusts the answer based on the emotions recognized by the emotion engine.
[1530] Terminal: Provides the received answer to the user. The answer is displayed visually on the display or provided aloud using speech synthesis technology. The answer is provided in an appropriate expression according to the user's emotion.
[1531] Specific examples
[1532] Example 1: Use at a restaurant
[1533] User: Looking at a new menu at a restaurant.
[1534] Device: Captures menus with the built-in camera and tracks the user's gaze and facial expressions.
[1535] Server: Recognizes when the user is looking at the menu and generates relevant question candidates (e.g., "What are the ingredients?", "How many calories?"). If the user's facial expression indicates interest, it prioritizes displaying relevant information.
[1536] Device: Candidate questions are displayed on the screen, and the user selects a question such as "How many calories?" with their gaze.
[1537] Server: Retrieves the calorie information of the menu item from the database and generates an answer. If the user looks surprised, the answer is provided in a softer tone.
[1538] Terminal: Provides answers visually or audibly.
[1539] Example 2: Use during learning
[1540] User: Studying by reading a textbook.
[1541] On the device: Built-in cameras and sensors identify where the user is reading, track their gaze and facial expressions, and an emotion engine recognizes the user's emotions (e.g., confusion, interest, etc.).
[1542] Server: Generates relevant question candidates (e.g., "What is the background of this incident?", "Who are the main characters?") based on the textbook content. If it determines that the user is confused, it prioritizes question candidates that include explanations.
[1543] Device: Candidate questions are displayed on the screen, and the user selects a question such as "What is the background to this incident?" with their gaze.
[1544] Server: Retrieves background information from textbook content and related materials and generates an answer. If the user is dissatisfied, the answer is expanded.
[1545] Terminal: Provides answers visually or audibly.
[1546] The system of this invention optimizes the provision of information by capturing the user's behavior, gaze, and emotions from multiple angles, and supports the user in resolving questions in their studies and daily life. By providing information according to the user's emotions, more effective learning and information acquisition are realized.
[1547] The processing flow will be explained below.
[1548] Step 1:
[1549] Device: The device uses built-in sensors and cameras to capture the user's movements, gaze, and facial expressions in real time. The sensors detect the user's movements (e.g., walking, standing, sitting, etc.), and the camera tracks the user's gaze direction and facial expressions. The captured data is temporarily stored on the device.
[1550] Step 2:
[1551] Device: Processes the collected behavior, gaze, and facial expression data, extracts necessary information, and sends it to the cloud server. This transmitted data includes the user's current behavioral patterns, gaze information, and facial expression data.
[1552] Step 3:
[1553] Server: The cloud server receives the transmitted data and analyzes it using data analysis methods and an emotion engine. In particular, it determines what information the user is paying attention to and what emotion (interest, confusion, delight, etc.) the user is feeling.
[1554] Step 4:
[1555] Server: Based on the analysis results, predicts the questions the user is likely to have in that situation. Generates a list of predicted questions that also reflects the user's emotional state. The generated list of questions is sent from the cloud server to the device.
[1556] Step 5:
[1557] Terminal: The received question list is displayed on the display. The displayed questions are displayed so that they can be visually recognized within the user's field of vision.
[1558] Step 6:
[1559] User: Selects a question displayed on the display by gazing at it. For example, if a user fixates their gaze on the question "How many calories does it have?" for a certain period of time, that question is recognized as selected.
[1560] Step 7:
[1561] The terminal sends the selected question to the server. The sent data includes the selected question, as well as the user's current situation and emotional information.
[1562] Step 8:
[1563] Server: The server analyzes the received question and generates an appropriate answer. It searches for relevant information from a database and creates an answer using natural language processing technology. It also adjusts the expression and level of detail of the answer based on the emotions recognized by the emotion engine.
[1564] Step 9:
[1565] Server: Sends the generated answer to the terminal, which may be presented as a visual display or audio.
[1566] Step 10:
[1567] Terminal: Provides the received answer to the user. The answer is displayed visually on the display or provided aloud using speech synthesis technology. The answer is provided in an appropriate expression according to the user's emotions.
[1568] Step 11:
[1569] Server: User responses and selection patterns are logged and used for later analysis, forming a feedback loop to improve the accuracy of future predictions.
[1570] These are the specific processing steps of this system. By comprehensively capturing the user's behavior, gaze, and emotions, and providing predicted questions and answers based on this information in real time, the system streamlines the user's learning process and information acquisition.
[1571] Example 2
[1572] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1573] Conventional information acquisition systems have had difficulty efficiently providing users with the information they seek. In particular, they have been unable to sense data such as the user's gaze and emotions in real time, pose appropriate questions, and provide answers tailored to the user. As a result, the user's information acquisition process has been subject to significant time and effort.
[1574] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes a data collection means including a sensor and a camera for detecting the user's movements, gaze data, and emotions, a data analysis means for analyzing the detected user's movement data, gaze data, and emotion data and generating predicted questions, a display means for displaying a question list generated based on the analysis means to the user and allowing the user to select a question with their gaze, an answer generation means for generating and providing an appropriate answer to the question selected by the user, and an emotion analysis means for adjusting the generated answer based on the user's emotion data. This makes it possible to efficiently obtain the information the user desires.
[1575] A "data collection means" is a device that includes sensors and cameras to detect the user's movements, gaze, and emotions.
[1576] The "data analysis means" is a device or software that analyzes the detected user's motion data, gaze data, and emotion data and generates predicted questions.
[1577] The "display means" is a device or software that displays the question list generated based on the analysis means to the user and allows the user to select a question by line of sight.
[1578] The "answer generation means" is a device or software that generates and provides an appropriate answer to a question selected by a user.
[1579] The "emotion analysis means" is a device or software that adjusts the response generated based on the user's emotional data.
[1580] This invention is a system that efficiently supports users in acquiring information by detecting a user's movements, gaze, and emotions, generating predicted questions, and providing appropriate answers to those questions. This system uses sensors and a camera to capture the user's movements and gaze, recognizes the user's emotions using an emotion engine, analyzes the data using a cloud server, generates appropriate questions and displays them to the user, and generates and provides answers to the selected questions.
[1581] The following components are primarily involved in the operation of the system:
[1582] Device: Uses a device (such as smart glasses) equipped with sensors and a camera to capture the user's movements, gaze, and emotions in real time. The sensor tracks the user's movements, and the camera tracks the direction of gaze and facial expressions, and this data is temporarily stored on the device. The device is also equipped with an emotion engine that analyzes the user's facial expressions and voice to recognize emotions.
[1583] Server: The cloud server receives the data sent from the device and uses data analysis methods to analyze the user's current behavior, gaze, and emotions. This analysis predicts the questions the user is likely to have in that situation and generates a list of questions. This list of questions is then sent from the cloud server to the device.
[1584] Display method: The device displays the question list received from the server on the display. The display format is an interactive format using pop-ups or eye tracking. For example, when a user is looking at a new menu, questions such as "How many calories are in this dish?" and "What are the ingredients?" are displayed.
[1585] User: The user selects a question by gazing at it. By fixing their gaze for a certain period of time, the system recognizes the user's selection and determines that the question has been selected.
[1586] Answer generation method: The selected question and context information (user's facial expression and current situation information) are sent from the device to the server. The server analyzes the question, searches for relevant information from a database, and generates an answer using natural language processing technology. It also adjusts the answer based on an emotion engine, for example, providing a softer tone if the user is surprised.
[1587] Providing an answer: The device provides the received answer to the user either visually on the display or audibly using speech synthesis technology. For example, it may visually display or audibly guide the user by saying, "This dish is 500 kcal."
[1588] Specific examples
[1589] Example 1: Use at a restaurant
[1590] 1. User: Looking at a new menu at a restaurant.
[1591] 2. Device: Capture menus with the built-in camera and track the user's gaze and facial expressions.
[1592] 3. Server: Recognizes that the user is looking at the menu and generates relevant question candidates ("What are the ingredients?", "How many calories?").
[1593] 4. Device: Question candidates are displayed on the screen, and the user selects a question such as "How many calories?" with their gaze.
[1594] 5. Server: Retrieves the calorie information of the menu item from the database and generates an answer. If the user looks surprised, the answer is provided in a softer tone.
[1595] 6. Terminal: Provides answers visually or audibly.
[1596] Example 2: Use during learning
[1597] 1. User: Reading and studying a textbook.
[1598] 2. Device: The built-in camera and sensors identify where the user is reading, track their gaze and facial expressions, and the emotion engine recognizes the user's emotions (e.g., confusion, interest, etc.).
[1599] 3. Server: Generates candidate questions ("What is the background of this incident?", "Who are the main characters?") based on the content of the textbook. If it determines that the user is confused, it prioritizes candidate questions that include explanations.
[1600] 4. Terminal: Candidate questions are displayed on the screen, and the user selects a question such as "What is the background to this incident?" with their gaze.
[1601] 5. Server: Retrieves background information from textbook content and related materials and generates an answer. If the user is dissatisfied, the answer is expanded.
[1602] 6. Terminal: Provides answers visually or audibly.
[1603] Prompt Sentence Examples
[1604] "A user is looking at a new menu item at a restaurant. If the user looks interested, suggest possible questions about the menu."
[1605] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1606] Step 1: Capture data with your device
[1607] Device: Activates built-in sensors and cameras to capture the user's movements, gaze, and facial expressions in real time. Inputs include the user's physical movements, gaze, and facial expressions, which are detected by the sensors and camera and temporarily stored as digital data. Specific movements include detecting walking and sitting, tracking gaze direction, and analyzing facial expressions. Outputs include movement data, gaze data, and facial expression data.
[1608] Step 2: Send data to the server
[1609] Terminal: Sends the captured data to the cloud server. The inputs are the movement data, gaze data, and facial expression data generated in step 1, and are sent to the server in real time. Specific operations include packaging and sending the data. The output is the data sent to the server.
[1610] Step 3: Data analysis and question generation by the server
[1611] Server: Analyzes the received data to determine the user's current behavior, gaze, and emotions. Inputs include motion data, gaze data, and facial expression data sent from the device, which are analyzed using machine learning algorithms and an emotion analysis engine. Specific operations include analyzing behavioral patterns, identifying gaze focus, and recognizing emotions. The output is a list of questions the user is likely to have in the situation.
[1612] Step 4: Sending the Question List from the Server to the Device
[1613] Server: Sends the question list generated by the analysis to the terminal. The input is the question list generated in step 3, which is sent to the terminal. Specifically, the server packages and sends the question list. The output is the question list sent to the terminal.
[1614] Step 5: Ask a question on the device
[1615] Terminal: The terminal displays the question list received from the server. The input is the question list sent from the server, which is displayed in a format that is easy for the user to view. Specific operations include displaying questions in a pop-up or interactive format. The output is the questions displayed on the screen.
[1616] Step 6: User selects question
[1617] User: Selects a question candidate displayed on the display with their gaze. The input is a list of questions displayed on the device, and the user fixates their gaze on a specific question for a certain period of time. Specific operations include processing the gaze tracking data and selecting a question. The output is the selected question.
[1618] Step 7: Sending the selected question to the server
[1619] Terminal: Sends the user's current situation and context information to the server along with the selected question. The inputs are the user's selected question and related context information (motion data, gaze data, and facial expression data), which are then sent to the server. Specific operations include packaging and sending the data. The output is the question and context information sent to the server.
[1620] Step 8: Server Generates Answer
[1621] Server: Analyzes the received question, retrieves relevant information from a database, and generates an answer. The input is the question and context information sent from the device, which is analyzed using natural language processing technology. Specific operations include database search, natural language generation, and response adjustment using an emotion engine. The output is the generated answer.
[1622] Step 9: Provide answers on your device
[1623] Terminal: Provides the received answer to the user. The input is the answer sent from the server, which is displayed on a screen or provided as voice using speech synthesis technology. Specific operations include visual display and speech synthesis. The output is the answer provided to the user.
[1624] (Application example 2)
[1625] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1626] Currently, when a customer has a question about a product in a physical store, it often takes time and effort to find a store employee to resolve the question. It can also be difficult for store staff to respond to all questions quickly and accurately. Furthermore, there is no system that can predict what customers are wondering about and provide information to answer those questions. This can lead to a poor customer experience and affect store sales.
[1627] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1628] In this invention, the server includes a data collection means including a sensor and a camera for detecting user behavior, a data analysis means for analyzing the detected user behavior data and emotion data and generating predicted questions, a display means for displaying the generated questions to the user and allowing the user to select one with their line of sight, and an answer generation means for generating and providing an appropriate answer to the question selected by the user. This makes it possible to provide quick and appropriate answers to questions that users have in physical stores.
[1629] "User" refers to an individual who uses the system.
[1630] An "action" is when a user performs a specific physical action or visual attention.
[1631] A "sensor" is a device that senses physical movements and environmental information and acquires data based on that information.
[1632] A "camera" is an optical device for capturing images or videos.
[1633] The "data collection means" is a means for collecting user behavioral data and emotional data using sensors and cameras.
[1634] "Emotion data" refers to emotional information sensed from the user's facial expressions, voice, etc.
[1635] "Data analysis means" refers to means for analyzing collected behavioral and emotional data and generating predicted questions.
[1636] The "display means" is a device that displays the generated questions to the user and allows the user to select one by line of sight.
[1637] The "answer generation means" is a means for generating and providing an appropriate answer to a question selected by a user.
[1638] "Natural language processing technology" is a technology that enables computers to understand and generate natural language.
[1639] A "generative AI model" is an artificial intelligence model that generates new information based on previously learned data.
[1640] The present invention provides a system that detects a user's behavior, gaze, and emotions, generates predicted questions based on the detected behavior, and provides appropriate answers to the questions. This system includes a data collection means, a data analysis means, a display means, and an answer generation means.
[1641] The system is implemented using the following hardware and software.
[1642] Hardware:
[1643] Smart glasses: Equipped with a built-in camera and display, they capture the user's gaze and facial expressions in real time.
[1644] Sensor: A device that detects user movements, such as an accelerometer or gyro sensor.
[1645] software:
[1646] OpenCV: A library for capturing and processing video data from cameras.
[1647] EmotionEngine: A system for analyzing emotions from a user's facial expressions and voice.
[1648] Cloud server: A server for data analysis and answer generation, utilizing natural language processing technology and generative AI models.
[1649] Natural language processing techniques: For example, using advanced language models such as GPT-3.
[1650] Process flow:
[1651] 1. Data collection methods:
[1652] The smart glasses' built-in cameras and sensors collect user behavioral and emotional data.
[1653] 2. Data analysis methods:
[1654] The collected data is analyzed on the cloud server to understand the user's current behavior and emotions and generate relevant questions.
[1655] 3. Display means:
[1656] The generated question list is displayed on the smart glasses display, and the user can select a question by eye gaze.
[1657] 4. Answer generation means:
[1658] Appropriate answers to questions selected by the user are generated on a cloud server and provided to the smart glasses either displayed or via voice.
[1659] Examples:
[1660] Example 1: Supermarket use:
[1661] User: Looking for food in the supermarket.
[1662] Smart glasses: Use a camera to track the user's gaze and analyze emotions.
[1663] Cloud server: Generates food-related question candidates (e.g., "What is the allergy information for this ingredient?").
[1664] Smart glasses: A list of questions is displayed and the user selects one by looking at the screen.
[1665] Cloud server: Generates answers to selected questions and displays them to the user.
[1666] Example prompt sentence:
[1667] "Please provide information about the ingredients of foods that may be of interest to users."
[1668] User gaze information: Fresh vegetable section
[1669] User sentiment: Showing interest
[1670] Questions generated: "What is the nutritional value of this vegetable?", "What dishes can it be used in?", "What is the allergy information?"
[1671] In this way, by using the system of the present invention, users can smoothly resolve their questions in physical stores, thereby improving the customer experience.
[1672] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1673] Step 1:
[1674] The device uses a built-in camera and sensors to capture the user's actions and gaze. The camera captures the user's gaze direction and facial expressions in real time, and the sensor detects the user's actions (e.g., reaching for a product). This captured data is collected by a data collection means. The input is the user's physical actions and gaze information, and the output is the captured action and gaze data.
[1675] Step 2:
[1676] The device sends the collected behavioral data and gaze data to a cloud server. The cloud server receives this data and analyzes it using data analysis means. The data analysis means analyzes the user's emotions from their gaze direction, movements, and facial expressions, and predicts questions that are likely to interest the user. For example, if the user is looking at a particular food item, it generates a question related to that food (e.g., "What are the ingredients in this food?"). The input is behavioral data and gaze data, and the output is a list of predicted questions.
[1677] Step 3:
[1678] The server generates questions based on the analysis results and sends them to the device in the form of a list. The device displays the received question list on the smart glasses display. The user visually checks the displayed question list and fixes their gaze on questions that interest them. The input is the predicted question list, and the output is the displayed question list.
[1679] Step 4:
[1680] When a user fixates their gaze on a question, the device detects the gaze information and determines that the user has selected the question. The selected question is then sent back to the cloud server, which then generates an appropriate answer for the question. The input is the question selected by the user, and the output is the generated answer.
[1681] Step 5:
[1682] The cloud server uses natural language processing technology and generative AI models (e.g., GPT-3) to generate answers to questions selected by the user. This answer is generated by searching for relevant information from a database and formatting it in an appropriate linguistic format. Furthermore, it takes into account the user's emotional information and generates an answer in a tone appropriate to the emotion. The input is the question selected by the user and emotional data, and the output is the generated answer.
[1683] Step 6:
[1684] The device receives the answer from the cloud server and provides it to the user either visually or audibly using speech synthesis technology. This allows the user to quickly obtain an answer to their question. The input is the generated answer, and the output is the answer provided to the user.
[1685] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1686] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1687] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1688] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1689] FIG. 9 illustrates an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and behaviors arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1690] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1691] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1692] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1693] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1694] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1695] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1696] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1697] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1698] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1699] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1700] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1701] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1702] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1703] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1704] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1705] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1706] The following is further disclosed regarding the above embodiment.
[1707] (Claim 1)
[1708] a data collection means including sensors and cameras for detecting user behavior;
[1709] a data analysis means for analyzing the detected user behavior data and generating predicted questions;
[1710] a display means for displaying the generated questions to a user and allowing the user to select a question by line of sight;
[1711] an answer generation means for generating and providing an appropriate answer to a question selected by a user;
[1712] A system including:
[1713] (Claim 2)
[1714] 10. The system of claim 1, wherein the sensor and camera include means for tracking the user's gaze.
[1715] (Claim 3)
[1716] 2. The system of claim 1, wherein the answer generating means includes means for generating answers using natural language processing techniques.
[1717] "Example 1"
[1718] (Claim 1)
[1719] a data collection means including sensors and cameras for detecting user behavior;
[1720] a communication means for transmitting the detected user behavior data to a cloud server;
[1721] data analysis means for analyzing the transmitted data and generating predicted questions;
[1722] a display means for displaying the generated question list to a user and allowing the user to select a question by line of sight;
[1723] an answer generation means for generating and providing an appropriate answer to a question selected by a user;
[1724] A system including:
[1725] (Claim 2)
[1726] 10. The system of claim 1, wherein the sensor and camera include means for tracking the user's gaze.
[1727] (Claim 3)
[1728] 10. The system of claim 1, wherein the answer generation means includes means for generating answers using natural language processing techniques and database searches.
[1729] "Application Example 1"
[1730] (Claim 1)
[1731] a data collection means including sensors and cameras for detecting user behavior;
[1732] a data analysis means for analyzing the detected user behavior data and generating predicted questions;
[1733] a display means for displaying the generated questions to a user and allowing the user to select a question by line of sight;
[1734] an answer generation means for generating and providing an appropriate answer to a question selected by a user;
[1735] In a physical store, a means for capturing gazes using a camera and a sensor built into the smart glasses and analyzing them on a cloud server in order to provide information about products;
[1736] a means for displaying the question list received from the cloud server on a display of the smart glasses and selecting an appropriate question based on the user's gaze;
[1737] means for the smart glasses to retrieve information for the selected question from a database and provide it to the user using voice synthesis technology or a display;
[1738] A system including:
[1739] (Claim 2)
[1740] 10. The system of claim 1, wherein the sensor and camera include means for tracking the user's gaze.
[1741] (Claim 3)
[1742] 2. The system of claim 1, wherein the answer generating means includes means for generating answers using natural language processing techniques.
[1743] "Example 2: Combining Emotion Engines"
[1744] (Claim 1)
[1745] a data collection means including sensors and cameras for detecting the user's movements, gaze, and emotions;
[1746] a data analysis means for analyzing the detected user's motion data, gaze data, and emotion data and generating predicted questions;
[1747] a display means for displaying the question list generated based on the analysis means to a user and allowing the user to select a question by line of sight;
[1748] an answer generation means for generating and providing an appropriate answer to a question selected by a user;
[1749] emotion analysis means for adjusting the generated answers based on the user's emotion data;
[1750] A system including:
[1751] (Claim 2)
[1752] 10. The system of claim 1, including means for tracking a user's gaze in real time using sensors and cameras, and for selecting a question based on the user's gaze.
[1753] (Claim 3)
[1754] 2. The system of claim 1, wherein the answer generating means includes means for generating an answer to a user-selected question in an easy-to-read format using natural language processing techniques.
[1755] "Application example 2 when combining emotion engines"
[1756] (Claim 1)
[1757] a data collection means including sensors and cameras for detecting user behavior;
[1758] a data analysis means for analyzing the detected user behavior data and emotion data and generating predicted questions;
[1759] a display means for displaying the generated questions to a user and allowing the user to select a question by line of sight;
[1760] an answer generation means for generating and providing an appropriate answer to a question selected by a user;
[1761] A system including:
[1762] (Claim 2)
[1763] 10. The system of claim 1, wherein the sensors and cameras include means for performing eye tracking and emotion analysis of the user.
[1764] (Claim 3)
[1765] 2. The system of claim 1, wherein the answer generation means includes means for generating answers using natural language processing techniques and a generative AI model. [Explanation of symbols]
[1766] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. a data collection means including sensors and cameras for detecting user behavior; a data analysis means for analyzing the detected user behavior data and generating predicted questions; a display means for displaying the generated questions to a user and allowing the user to select a question by line of sight; an answer generation means for generating and providing an appropriate answer to a question selected by a user; A system including:
2. The system of claim 1 , wherein the sensor and camera include means for tracking the user's gaze.
3. 2. The system of claim 1, wherein the answer generating means includes means for generating answers using natural language processing techniques.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A