system
The system addresses inefficiencies in eye-tracking technologies by tracking gaze, retrieving relevant information, and enhancing suggestion accuracy through feedback, enabling efficient and optimized user interactions.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-10-21
- Publication Date
- 2026-05-07
AI Technical Summary
Existing technologies are inadequate in effectively utilizing eye-tracking data for information presentation and improving the accuracy of suggestions based on user feedback, leading to inefficiencies in acquiring necessary information and optimizing user actions.
A system that tracks a user's gaze in real-time, identifies the object of attention, retrieves relevant information from a database, and displays it prioritized, while incorporating feedback to improve suggestion accuracy, using devices like smart glasses or goggles equipped with sensors and a server for data processing.
Enables efficient acquisition of information and optimization of user actions by providing hands-free access to relevant data and improving suggestion accuracy through continuous user feedback loops.
Smart Images

Figure 2026074883000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, and includes steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] The present invention provides means for efficiently recognizing an object of attention and quickly obtaining related information for a user facing a lot of information and tasks. In particular, it is an object to solve the problem that it is difficult to immediately obtain necessary information in a situation of excessive information, and to provide support for a user to effectively optimize actions.
Means for Solving the Problems
[0005] This invention provides a means for tracking a user's gaze in real time and identifying the object the user is fixated on based on that gaze. Furthermore, it quickly retrieves information related to the identified object from a database and displays it to the user in order of priority, thereby enabling effective acquisition of necessary information. In addition, it provides a means for continuous user support by incorporating a function that suggests optimizations to user behavior and improves the accuracy of those suggestions based on subsequent feedback.
[0006] A "user" refers to a person who uses this system to receive information and action suggestions based on their gaze.
[0007] "Eye tracking" is a technology that detects a user's eye movements in real time and identifies the direction and object they are looking at.
[0008] A "point of attention" refers to an object or information that a user's gaze is directed towards for a certain period of time, indicating a particular interest or concern.
[0009] "Related information" refers to data and knowledge directly related to the subject of focus, and is presented in accordance with the user's needs.
[0010] "Behavioral optimization" is the process of suggesting the necessary actions and steps to enable users to achieve their goals efficiently and effectively.
[0011] "Feedback" refers to the evaluations and reactions that users give to the information and suggestions they provide, and is used to improve the accuracy of the system.
[0012] "Recommendation accuracy" is an indicator that shows the degree to which a system is able to make effective recommendations that are tailored to the user's situation and needs. [Brief explanation of the drawing]
[0013] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2]It is a conceptual diagram showing an example of the main functions of a data processing device and a smart device according to the first embodiment. [Figure 3] It is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] It is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which a plurality of emotions are mapped. [Figure 10] It shows an emotion map to which a plurality of emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Example 2 when an emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when an emotion engine is combined.
Embodiments for Carrying Out the Invention
[0014] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.
[0015] First, the terms used in the following description will be explained.
[0016] In the following embodiments, the numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.
[0017] In the following embodiments, the numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0018] In the following embodiments, the numbered storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, and the like.
[0019] In the following embodiments, the numbered communication I / F (Interface) is an interface including a communication processor and an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark), and the like.
[0020] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0021] [First Embodiment]
[0022] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0023] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0024] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0025] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0026] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0027] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0028] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0029] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0030] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0031] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0032] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0033] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0034] To implement this invention, an infrastructure is provided that allows for eye tracking when the user wears a specific goggle device. The goggles are equipped with sensors to detect gaze in real time and collect the data. The server has the function of using this gaze data to identify the object the user is fixated on and immediately retrieving relevant information from a database.
[0035] The terminal processes relevant information obtained from the server and generates action suggestions tailored to the user's situation. For example, if the user is working, it analyzes the progress of their work based on their gaze and recommends high-priority tasks. If the user is searching for something, it uses their gaze to display lists or areas that have not yet been checked.
[0036] In this system, information is displayed on the goggles' screen according to the user's gaze direction, allowing users to easily manage information using only their eye movements. This enables users to acquire information and optimize their actions without using their hands. Furthermore, users can provide feedback on the information and suggestions they receive, and this feedback is collected on a server to improve the accuracy of future suggestions.
[0037] As a concrete example, consider a scenario where a user uses this system while working at a desk. The server detects when the user's gaze is fixed on a document for a certain period of time and searches for and provides the latest reports and materials related to that document. The terminal analyzes these materials and suggests the next steps the user should complete. This allows the user to work more efficiently.
[0038] Thus, the present invention can significantly improve the efficiency of user information acquisition and actions through a user interface that utilizes eye gaze.
[0039] The following describes the processing flow.
[0040] Step 1:
[0041] The server acquires user gaze data from the eye-tracking sensors built into the goggles. This data includes the direction, movement, and duration of gaze fixation.
[0042] Step 2:
[0043] The server analyzes the acquired gaze data to identify the area the user is fixating on. It estimates the object of gaze based on the duration of gaze fixation and the degree of pupil convergence.
[0044] Step 3:
[0045] The device analyzes images of objects based on the identified gaze area. Using image recognition technology, it detects textual information and specific objects, and identifies the objects from a database.
[0046] Step 4:
[0047] The server searches the database for information related to the identified object and retrieves that information. The retrieved information is then organized according to its content and relevance.
[0048] Step 5:
[0049] The device generates action suggestions tailored to the user's situation based on the acquired relevant information. If the user is working, it will generate suggestions for priority tasks; if the user is wandering or exploring, it will generate suggestions indicating unexplored areas.
[0050] Step 6:
[0051] The user reviews relevant information and action suggestions displayed on the goggles' screen. Based on the suggestions, they select an action and continue working.
[0052] Step 7:
[0053] Users provide feedback to the server regarding the presented information and suggestions. This feedback is used to improve the accuracy of the next suggestion generation process.
[0054] (Example 1)
[0055] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0056] There is a growing need for eye-tracking interfaces to enable users to efficiently acquire information and optimize their actions. However, existing technologies are inadequate in their ability to effectively utilize eye-tracking data for information presentation and in their use of feedback to improve the accuracy of suggestions. The challenge lies in resolving this issue and improving the user experience.
[0057] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0058] In this invention, the server includes means for providing a user interface equipped with a gaze detection device, means for identifying an object the user is fixated on and retrieving related information from a storage device, means for processing information acquired by the terminal and generating suggestions to optimize the user's actions, and means for recording the user's response to the generated suggestions and improving the accuracy of the generated products. This enables the user to acquire information and optimize their actions based on their gaze.
[0059] A "gaze detection device" is a device that senses the direction and movement of a user's gaze in real time and collects that information.
[0060] A "user interface" is a device or software designed to allow users to interact with an information system intuitively.
[0061] "Means of identifying the target" refers to the process of determining and identifying the object or area that the user is focusing on, based on gaze data.
[0062] "Means for retrieving information from a storage device" refers to methods for quickly obtaining information related to a specified object from a database or similar source.
[0063] "Means of generating suggestions" refers to the process of presenting the user with the most suitable actions or solutions based on the information acquired.
[0064] "Means for recording responses" refers to a function that saves the user's feedback on the suggestions presented, in order to improve the accuracy of the system's suggestions in the future.
[0065] To implement this invention, the user must first wear a device that detects their gaze. This device is equipped with sensors to collect the user's gaze information in real time. This data is transmitted to a server via a network.
[0066] The server receives and analyzes the gaze data, using advanced data processing software. From the gaze data, it identifies the object the user is fixated on and searches for related information within the storage device. Database query techniques and search algorithms are utilized in this process.
[0067] The device receives information sent from the server and generates action suggestions that are best suited to the user's current situation. This is done using a generative AI model, which takes eye-tracking data and related information as prompt text to generate specific suggestions.
[0068] For example, if a user fixates their gaze on a specific document while working at their desk, the server searches for the latest research and reference materials related to that document. The terminal analyzes this information and suggests the user's next course of action, such as "read the next chapter" or "review relevant meeting materials." The information is displayed on the goggles' screen according to the user's gaze, and user feedback can also be collected.
[0069] As an example of a prompt, you can input a message into the AI model in the format of, "Based on the user's eye-tracking data, generate optimal action suggestions to maximize work efficiency."
[0070] Thus, this invention enables efficient information provision and behavioral optimization by utilizing the user's gaze.
[0071] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0072] Step 1:
[0073] The user wears a device that detects their gaze. This is a preparatory step to enable the collection of gaze data. The gaze sensor detects the user's gaze direction in real time and generates data. The input is the direction of the user's gaze, and the output is the coordinates and movement data of the gaze.
[0074] Step 2:
[0075] The server receives gaze data. Here, the received gaze data is analyzed to identify the object the user is fixated on. In this process, the gaze data is converted into digital coordinates, and the database is queried to find matching information. The input is gaze coordinate data, and the output is the name of the identified object and related information.
[0076] Step 3:
[0077] The server searches the database for relevant information based on the identified target. The server efficiently accesses the database and quickly retrieves the relevant information. The data processing performed here is carried out using an information retrieval algorithm. The input is the name of the identified target, and the output is information related to that target.
[0078] Step 4:
[0079] The terminal receives information from the server and generates action suggestions tailored to the user's current situation. Using a generation AI model, it takes user eye-tracking data and related information as prompts and generates optimal suggestions based on those prompts. The input is related information and eye-tracking data, and the output is action suggestions.
[0080] Step 5:
[0081] The terminal displays generated action suggestions on the user's goggles display. The user uses their gaze to confirm the presented information and provide feedback. The input is the action suggestion, and the output is the information displayed to the user. In addition, the user's feedback is sent to the server as new data input.
[0082] (Application Example 1)
[0083] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0084] In modern factory environments, rapid and accurate information provision is essential to maintaining and improving worker efficiency and quality. However, traditional methods require manual work to access work instructions and documents, interrupting workers' work. Furthermore, as the amount of information increases, it becomes difficult for workers to quickly find the information they need. In such situations, work efficiency decreases, potentially negatively impacting productivity. In addition, there is a lack of systems for collecting feedback on given instructions and suggestions, which significantly limits the improvement of the accuracy of suggestions.
[0085] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0086] In this invention, the server includes means for tracking the direction in which the user is looking, means for identifying the object being looked at and retrieving related information, and means for suggesting an optimization of actions based on the identified information. This makes it possible to instantly obtain necessary information using eye movements in the work environment and optimize actions without using hands.
[0087] "Means for tracking the direction in which a user is looking" refers to a device or system that detects the direction of a user's gaze in real time and collects that information as digital data.
[0088] "Means for identifying objects under gaze and retrieving related information" refers to a device or system that uses sensors or other means to identify objects that fall within a user's line of sight and retrieves information related to those objects from a database or similar source.
[0089] A "means for proposing optimized behavior" refers to a device or system that has the function of generating and presenting the most effective action plan to the user based on acquired information and circumstances.
[0090] "Means for collecting feedback on provided proposals and improving the accuracy of subsequent proposals" refers to a device or system that has the function of collecting responses and opinions from users to previously received proposals and analyzing them to improve the accuracy of future proposals.
[0091] "Means for displaying procedural information and instructions related to an object of gaze in a work environment" refers to a device or system that has the function of displaying procedures and instructions related to an object or task identified by the user's gaze within the user's field of view in a workplace such as a factory.
[0092] "Means for instantly searching a data repository and presenting it in order of importance" refers to a device or system that has the function of quickly searching a database containing information related to an object being viewed and visually presenting that information to the user based on its importance.
[0093] The system for implementing this invention consists of smart glasses worn by the user and a server. The smart glasses are equipped with eye-tracking sensors that recognize the direction in which the user is looking in real time. This eye-tracking data is transmitted to the server via a network. The server analyzes the received eye-tracking data to identify the object being looked at. This analysis uses indicators such as gaze fixation time and concentration level.
[0094] The server searches its data repository for information related to the identified object. The information retrieval is performed quickly, and the results are ranked by importance. This ensures that the most relevant data is presented to the user preferentially. The presented information and suggestions are visually displayed on the smart glasses' screen, allowing the user to efficiently obtain information hands-free.
[0095] Furthermore, the server collects feedback from users on the information and suggestions presented. This feedback is used to improve the accuracy of future suggestions. The feedback includes information such as which information was helpful and which suggestions were appropriate.
[0096] As a concrete example, consider a scenario where a factory worker uses this system. When the worker looks at a new control panel, the smart glasses automatically display the assembly procedure. After completing the task, the worker can provide feedback on whether the displayed procedure was helpful, and this information is used to improve the accuracy of future procedure suggestions. In this way, eye-tracking interfaces significantly improve efficiency and accuracy in factory operations.
[0097] Example of a prompt:
[0098] "When workers are looking at new parts, please display the relevant manual on their glasses."
[0099] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0100] Step 1:
[0101] Sensors built into the smart glasses track the user's gaze and collect eye-tracking data. The direction of gaze, duration of fixed gaze, and level of concentration are measured as input. This data is converted into a digital format and sent to a server.
[0102] Step 2:
[0103] The server analyzes the received gaze data to identify the object being looked at. It receives gaze data as input, compares it with object information registered in the database, and outputs the specific object the gaze was directed at. This analysis uses a specific algorithm to process the data.
[0104] Step 3:
[0105] The server searches the data repository for information related to the identified object. It takes the object name as input and retrieves relevant information via a query. The search results are ranked based on importance, and the most relevant information is provided as output. This search and ranking process utilizes database queries and importance ranking algorithms.
[0106] Step 4:
[0107] Information transmitted from the server is displayed on the user's smart glasses. This displayed information includes relevant procedural data and important instructions. The user can intuitively confirm the information obtained through their gaze. Based on the relevant information received from the server as input, visual data is output to the smart glasses' display.
[0108] Step 5:
[0109] Users provide feedback on the displayed suggestions. Feedback is entered through the smart glasses interface, and the server collects it to improve the accuracy of future suggestions. Natural language processing technology is used to analyze the feedback, inputting the feedback data and performing specific data processing and calculations to identify areas for improvement.
[0110] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0111] This invention is a system that combines tracking the user's gaze direction and identifying the object of their gaze with an emotion engine that enables emotion recognition. The user wears a special pair of goggles, which contain an eye-tracking sensor and a biosensor that functions as the emotion engine.
[0112] The server acquires eye-tracking data and biometric data, and analyzes the user's gaze direction and emotional state. Eye-tracking analysis identifies the object of gaze, and the emotion engine uses biometric signals to estimate the user's emotional state, such as stress levels and concentration levels.
[0113] The device searches a database for information corresponding to the identified object of attention and the user's emotional state. The emotion engine then uses this information to generate information and suggestions best suited to the user's emotions. For example, if the user is stressed, it prioritizes displaying information that promotes relaxation; if the user is focused, it prioritizes displaying suggestions that enhance efficiency.
[0114] This system provides users with information and action suggestions through the goggles' display. The device collects user feedback and sends this data to a server to improve the accuracy of the suggestions. As a specific example, consider a scenario where a user uses the system during a meeting. The server detects that the user is intently looking at meeting materials and prioritizes displaying important information related to those materials. Furthermore, if the emotion engine detects a high level of concentration, it provides additional information and action plans to support further productivity improvements.
[0115] This technology allows users to instantly obtain necessary information using their gaze and emotions, enabling them to intuitively and efficiently optimize their actions.
[0116] The following describes the processing flow.
[0117] Step 1:
[0118] The server acquires the user's gaze data and biometric data in real time from the eye-tracking sensor and biosensor built into the goggles. This allows the server to observe the user's gaze direction and physiological responses.
[0119] Step 2:
[0120] The server identifies the direction and object of gaze from the eye-tracking data. This identification is achieved by analyzing the duration of gaze fixation and pupil movement.
[0121] Step 3:
[0122] The server analyzes the user's emotional state based on biometric data. The emotion engine analyzes biometric signals such as heart rate and skin temperature to estimate whether the user is excited, relaxed, focused, or otherwise in a particular state.
[0123] Step 4:
[0124] The device searches a database for information related to the gaze target identified by the eye-tracking data. During this process, high-priority information is selected and retrieved immediately.
[0125] Step 5:
[0126] The server creates appropriate action suggestions for the user based on the emotional state and related information analyzed by the emotion engine. For example, if the server determines that the user is stressed, it will suggest relaxation methods or encourage deep breathing.
[0127] Step 6:
[0128] Users view relevant information and action suggestions through the goggles' display. Information display based on eye gaze and emotion reduces user burden and allows for intuitive operation.
[0129] Step 7:
[0130] Users send feedback to the server regarding the information and suggestions provided. This feedback is recorded and analyzed by the server to improve accuracy in future instances.
[0131] (Example 2)
[0132] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0133] In modern society, users are required to quickly and accurately retrieve necessary information from a vast amount of data. However, conventional information retrieval methods have the problem of not being able to provide dynamic and optimal information that responds to changes in the user's gaze and emotions. To solve this problem, a system is needed that analyzes the user's gaze direction and emotional state in real time and provides information based on that analysis immediately.
[0134] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0135] In this invention, the server includes means for tracking the user's gaze and analyzing its direction, means for estimating the emotional state from biometric information, and means for retrieving information from a database based on the estimated gaze target and emotional state. This makes it possible to analyze the user's current gaze direction and emotional state in real time and quickly provide optimal information accordingly.
[0136] "Eye-tracking" is a technology that detects the movement of a user's eyeballs and understands the direction of their gaze in real time.
[0137] "Biometric information" refers to physiological data such as the user's heart rate and skin electrical responses, and is used to infer emotional states and other related information.
[0138] "Estimating emotional state" is a method of analyzing a user's emotional state from biometric information and evaluating factors such as stress levels and concentration levels.
[0139] "Searching for information in a database" is the process of retrieving relevant data from a collection of aggregated information based on specific criteria.
[0140] "Generating suggestions" is the process of creating optimal action instructions and information presentations for a user based on their analyzed state.
[0141] "Collecting feedback" is the activity of gathering user reactions and behaviors as data and analyzing it to improve the accuracy of future information provision.
[0142] This system starts operating when the user puts on special goggles. The goggles have built-in eye-tracking sensors and biosensors, which are used to acquire the user's gaze data and biometric information in real time. The eye-tracking sensors sense the movement of the user's eyeballs and provide technology to identify the direction of their gaze. The biosensors collect physiological data such as heart rate and skin electrical activity, and the server uses a generated AI model to estimate the user's emotional state.
[0143] The server analyzes the received data and evaluates what the user is looking at and their emotional state (e.g., stress level, concentration level). The analysis results are sent to the terminal, which then searches the database for relevant information based on the analyzed gaze and emotional state. This database search is performed quickly and efficiently, providing information that best suits the user's needs.
[0144] In this system, the generative AI model plays a crucial role in learning from the user's past behavior and feedback to generate personalized suggestions for each individual user. These suggestions and information are presented through the goggles' display.
[0145] As a concrete example, let's assume a user is attending a public lecture. In this scenario, the server tracks the user's gaze and detects when they are focusing on a particular slide. The sentiment engine estimates that the user is highly interested in the lecture's content, and the device provides relevant materials and additional background information. This allows the user to understand the lecture more deeply.
[0146] An example of a prompt is, "Design a system that generates and presents personalized information in real time based on user emotions and eye-tracking data." This example demonstrates how a system can improve the user experience.
[0147] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0148] Step 1:
[0149] The user puts on a special goggle and begins data collection. An eye-tracking sensor built into the goggles detects the user's eye movements in real time and acquires the direction as digital data. In addition, biosensors collect physiological data such as the user's heart rate and skin electrical activity. The input is the user's gaze direction and biometric information, and the output is provided as a data stream in its original form.
[0150] Step 2:
[0151] The device transmits gaze data and biometric information acquired from the goggles to the server via wireless communication (e.g., Bluetooth or Wi-Fi). The input data consists of gaze position data and biometric signal data, and the output is the result of transmitting that data to the server.
[0152] Step 3:
[0153] The server analyzes the received gaze data to determine the user's gaze direction. In this process, the server uses an algorithm to identify the specific object the user is fixated on based on the gaze data. The input is gaze data, and the output is the identified object of fixation.
[0154] Step 4:
[0155] The server simultaneously analyzes biometric information and uses a generative AI model to estimate the user's emotional state. This involves performing the necessary data calculations to determine emotional states such as stress levels and concentration levels. The input is biometric information, and the output is an evaluation of the user's emotional state.
[0156] Step 5:
[0157] The terminal searches the database for relevant information based on the object of attention and emotional state received from the server. At this stage, the terminal selects high-priority information and prepares it to be provided to the user. The input is the identified object of attention and emotional state, and the output is the relevant information retrieved from the database.
[0158] Step 6:
[0159] The device provides the user with generated suggestions and information through the goggles' display. The user can visually confirm this. The input is the searched information and suggestions, and the output appears as the displayed content.
[0160] Step 7:
[0161] The device records the user's reactions to the information received through the goggles and sends them to the server as feedback. This feedback is used to improve future suggestions. The input is the user's reactions and feedback, and the output is the feedback data sent to the server.
[0162] (Application Example 2)
[0163] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0164] The challenge is to provide a system that supports efficient and accurate purchasing decisions by eliminating the confusion caused by information overload and numerous choices when users select products in physical stores. Furthermore, it is necessary to improve the shopping experience by providing personalized information tailored to the user's interests and emotional state.
[0165] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0166] In this invention, the server includes means for tracking the direction in which the user is looking, means for identifying the object being looked at and retrieving related information, means for suggesting optimized actions based on the identified information, means for collecting feedback on the provided suggestions and improving the accuracy of subsequent suggestions, means for detecting the user's emotional state and displaying information corresponding to that emotion, and means for prioritizing and displaying multiple pieces of information by combining gaze and emotion. This enables intuitive and effective information provision based on the user's gaze and emotions, significantly improving the user's purchasing experience.
[0167] "Means of tracking the direction a user is looking" refers to devices or technologies that detect the direction a user is looking and acquire it as digital data.
[0168] "Means for identifying objects being watched and retrieving related information" refers to technology that identifies objects or areas in the user's line of sight and extracts and provides information about them from a digital database.
[0169] "Means of suggesting optimized actions based on identified information" refers to technologies that present recommendations and action plans to support user decision-making based on acquired information.
[0170] "Means of collecting feedback on provided proposals and improving the accuracy of future proposals" refers to technologies that collect user reactions and opinions, analyze them, and increase the success rate of future proposals.
[0171] "A means of detecting a user's emotional state and displaying information corresponding to that emotion" refers to a technology that analyzes a user's emotions from their biological responses and behavior and provides information that is optimal for that state.
[0172] "A method for prioritizing and displaying multiple pieces of information by combining eye movements and emotions" is a technology that evaluates the importance of information and prioritizes its display by simultaneously analyzing the user's eye movements and emotional data.
[0173] The system that realizes this application example uses a special goggle that incorporates means to track the direction the user is looking and means to detect the user's emotional state. The goggle is equipped with an eye-tracking sensor and a biosensor, which collect data in real time.
[0174] The server is responsible for analyzing gaze data and biometric data transmitted from the goggles. It uses gaze data to identify the object of gaze and an emotion engine to estimate the user's emotional state based on biometric signals. Because the server is equipped with advanced data analysis algorithms, it can accurately assess the user's emotional state, such as stress levels and concentration.
[0175] The device searches a vast database for relevant information based on analysis results sent from the server and presents content tailored to the user's emotions. For example, if a user is looking at something they are interested in, product information and discount campaign information related to that item will be displayed immediately.
[0176] As a concrete example, imagine a user wearing goggles in a physical store and browsing the electronics section. The server detects from eye-tracking data and biometric information that the user is focusing on a specific product and has a deep interest in it. The device then displays promotions and further product information related to this product on the goggles' display, supporting the user's purchasing decision. Through this entire process, the user can enjoy a more intuitive and personalized shopping experience.
[0177] An example of a prompt message is, "Identify where the user is focusing their attention on their smartphone, analyze their feelings towards that product, and generate potential promotions to display." Based on this prompt message, a generation AI model is used to provide optimal information suggestions tailored to the user's focus and emotional state.
[0178] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0179] Step 1:
[0180] The user puts on a special pair of goggles. The eye-tracking sensor and biosensor built into the goggles are activated, collecting the user's eye-tracking data and biometric information in real time. This data is transmitted to a server via wireless communication.
[0181] Step 2:
[0182] The server analyzes the received gaze data to identify the object the user is fixated on. Specifically, it calculates the gaze vector and evaluates its intensity and duration to identify a particular object or area. As a result of the analysis, it outputs the ID and location information of the object being fixed on.
[0183] Step 3:
[0184] The server evaluates the user's emotional state based on biometric data. By analyzing biosignals such as heart rate, skin temperature, and electrical activity, it estimates whether the user is relaxed or focused. The emotion engine performs this process and outputs the evaluation result as the user's emotional state.
[0185] Step 4:
[0186] The device searches a database for relevant information based on the ID and location information of the object being watched, obtained from the server, and the results of the emotional state evaluation. At this time, a generative AI model is used to select high-priority information and information most appropriate to the emotion. Based on the search results, display candidates are generated and formatted into a list of information to be presented.
[0187] Step 5:
[0188] The device displays information based on a formatted list, optimized for the user's goggles display. Prioritized information is highlighted for greater importance. The user is also provided with options to directly interact with the displayed information.
[0189] Step 6:
[0190] Users provide feedback on the information and suggestions presented. This feedback is sent from the device to the server and stored in a database to improve accuracy for future use. For example, the user's actions and reactions to each piece of information are recorded as logs.
[0191] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0192] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0193] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0194] [Second Embodiment]
[0195] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0196] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0197] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0198] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0199] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0200] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0201] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0202] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0203] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0204] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0205] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0206] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0207] To implement this invention, an infrastructure is provided that allows for eye tracking when the user wears a specific goggle device. The goggles are equipped with sensors to detect gaze in real time and collect the data. The server has the function of using this gaze data to identify the object the user is fixated on and immediately retrieving relevant information from a database.
[0208] The terminal processes relevant information obtained from the server and generates action suggestions tailored to the user's situation. For example, if the user is working, it analyzes the progress of their work based on their gaze and recommends high-priority tasks. If the user is searching for something, it uses their gaze to display lists or areas that have not yet been checked.
[0209] In this system, information is displayed on the goggles' screen according to the user's gaze direction, allowing users to easily manage information using only their eye movements. This enables users to acquire information and optimize their actions without using their hands. Furthermore, users can provide feedback on the information and suggestions they receive, and this feedback is collected on a server to improve the accuracy of future suggestions.
[0210] As a concrete example, consider a scenario where a user uses this system while working at a desk. The server detects when the user's gaze is fixed on a document for a certain period of time and searches for and provides the latest reports and materials related to that document. The terminal analyzes these materials and suggests the next steps the user should complete. This allows the user to work more efficiently.
[0211] Thus, the present invention can significantly improve the efficiency of user information acquisition and actions through a user interface that utilizes eye gaze.
[0212] The following describes the processing flow.
[0213] Step 1:
[0214] The server acquires user gaze data from the eye-tracking sensors built into the goggles. This data includes the direction, movement, and duration of gaze fixation.
[0215] Step 2:
[0216] The server analyzes the acquired gaze data to identify the area the user is fixating on. It estimates the object of gaze based on the duration of gaze fixation and the degree of pupil convergence.
[0217] Step 3:
[0218] The device analyzes images of objects based on the identified gaze area. Using image recognition technology, it detects textual information and specific objects, and identifies the objects from a database.
[0219] Step 4:
[0220] The server searches the database for information related to the identified object and retrieves that information. The retrieved information is then organized according to its content and relevance.
[0221] Step 5:
[0222] The device generates action suggestions tailored to the user's situation based on the acquired relevant information. If the user is working, it will generate suggestions for priority tasks; if the user is wandering or exploring, it will generate suggestions indicating unexplored areas.
[0223] Step 6:
[0224] The user reviews relevant information and action suggestions displayed on the goggles' screen. Based on the suggestions, they select an action and continue working.
[0225] Step 7:
[0226] Users provide feedback to the server regarding the presented information and suggestions. This feedback is used to improve the accuracy of the next suggestion generation process.
[0227] (Example 1)
[0228] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0229] There is a growing need for eye-tracking interfaces to enable users to efficiently acquire information and optimize their actions. However, existing technologies are inadequate in their ability to effectively utilize eye-tracking data for information presentation and in their use of feedback to improve the accuracy of suggestions. The challenge lies in resolving this issue and improving the user experience.
[0230] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0231] In this invention, the server includes means for providing a user interface equipped with a gaze detection device, means for identifying an object the user is fixated on and retrieving related information from a storage device, means for processing information acquired by the terminal and generating suggestions to optimize the user's actions, and means for recording the user's response to the generated suggestions and improving the accuracy of the generated products. This enables the user to acquire information and optimize their actions based on their gaze.
[0232] A "gaze detection device" is a device that senses the direction and movement of a user's gaze in real time and collects that information.
[0233] A "user interface" is a device or software designed to allow users to interact with an information system intuitively.
[0234] "Means of identifying the target" refers to the process of determining and identifying the object or area that the user is focusing on, based on gaze data.
[0235] "Means for retrieving information from a storage device" refers to methods for quickly obtaining information related to a specified object from a database or similar source.
[0236] "Means of generating suggestions" refers to the process of presenting the user with the most suitable actions or solutions based on the information acquired.
[0237] "Means for recording responses" refers to a function that saves the user's feedback on the suggestions presented, in order to improve the accuracy of the system's suggestions in the future.
[0238] To implement this invention, the user must first wear a device that detects their gaze. This device is equipped with sensors to collect the user's gaze information in real time. This data is transmitted to a server via a network.
[0239] The server receives and analyzes the gaze data, using advanced data processing software. From the gaze data, it identifies the object the user is fixated on and searches for related information within the storage device. Database query techniques and search algorithms are utilized in this process.
[0240] The device receives information sent from the server and generates action suggestions that are best suited to the user's current situation. This is done using a generative AI model, which takes eye-tracking data and related information as prompt text to generate specific suggestions.
[0241] For example, if a user fixates their gaze on a specific document while working at their desk, the server searches for the latest research and reference materials related to that document. The terminal analyzes this information and suggests the user's next course of action, such as "read the next chapter" or "review relevant meeting materials." The information is displayed on the goggles' screen according to the user's gaze, and user feedback can also be collected.
[0242] As an example of a prompt, you can input a message into the AI model in the format of, "Based on the user's eye-tracking data, generate optimal action suggestions to maximize work efficiency."
[0243] Thus, this invention enables efficient information provision and behavioral optimization by utilizing the user's gaze.
[0244] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0245] Step 1:
[0246] The user wears a device that detects their gaze. This is a preparatory step to enable the collection of gaze data. The gaze sensor detects the user's gaze direction in real time and generates data. The input is the direction of the user's gaze, and the output is the coordinates and movement data of the gaze.
[0247] Step 2:
[0248] The server receives gaze data. Here, the received gaze data is analyzed to identify the object the user is fixated on. In this process, the gaze data is converted into digital coordinates, and the database is queried to find matching information. The input is gaze coordinate data, and the output is the name of the identified object and related information.
[0249] Step 3:
[0250] The server searches the database for relevant information based on the identified target. The server efficiently accesses the database and quickly retrieves the relevant information. The data processing performed here is carried out using an information retrieval algorithm. The input is the name of the identified target, and the output is information related to that target.
[0251] Step 4:
[0252] The terminal receives information from the server and generates action suggestions tailored to the user's current situation. Using a generation AI model, it takes user eye-tracking data and related information as prompts and generates optimal suggestions based on those prompts. The input is related information and eye-tracking data, and the output is action suggestions.
[0253] Step 5:
[0254] The terminal displays generated action suggestions on the user's goggles display. The user uses their gaze to confirm the presented information and provide feedback. The input is the action suggestion, and the output is the information displayed to the user. In addition, the user's feedback is sent to the server as new data input.
[0255] (Application Example 1)
[0256] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0257] In modern factory environments, rapid and accurate information provision is essential to maintaining and improving worker efficiency and quality. However, traditional methods require manual work to access work instructions and documents, interrupting workers' work. Furthermore, as the amount of information increases, it becomes difficult for workers to quickly find the information they need. In such situations, work efficiency decreases, potentially negatively impacting productivity. In addition, there is a lack of systems for collecting feedback on given instructions and suggestions, which significantly limits the improvement of the accuracy of suggestions.
[0258] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0259] In this invention, the server includes means for tracking the direction in which the user is looking, means for identifying the object being looked at and retrieving related information, and means for suggesting an optimization of actions based on the identified information. This makes it possible to instantly obtain necessary information using eye movements in the work environment and optimize actions without using hands.
[0260] "Means for tracking the direction in which a user is looking" refers to a device or system that detects the direction of a user's gaze in real time and collects that information as digital data.
[0261] "Means for identifying objects under gaze and retrieving related information" refers to a device or system that uses sensors or other means to identify objects that fall within a user's line of sight and retrieves information related to those objects from a database or similar source.
[0262] A "means for proposing optimized behavior" refers to a device or system that has the function of generating and presenting the most effective action plan to the user based on acquired information and circumstances.
[0263] "Means for collecting feedback on provided proposals and improving the accuracy of subsequent proposals" refers to a device or system that has the function of collecting responses and opinions from users to previously received proposals and analyzing them to improve the accuracy of future proposals.
[0264] "Means for displaying procedural information and instructions related to an object of gaze in a work environment" refers to a device or system that has the function of displaying procedures and instructions related to an object or task identified by the user's gaze within the user's field of view in a workplace such as a factory.
[0265] "Means for instantly searching a data repository and presenting it in order of importance" refers to a device or system that has the function of quickly searching a database containing information related to an object being viewed and visually presenting that information to the user based on its importance.
[0266] The system for implementing this invention consists of smart glasses worn by the user and a server. The smart glasses are equipped with eye-tracking sensors that recognize the direction in which the user is looking in real time. This eye-tracking data is transmitted to the server via a network. The server analyzes the received eye-tracking data to identify the object being looked at. This analysis uses indicators such as gaze fixation time and concentration level.
[0267] The server searches its data repository for information related to the identified object. The information retrieval is performed quickly, and the results are ranked by importance. This ensures that the most relevant data is presented to the user preferentially. The presented information and suggestions are visually displayed on the smart glasses' screen, allowing the user to efficiently obtain information hands-free.
[0268] Furthermore, the server collects feedback from users on the information and suggestions presented. This feedback is used to improve the accuracy of future suggestions. The feedback includes information such as which information was helpful and which suggestions were appropriate.
[0269] As a concrete example, consider a scenario where a factory worker uses this system. When the worker looks at a new control panel, the smart glasses automatically display the assembly procedure. After completing the task, the worker can provide feedback on whether the displayed procedure was helpful, and this information is used to improve the accuracy of future procedure suggestions. In this way, eye-tracking interfaces significantly improve efficiency and accuracy in factory operations.
[0270] Example of a prompt:
[0271] "When workers are looking at new parts, please display the relevant manual on their glasses."
[0272] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0273] Step 1:
[0274] Sensors built into the smart glasses track the user's gaze and collect eye-tracking data. The direction of gaze, duration of fixed gaze, and level of concentration are measured as input. This data is converted into a digital format and sent to a server.
[0275] Step 2:
[0276] The server analyzes the received gaze data to identify the object being looked at. It receives gaze data as input, compares it with object information registered in the database, and outputs the specific object the gaze was directed at. This analysis uses a specific algorithm to process the data.
[0277] Step 3:
[0278] The server searches the data repository for information related to the identified object. It takes the object name as input and retrieves relevant information via a query. The search results are ranked based on importance, and the most relevant information is provided as output. This search and ranking process utilizes database queries and importance ranking algorithms.
[0279] Step 4:
[0280] Information transmitted from the server is displayed on the user's smart glasses. This displayed information includes relevant procedural data and important instructions. The user can intuitively confirm the information obtained through their gaze. Based on the relevant information received from the server as input, visual data is output to the smart glasses' display.
[0281] Step 5:
[0282] Users provide feedback on the displayed suggestions. Feedback is entered through the smart glasses interface, and the server collects it to improve the accuracy of future suggestions. Natural language processing technology is used to analyze the feedback, inputting the feedback data and performing specific data processing and calculations to identify areas for improvement.
[0283] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0284] The present invention is a system that combines an emotion engine enabling emotion recognition in addition to tracking the user's gaze direction and identifying the gaze target. The user wears dedicated goggles, which incorporate a gaze tracking sensor and a biosensor for functioning as an emotion engine.
[0285] The server acquires gaze data and biometric information data, and analyzes the user's gaze direction and emotional state. The gaze target is identified through gaze analysis, and the emotion engine estimates the user's emotional state, such as stress or concentration, using biosignals.
[0286] The terminal searches the database for information corresponding to the identified gaze target and emotional state. The emotion engine generates information and proposals most suitable for the user's emotion based on this information. For example, when the user is feeling stressed, information that can help relax is preferentially displayed, and when the user is concentrated, proposals to enhance efficiency are preferentially shown.
[0287] This system provides information and action proposals to the user through the display of the goggles. The terminal collects the user's feedback, sends this data to the server, and uses it to improve the accuracy of the proposals. As a specific example, consider the case where the user uses the system during a meeting. The server detects that the user is looking at the meeting materials and preferentially displays important information related to those materials. Furthermore, when the emotion engine senses a high level of concentration, it supports further productivity improvement by providing additional information and action plans.
[0288] With this technology, the user can immediately obtain the necessary information using their gaze and emotion, and can intuitively and efficiently optimize their actions.
[0289] The following describes the processing flow.
[0290] Step 1:
[0291] The server acquires the user's gaze data and biometric data in real time from the eye-tracking sensor and biosensor built into the goggles. This allows the server to observe the user's gaze direction and physiological responses.
[0292] Step 2:
[0293] The server identifies the direction and object of gaze from the eye-tracking data. This identification is achieved by analyzing the duration of gaze fixation and pupil movement.
[0294] Step 3:
[0295] The server analyzes the user's emotional state based on biometric data. The emotion engine analyzes biometric signals such as heart rate and skin temperature to estimate whether the user is excited, relaxed, focused, or otherwise in a particular state.
[0296] Step 4:
[0297] The device searches a database for information related to the gaze target identified by the eye-tracking data. During this process, high-priority information is selected and retrieved immediately.
[0298] Step 5:
[0299] The server creates appropriate action suggestions for the user based on the emotional state and related information analyzed by the emotion engine. For example, if the server determines that the user is stressed, it will suggest relaxation methods or encourage deep breathing.
[0300] Step 6:
[0301] Users view relevant information and action suggestions through the goggles' display. Information display based on eye gaze and emotion reduces user burden and allows for intuitive operation.
[0302] Step 7:
[0303] The user sends feedback on the provided information and suggestions to the server. This feedback is recorded and analyzed by the server to improve the accuracy in subsequent times.
[0304] (Example 2)
[0305] Next, Example 2 will be described. In the following description, the data processing device 12 is referred to as the "server", and the smart glasses 214 are referred to as the "terminal".
[0306] In modern society, users are required to have the ability to quickly and accurately obtain the necessary information from a vast amount of information. However, in the conventional information acquisition methods, there has been a problem that it is difficult to provide dynamic and optimal information according to the changes in the user's line of sight and emotions. To solve this problem, a system that analyzes the user's line of sight direction and emotional state in real time and provides information based on it immediately is necessary.
[0307] The specific processing by the specific processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0308] In this invention, the server includes means for tracking the user's line of sight and analyzing its direction, means for estimating the emotional state from the biological information, and means for searching for information from the database based on the estimated fixation target and emotional state. Thereby, it becomes possible to analyze the user's current line of sight direction and emotional state in real time and quickly provide optimal information according to it.
[0309] "Tracking the line of sight" is a technology for detecting the movement of the user's eyeballs and grasping the fixation direction in real time.
[0310] "Biological information" refers to physiological data such as the user's heart rate and galvanic skin response, and is information used to infer the emotional state and the like.
[0311] "Estimating emotional state" is a method of analyzing a user's emotional state from biometric information and evaluating factors such as stress levels and concentration levels.
[0312] "Searching for information in a database" is the process of retrieving relevant data from a collection of aggregated information based on specific criteria.
[0313] "Generating suggestions" is the process of creating optimal action instructions and information presentations for a user based on their analyzed state.
[0314] "Collecting feedback" is the activity of gathering user reactions and behaviors as data and analyzing it to improve the accuracy of future information provision.
[0315] This system starts operating when the user puts on special goggles. The goggles have built-in eye-tracking sensors and biosensors, which are used to acquire the user's gaze data and biometric information in real time. The eye-tracking sensors sense the movement of the user's eyeballs and provide technology to identify the direction of their gaze. The biosensors collect physiological data such as heart rate and skin electrical activity, and the server uses a generated AI model to estimate the user's emotional state.
[0316] The server analyzes the received data and evaluates what the user is looking at and their emotional state (e.g., stress level, concentration level). The analysis results are sent to the terminal, which then searches the database for relevant information based on the analyzed gaze and emotional state. This database search is performed quickly and efficiently, providing information that best suits the user's needs.
[0317] In this system, the generative AI model plays a crucial role in learning from the user's past behavior and feedback to generate personalized suggestions for each individual user. These suggestions and information are presented through the goggles' display.
[0318] As a concrete example, let's assume a user is attending a public lecture. In this scenario, the server tracks the user's gaze and detects when they are focusing on a particular slide. The sentiment engine estimates that the user is highly interested in the lecture's content, and the device provides relevant materials and additional background information. This allows the user to understand the lecture more deeply.
[0319] An example of a prompt is, "Design a system that generates and presents personalized information in real time based on user emotions and eye-tracking data." This example demonstrates how a system can improve the user experience.
[0320] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0321] Step 1:
[0322] The user puts on a special goggle and begins data collection. An eye-tracking sensor built into the goggles detects the user's eye movements in real time and acquires the direction as digital data. In addition, biosensors collect physiological data such as the user's heart rate and skin electrical activity. The input is the user's gaze direction and biometric information, and the output is provided as a data stream in its original form.
[0323] Step 2:
[0324] The device transmits gaze data and biometric information acquired from the goggles to the server via wireless communication (e.g., Bluetooth or Wi-Fi). The input data consists of gaze position data and biometric signal data, and the output is the result of transmitting that data to the server.
[0325] Step 3:
[0326] The server analyzes the received gaze data to determine the user's gaze direction. In this process, the server uses an algorithm to identify the specific object the user is fixated on based on the gaze data. The input is gaze data, and the output is the identified object of fixation.
[0327] Step 4:
[0328] The server simultaneously analyzes biometric information and uses a generative AI model to estimate the user's emotional state. This involves performing the necessary data calculations to determine emotional states such as stress levels and concentration levels. The input is biometric information, and the output is an evaluation of the user's emotional state.
[0329] Step 5:
[0330] The terminal searches the database for relevant information based on the object of attention and emotional state received from the server. At this stage, the terminal selects high-priority information and prepares it to be provided to the user. The input is the identified object of attention and emotional state, and the output is the relevant information retrieved from the database.
[0331] Step 6:
[0332] The device provides the user with generated suggestions and information through the goggles' display. The user can visually confirm this. The input is the searched information and suggestions, and the output appears as the displayed content.
[0333] Step 7:
[0334] The device records the user's reactions to the information received through the goggles and sends them to the server as feedback. This feedback is used to improve future suggestions. The input is the user's reactions and feedback, and the output is the feedback data sent to the server.
[0335] (Application Example 2)
[0336] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0337] The challenge is to provide a system that supports efficient and accurate purchasing decisions by eliminating the confusion caused by information overload and numerous choices when users select products in physical stores. Furthermore, it is necessary to improve the shopping experience by providing personalized information tailored to the user's interests and emotional state.
[0338] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0339] In this invention, the server includes means for tracking the direction in which the user is looking, means for identifying the object being looked at and retrieving related information, means for suggesting optimized actions based on the identified information, means for collecting feedback on the provided suggestions and improving the accuracy of subsequent suggestions, means for detecting the user's emotional state and displaying information corresponding to that emotion, and means for prioritizing and displaying multiple pieces of information by combining gaze and emotion. This enables intuitive and effective information provision based on the user's gaze and emotions, significantly improving the user's purchasing experience.
[0340] "Means of tracking the direction a user is looking" refers to devices or technologies that detect the direction a user is looking and acquire it as digital data.
[0341] "Means for identifying objects being watched and retrieving related information" refers to technology that identifies objects or areas in the user's line of sight and extracts and provides information about them from a digital database.
[0342] "Means of suggesting optimized actions based on identified information" refers to technologies that present recommendations and action plans to support user decision-making based on acquired information.
[0343] "Means of collecting feedback on provided proposals and improving the accuracy of future proposals" refers to technologies that collect user reactions and opinions, analyze them, and increase the success rate of future proposals.
[0344] "A means of detecting a user's emotional state and displaying information corresponding to that emotion" refers to a technology that analyzes a user's emotions from their biological responses and behavior and provides information that is optimal for that state.
[0345] "A method for prioritizing and displaying multiple pieces of information by combining eye movements and emotions" is a technology that evaluates the importance of information and prioritizes its display by simultaneously analyzing the user's eye movements and emotional data.
[0346] The system that realizes this application example uses a special goggle that incorporates means to track the direction the user is looking and means to detect the user's emotional state. The goggle is equipped with an eye-tracking sensor and a biosensor, which collect data in real time.
[0347] The server is responsible for analyzing gaze data and biometric data transmitted from the goggles. It uses gaze data to identify the object of gaze and an emotion engine to estimate the user's emotional state based on biometric signals. Because the server is equipped with advanced data analysis algorithms, it can accurately assess the user's emotional state, such as stress levels and concentration.
[0348] The device searches a vast database for relevant information based on analysis results sent from the server and presents content tailored to the user's emotions. For example, if a user is looking at something they are interested in, product information and discount campaign information related to that item will be displayed immediately.
[0349] As a concrete example, imagine a user wearing goggles in a physical store and browsing the electronics section. The server detects from eye-tracking data and biometric information that the user is focusing on a specific product and has a deep interest in it. The device then displays promotions and further product information related to this product on the goggles' display, supporting the user's purchasing decision. Through this entire process, the user can enjoy a more intuitive and personalized shopping experience.
[0350] An example of a prompt message is, "Identify where the user is focusing their attention on their smartphone, analyze their feelings towards that product, and generate potential promotions to display." Based on this prompt message, a generation AI model is used to provide optimal information suggestions tailored to the user's focus and emotional state.
[0351] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0352] Step 1:
[0353] The user puts on a special pair of goggles. The eye-tracking sensor and biosensor built into the goggles are activated, collecting the user's eye-tracking data and biometric information in real time. This data is transmitted to a server via wireless communication.
[0354] Step 2:
[0355] The server analyzes the received gaze data to identify the object the user is fixated on. Specifically, it calculates the gaze vector and evaluates its intensity and duration to identify a particular object or area. As a result of the analysis, it outputs the ID and location information of the object being fixed on.
[0356] Step 3:
[0357] The server evaluates the user's emotional state based on biometric data. By analyzing biosignals such as heart rate, skin temperature, and electrical activity, it estimates whether the user is relaxed or focused. The emotion engine performs this process and outputs the evaluation result as the user's emotional state.
[0358] Step 4:
[0359] The device searches a database for relevant information based on the ID and location information of the object being watched, obtained from the server, and the results of the emotional state evaluation. At this time, a generative AI model is used to select high-priority information and information most appropriate to the emotion. Based on the search results, display candidates are generated and formatted into a list of information to be presented.
[0360] Step 5:
[0361] The device displays information based on a formatted list, optimized for the user's goggles display. Prioritized information is highlighted for greater importance. The user is also provided with options to directly interact with the displayed information.
[0362] Step 6:
[0363] Users provide feedback on the information and suggestions presented. This feedback is sent from the device to the server and stored in a database to improve accuracy for future use. For example, the user's actions and reactions to each piece of information are recorded as logs.
[0364] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0365] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0366] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0367] [Third Embodiment]
[0368] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0369] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0370] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0371] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0372] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0373] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0374] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0375] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0376] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0377] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0378] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0379] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0380] To implement this invention, an infrastructure is provided that allows for eye tracking when the user wears a specific goggle device. The goggles are equipped with sensors to detect gaze in real time and collect the data. The server has the function of using this gaze data to identify the object the user is fixated on and immediately retrieving relevant information from a database.
[0381] The terminal processes relevant information obtained from the server and generates action suggestions tailored to the user's situation. For example, if the user is working, it analyzes the progress of their work based on their gaze and recommends high-priority tasks. If the user is searching for something, it uses their gaze to display lists or areas that have not yet been checked.
[0382] In this system, information is displayed on the goggles' screen according to the user's gaze direction, allowing users to easily manage information using only their eye movements. This enables users to acquire information and optimize their actions without using their hands. Furthermore, users can provide feedback on the information and suggestions they receive, and this feedback is collected on a server to improve the accuracy of future suggestions.
[0383] As a concrete example, consider a scenario where a user uses this system while working at a desk. The server detects when the user's gaze is fixed on a document for a certain period of time and searches for and provides the latest reports and materials related to that document. The terminal analyzes these materials and suggests the next steps the user should complete. This allows the user to work more efficiently.
[0384] Thus, the present invention can significantly improve the efficiency of user information acquisition and actions through a user interface that utilizes eye gaze.
[0385] The following describes the processing flow.
[0386] Step 1:
[0387] The server acquires user gaze data from the eye-tracking sensors built into the goggles. This data includes the direction, movement, and duration of gaze fixation.
[0388] Step 2:
[0389] The server analyzes the acquired gaze data to identify the area the user is fixating on. It estimates the object of gaze based on the duration of gaze fixation and the degree of pupil convergence.
[0390] Step 3:
[0391] The device analyzes images of objects based on the identified gaze area. Using image recognition technology, it detects textual information and specific objects, and identifies the objects from a database.
[0392] Step 4:
[0393] The server searches the database for information related to the identified object and retrieves that information. The retrieved information is then organized according to its content and relevance.
[0394] Step 5:
[0395] The device generates action suggestions tailored to the user's situation based on the acquired relevant information. If the user is working, it will generate suggestions for priority tasks; if the user is wandering or exploring, it will generate suggestions indicating unexplored areas.
[0396] Step 6:
[0397] The user reviews relevant information and action suggestions displayed on the goggles' screen. Based on the suggestions, they select an action and continue working.
[0398] Step 7:
[0399] Users provide feedback to the server regarding the presented information and suggestions. This feedback is used to improve the accuracy of the next suggestion generation process.
[0400] (Example 1)
[0401] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0402] There is a growing need for eye-tracking interfaces to enable users to efficiently acquire information and optimize their actions. However, existing technologies are inadequate in their ability to effectively utilize eye-tracking data for information presentation and in their use of feedback to improve the accuracy of suggestions. The challenge lies in resolving this issue and improving the user experience.
[0403] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0404] In this invention, the server includes means for providing a user interface equipped with a gaze detection device, means for identifying an object the user is fixated on and retrieving related information from a storage device, means for processing information acquired by the terminal and generating suggestions to optimize the user's actions, and means for recording the user's response to the generated suggestions and improving the accuracy of the generated products. This enables the user to acquire information and optimize their actions based on their gaze.
[0405] A "gaze detection device" is a device that senses the direction and movement of a user's gaze in real time and collects that information.
[0406] A "user interface" is a device or software designed to allow users to interact with an information system intuitively.
[0407] "Means of identifying the target" refers to the process of determining and identifying the object or area that the user is focusing on, based on gaze data.
[0408] "Means for retrieving information from a storage device" refers to methods for quickly obtaining information related to a specified object from a database or similar source.
[0409] "Means of generating suggestions" refers to the process of presenting the user with the most suitable actions or solutions based on the information acquired.
[0410] "Means for recording responses" refers to a function that saves the user's feedback on the suggestions presented, in order to improve the accuracy of the system's suggestions in the future.
[0411] To implement this invention, the user must first wear a device that detects their gaze. This device is equipped with sensors to collect the user's gaze information in real time. This data is transmitted to a server via a network.
[0412] The server receives and analyzes the gaze data, using advanced data processing software. From the gaze data, it identifies the object the user is fixated on and searches for related information within the storage device. Database query techniques and search algorithms are utilized in this process.
[0413] The device receives information sent from the server and generates action suggestions that are best suited to the user's current situation. This is done using a generative AI model, which takes eye-tracking data and related information as prompt text to generate specific suggestions.
[0414] For example, if a user fixates their gaze on a specific document while working at their desk, the server searches for the latest research and reference materials related to that document. The terminal analyzes this information and suggests the user's next course of action, such as "read the next chapter" or "review relevant meeting materials." The information is displayed on the goggles' screen according to the user's gaze, and user feedback can also be collected.
[0415] As an example of a prompt, you can input a message into the AI model in the format of, "Based on the user's eye-tracking data, generate optimal action suggestions to maximize work efficiency."
[0416] Thus, this invention enables efficient information provision and behavioral optimization by utilizing the user's gaze.
[0417] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0418] Step 1:
[0419] The user wears a device that detects their gaze. This is a preparatory step to enable the collection of gaze data. The gaze sensor detects the user's gaze direction in real time and generates data. The input is the direction of the user's gaze, and the output is the coordinates and movement data of the gaze.
[0420] Step 2:
[0421] The server receives gaze data. Here, the received gaze data is analyzed to identify the object the user is fixated on. In this process, the gaze data is converted into digital coordinates, and the database is queried to find matching information. The input is gaze coordinate data, and the output is the name of the identified object and related information.
[0422] Step 3:
[0423] The server searches the database for relevant information based on the identified target. The server efficiently accesses the database and quickly retrieves the relevant information. The data processing performed here is carried out using an information retrieval algorithm. The input is the name of the identified target, and the output is information related to that target.
[0424] Step 4:
[0425] The terminal receives information from the server and generates action suggestions tailored to the user's current situation. Using a generation AI model, it takes user eye-tracking data and related information as prompts and generates optimal suggestions based on those prompts. The input is related information and eye-tracking data, and the output is action suggestions.
[0426] Step 5:
[0427] The terminal displays generated action suggestions on the user's goggles display. The user uses their gaze to confirm the presented information and provide feedback. The input is the action suggestion, and the output is the information displayed to the user. In addition, the user's feedback is sent to the server as new data input.
[0428] (Application Example 1)
[0429] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0430] In modern factory environments, rapid and accurate information provision is essential to maintaining and improving worker efficiency and quality. However, traditional methods require manual work to access work instructions and documents, interrupting workers' work. Furthermore, as the amount of information increases, it becomes difficult for workers to quickly find the information they need. In such situations, work efficiency decreases, potentially negatively impacting productivity. In addition, there is a lack of systems for collecting feedback on given instructions and suggestions, which significantly limits the improvement of the accuracy of suggestions.
[0431] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0432] In this invention, the server includes means for tracking the direction in which the user is looking, means for identifying the object being looked at and retrieving related information, and means for suggesting an optimization of actions based on the identified information. This makes it possible to instantly obtain necessary information using eye movements in the work environment and optimize actions without using hands.
[0433] "Means for tracking the direction in which a user is looking" refers to a device or system that detects the direction of a user's gaze in real time and collects that information as digital data.
[0434] "Means for identifying objects under gaze and retrieving related information" refers to a device or system that uses sensors or other means to identify objects that fall within a user's line of sight and retrieves information related to those objects from a database or similar source.
[0435] A "means for proposing optimized behavior" refers to a device or system that has the function of generating and presenting the most effective action plan to the user based on acquired information and circumstances.
[0436] "Means for collecting feedback on provided proposals and improving the accuracy of subsequent proposals" refers to a device or system that has the function of collecting responses and opinions from users to previously received proposals and analyzing them to improve the accuracy of future proposals.
[0437] "Means for displaying procedural information and instructions related to an object of gaze in a work environment" refers to a device or system that has the function of displaying procedures and instructions related to an object or task identified by the user's gaze within the user's field of view in a workplace such as a factory.
[0438] "Means for instantly searching a data repository and presenting it in order of importance" refers to a device or system that has the function of quickly searching a database containing information related to an object being viewed and visually presenting that information to the user based on its importance.
[0439] The system for implementing this invention consists of smart glasses worn by the user and a server. The smart glasses are equipped with eye-tracking sensors that recognize the direction in which the user is looking in real time. This eye-tracking data is transmitted to the server via a network. The server analyzes the received eye-tracking data to identify the object being looked at. This analysis uses indicators such as gaze fixation time and concentration level.
[0440] The server searches its data repository for information related to the identified object. The information retrieval is performed quickly, and the results are ranked by importance. This ensures that the most relevant data is presented to the user preferentially. The presented information and suggestions are visually displayed on the smart glasses' screen, allowing the user to efficiently obtain information hands-free.
[0441] Furthermore, the server collects feedback from users on the information and suggestions presented. This feedback is used to improve the accuracy of future suggestions. The feedback includes information such as which information was helpful and which suggestions were appropriate.
[0442] As a concrete example, consider a scenario where a factory worker uses this system. When the worker looks at a new control panel, the smart glasses automatically display the assembly procedure. After completing the task, the worker can provide feedback on whether the displayed procedure was helpful, and this information is used to improve the accuracy of future procedure suggestions. In this way, eye-tracking interfaces significantly improve efficiency and accuracy in factory operations.
[0443] Example of a prompt:
[0444] "When workers are looking at new parts, please display the relevant manual on their glasses."
[0445] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0446] Step 1:
[0447] Sensors built into the smart glasses track the user's gaze and collect eye-tracking data. The direction of gaze, duration of fixed gaze, and level of concentration are measured as input. This data is converted into a digital format and sent to a server.
[0448] Step 2:
[0449] The server analyzes the received gaze data to identify the object being looked at. It receives gaze data as input, compares it with object information registered in the database, and outputs the specific object the gaze was directed at. This analysis uses a specific algorithm to process the data.
[0450] Step 3:
[0451] The server searches the data repository for information related to the identified object. It takes the object name as input and retrieves relevant information via a query. The search results are ranked based on importance, and the most relevant information is provided as output. This search and ranking process utilizes database queries and importance ranking algorithms.
[0452] Step 4:
[0453] Information transmitted from the server is displayed on the user's smart glasses. This displayed information includes relevant procedural data and important instructions. The user can intuitively confirm the information obtained through their gaze. Based on the relevant information received from the server as input, visual data is output to the smart glasses' display.
[0454] Step 5:
[0455] Users provide feedback on the displayed suggestions. Feedback is entered through the smart glasses interface, and the server collects it to improve the accuracy of future suggestions. Natural language processing technology is used to analyze the feedback, inputting the feedback data and performing specific data processing and calculations to identify areas for improvement.
[0456] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0457] This invention is a system that combines tracking the user's gaze direction and identifying the object of their gaze with an emotion engine that enables emotion recognition. The user wears a special pair of goggles, which contain an eye-tracking sensor and a biosensor that functions as the emotion engine.
[0458] The server acquires eye-tracking data and biometric data, and analyzes the user's gaze direction and emotional state. Eye-tracking analysis identifies the object of gaze, and the emotion engine uses biometric signals to estimate the user's emotional state, such as stress levels and concentration levels.
[0459] The device searches a database for information corresponding to the identified object of attention and the user's emotional state. The emotion engine then uses this information to generate information and suggestions best suited to the user's emotions. For example, if the user is stressed, it prioritizes displaying information that promotes relaxation; if the user is focused, it prioritizes displaying suggestions that enhance efficiency.
[0460] This system provides users with information and action suggestions through the goggles' display. The device collects user feedback and sends this data to a server to improve the accuracy of the suggestions. As a specific example, consider a scenario where a user uses the system during a meeting. The server detects that the user is intently looking at meeting materials and prioritizes displaying important information related to those materials. Furthermore, if the emotion engine detects a high level of concentration, it provides additional information and action plans to support further productivity improvements.
[0461] This technology allows users to instantly obtain necessary information using their gaze and emotions, enabling them to intuitively and efficiently optimize their actions.
[0462] The following describes the processing flow.
[0463] Step 1:
[0464] The server acquires the user's gaze data and biometric data in real time from the eye-tracking sensor and biosensor built into the goggles. This allows the server to observe the user's gaze direction and physiological responses.
[0465] Step 2:
[0466] The server identifies the direction and object of gaze from the eye-tracking data. This identification is achieved by analyzing the duration of gaze fixation and pupil movement.
[0467] Step 3:
[0468] The server analyzes the user's emotional state based on biometric data. The emotion engine analyzes biometric signals such as heart rate and skin temperature to estimate whether the user is excited, relaxed, focused, or otherwise in a particular state.
[0469] Step 4:
[0470] The device searches a database for information related to the gaze target identified by the eye-tracking data. During this process, high-priority information is selected and retrieved immediately.
[0471] Step 5:
[0472] The server creates appropriate action suggestions for the user based on the emotional state and related information analyzed by the emotion engine. For example, if the server determines that the user is stressed, it will suggest relaxation methods or encourage deep breathing.
[0473] Step 6:
[0474] Users view relevant information and action suggestions through the goggles' display. Information display based on eye gaze and emotion reduces user burden and allows for intuitive operation.
[0475] Step 7:
[0476] Users send feedback to the server regarding the information and suggestions provided. This feedback is recorded and analyzed by the server to improve accuracy in future instances.
[0477] (Example 2)
[0478] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0479] In modern society, users are required to quickly and accurately retrieve necessary information from a vast amount of data. However, conventional information retrieval methods have the problem of not being able to provide dynamic and optimal information that responds to changes in the user's gaze and emotions. To solve this problem, a system is needed that analyzes the user's gaze direction and emotional state in real time and provides information based on that analysis immediately.
[0480] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0481] In this invention, the server includes means for tracking the user's gaze and analyzing its direction, means for estimating the emotional state from biometric information, and means for retrieving information from a database based on the estimated gaze target and emotional state. This makes it possible to analyze the user's current gaze direction and emotional state in real time and quickly provide optimal information accordingly.
[0482] "Eye-tracking" is a technology that detects the movement of a user's eyeballs and understands the direction of their gaze in real time.
[0483] "Biometric information" refers to physiological data such as the user's heart rate and skin electrical responses, and is used to infer emotional states and other related information.
[0484] "Estimating emotional state" is a method of analyzing a user's emotional state from biometric information and evaluating factors such as stress levels and concentration levels.
[0485] "Searching for information in a database" is the process of retrieving relevant data from a collection of aggregated information based on specific criteria.
[0486] "Generating suggestions" is the process of creating optimal action instructions and information presentations for a user based on their analyzed state.
[0487] "Collecting feedback" is the activity of gathering user reactions and behaviors as data and analyzing it to improve the accuracy of future information provision.
[0488] This system starts operating when the user puts on special goggles. The goggles have built-in eye-tracking sensors and biosensors, which are used to acquire the user's gaze data and biometric information in real time. The eye-tracking sensors sense the movement of the user's eyeballs and provide technology to identify the direction of their gaze. The biosensors collect physiological data such as heart rate and skin electrical activity, and the server uses a generated AI model to estimate the user's emotional state.
[0489] The server analyzes the received data and evaluates what the user is looking at and their emotional state (e.g., stress level, concentration level). The analysis results are sent to the terminal, which then searches the database for relevant information based on the analyzed gaze and emotional state. This database search is performed quickly and efficiently, providing information that best suits the user's needs.
[0490] In this system, the generative AI model plays a crucial role in learning from the user's past behavior and feedback to generate personalized suggestions for each individual user. These suggestions and information are presented through the goggles' display.
[0491] As a concrete example, let's assume a user is attending a public lecture. In this scenario, the server tracks the user's gaze and detects when they are focusing on a particular slide. The sentiment engine estimates that the user is highly interested in the lecture's content, and the device provides relevant materials and additional background information. This allows the user to understand the lecture more deeply.
[0492] An example of a prompt is, "Design a system that generates and presents personalized information in real time based on user emotions and eye-tracking data." This example demonstrates how a system can improve the user experience.
[0493] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0494] Step 1:
[0495] The user puts on a special goggle and begins data collection. An eye-tracking sensor built into the goggles detects the user's eye movements in real time and acquires the direction as digital data. In addition, biosensors collect physiological data such as the user's heart rate and skin electrical activity. The input is the user's gaze direction and biometric information, and the output is provided as a data stream in its original form.
[0496] Step 2:
[0497] The device transmits gaze data and biometric information acquired from the goggles to the server via wireless communication (e.g., Bluetooth or Wi-Fi). The input data consists of gaze position data and biometric signal data, and the output is the result of transmitting that data to the server.
[0498] Step 3:
[0499] The server analyzes the received gaze data to determine the user's gaze direction. In this process, the server uses an algorithm to identify the specific object the user is fixated on based on the gaze data. The input is gaze data, and the output is the identified object of fixation.
[0500] Step 4:
[0501] The server simultaneously analyzes biometric information and uses a generative AI model to estimate the user's emotional state. This involves performing the necessary data calculations to determine emotional states such as stress levels and concentration levels. The input is biometric information, and the output is an evaluation of the user's emotional state.
[0502] Step 5:
[0503] The terminal searches the database for relevant information based on the object of attention and emotional state received from the server. At this stage, the terminal selects high-priority information and prepares it to be provided to the user. The input is the identified object of attention and emotional state, and the output is the relevant information retrieved from the database.
[0504] Step 6:
[0505] The device provides the user with generated suggestions and information through the goggles' display. The user can visually confirm this. The input is the searched information and suggestions, and the output appears as the displayed content.
[0506] Step 7:
[0507] The device records the user's reactions to the information received through the goggles and sends them to the server as feedback. This feedback is used to improve future suggestions. The input is the user's reactions and feedback, and the output is the feedback data sent to the server.
[0508] (Application Example 2)
[0509] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0510] The challenge is to provide a system that supports efficient and accurate purchasing decisions by eliminating the confusion caused by information overload and numerous choices when users select products in physical stores. Furthermore, it is necessary to improve the shopping experience by providing personalized information tailored to the user's interests and emotional state.
[0511] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0512] In this invention, the server includes means for tracking the direction in which the user is looking, means for identifying the object being looked at and retrieving related information, means for suggesting optimized actions based on the identified information, means for collecting feedback on the provided suggestions and improving the accuracy of subsequent suggestions, means for detecting the user's emotional state and displaying information corresponding to that emotion, and means for prioritizing and displaying multiple pieces of information by combining gaze and emotion. This enables intuitive and effective information provision based on the user's gaze and emotions, significantly improving the user's purchasing experience.
[0513] "Means of tracking the direction a user is looking" refers to devices or technologies that detect the direction a user is looking and acquire it as digital data.
[0514] "Means for identifying objects being watched and retrieving related information" refers to technology that identifies objects or areas in the user's line of sight and extracts and provides information about them from a digital database.
[0515] "Means of suggesting optimized actions based on identified information" refers to technologies that present recommendations and action plans to support user decision-making based on acquired information.
[0516] "Means of collecting feedback on provided proposals and improving the accuracy of future proposals" refers to technologies that collect user reactions and opinions, analyze them, and increase the success rate of future proposals.
[0517] "A means of detecting a user's emotional state and displaying information corresponding to that emotion" refers to a technology that analyzes a user's emotions from their biological responses and behavior and provides information that is optimal for that state.
[0518] "A method for prioritizing and displaying multiple pieces of information by combining eye movements and emotions" is a technology that evaluates the importance of information and prioritizes its display by simultaneously analyzing the user's eye movements and emotional data.
[0519] The system that realizes this application example uses a special goggle that incorporates means to track the direction the user is looking and means to detect the user's emotional state. The goggle is equipped with an eye-tracking sensor and a biosensor, which collect data in real time.
[0520] The server is responsible for analyzing gaze data and biometric data transmitted from the goggles. It uses gaze data to identify the object of gaze and an emotion engine to estimate the user's emotional state based on biometric signals. Because the server is equipped with advanced data analysis algorithms, it can accurately assess the user's emotional state, such as stress levels and concentration.
[0521] The device searches a vast database for relevant information based on analysis results sent from the server and presents content tailored to the user's emotions. For example, if a user is looking at something they are interested in, product information and discount campaign information related to that item will be displayed immediately.
[0522] As a concrete example, imagine a user wearing goggles in a physical store and browsing the electronics section. The server detects from eye-tracking data and biometric information that the user is focusing on a specific product and has a deep interest in it. The device then displays promotions and further product information related to this product on the goggles' display, supporting the user's purchasing decision. Through this entire process, the user can enjoy a more intuitive and personalized shopping experience.
[0523] An example of a prompt message is, "Identify where the user is focusing their attention on their smartphone, analyze their feelings towards that product, and generate potential promotions to display." Based on this prompt message, a generation AI model is used to provide optimal information suggestions tailored to the user's focus and emotional state.
[0524] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0525] Step 1:
[0526] The user puts on a special pair of goggles. The eye-tracking sensor and biosensor built into the goggles are activated, collecting the user's eye-tracking data and biometric information in real time. This data is transmitted to a server via wireless communication.
[0527] Step 2:
[0528] The server analyzes the received gaze data to identify the object the user is fixated on. Specifically, it calculates the gaze vector and evaluates its intensity and duration to identify a particular object or area. As a result of the analysis, it outputs the ID and location information of the object being fixed on.
[0529] Step 3:
[0530] The server evaluates the user's emotional state based on biometric data. By analyzing biosignals such as heart rate, skin temperature, and electrical activity, it estimates whether the user is relaxed or focused. The emotion engine performs this process and outputs the evaluation result as the user's emotional state.
[0531] Step 4:
[0532] The device searches a database for relevant information based on the ID and location information of the object being watched, obtained from the server, and the results of the emotional state evaluation. At this time, a generative AI model is used to select high-priority information and information most appropriate to the emotion. Based on the search results, display candidates are generated and formatted into a list of information to be presented.
[0533] Step 5:
[0534] The device displays information based on a formatted list, optimized for the user's goggles display. Prioritized information is highlighted for greater importance. The user is also provided with options to directly interact with the displayed information.
[0535] Step 6:
[0536] Users provide feedback on the information and suggestions presented. This feedback is sent from the device to the server and stored in a database to improve accuracy for future use. For example, the user's actions and reactions to each piece of information are recorded as logs.
[0537] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0538] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0539] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[0540] [Fourth Embodiment]
[0541] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0542] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0543] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0544] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0545] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0546] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0547] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0548] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0549] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0550] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0551] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0552] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0553] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0554] To implement this invention, an infrastructure is provided that allows for eye tracking when the user wears a specific goggle device. The goggles are equipped with sensors to detect gaze in real time and collect the data. The server has the function of using this gaze data to identify the object the user is fixated on and immediately retrieving relevant information from a database.
[0555] The terminal processes relevant information obtained from the server and generates action suggestions tailored to the user's situation. For example, if the user is working, it analyzes the progress of their work based on their gaze and recommends high-priority tasks. If the user is searching for something, it uses their gaze to display lists or areas that have not yet been checked.
[0556] In this system, information is displayed on the goggles' screen according to the user's gaze direction, allowing users to easily manage information using only their eye movements. This enables users to acquire information and optimize their actions without using their hands. Furthermore, users can provide feedback on the information and suggestions they receive, and this feedback is collected on a server to improve the accuracy of future suggestions.
[0557] As a concrete example, consider a scenario where a user uses this system while working at a desk. The server detects when the user's gaze is fixed on a document for a certain period of time and searches for and provides the latest reports and materials related to that document. The terminal analyzes these materials and suggests the next steps the user should complete. This allows the user to work more efficiently.
[0558] Thus, the present invention can significantly improve the efficiency of user information acquisition and actions through a user interface that utilizes eye gaze.
[0559] The following describes the processing flow.
[0560] Step 1:
[0561] The server acquires user gaze data from the eye-tracking sensors built into the goggles. This data includes the direction, movement, and duration of gaze fixation.
[0562] Step 2:
[0563] The server analyzes the acquired gaze data to identify the area the user is fixating on. It estimates the object of gaze based on the duration of gaze fixation and the degree of pupil convergence.
[0564] Step 3:
[0565] The device analyzes images of objects based on the identified gaze area. Using image recognition technology, it detects textual information and specific objects, and identifies the objects from a database.
[0566] Step 4:
[0567] The server searches the database for information related to the identified object and retrieves that information. The retrieved information is then organized according to its content and relevance.
[0568] Step 5:
[0569] The device generates action suggestions tailored to the user's situation based on the acquired relevant information. If the user is working, it will generate suggestions for priority tasks; if the user is wandering or exploring, it will generate suggestions indicating unexplored areas.
[0570] Step 6:
[0571] The user reviews relevant information and action suggestions displayed on the goggles' screen. Based on the suggestions, they select an action and continue working.
[0572] Step 7:
[0573] Users provide feedback to the server regarding the presented information and suggestions. This feedback is used to improve the accuracy of the next suggestion generation process.
[0574] (Example 1)
[0575] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0576] There is a growing need for eye-tracking interfaces to enable users to efficiently acquire information and optimize their actions. However, existing technologies are inadequate in their ability to effectively utilize eye-tracking data for information presentation and in their use of feedback to improve the accuracy of suggestions. The challenge lies in resolving this issue and improving the user experience.
[0577] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0578] In this invention, the server includes means for providing a user interface equipped with a gaze detection device, means for identifying an object the user is fixated on and retrieving related information from a storage device, means for processing information acquired by the terminal and generating suggestions to optimize the user's actions, and means for recording the user's response to the generated suggestions and improving the accuracy of the generated products. This enables the user to acquire information and optimize their actions based on their gaze.
[0579] A "gaze detection device" is a device that senses the direction and movement of a user's gaze in real time and collects that information.
[0580] A "user interface" is a device or software designed to allow users to interact with an information system intuitively.
[0581] "Means of identifying the target" refers to the process of determining and identifying the object or area that the user is focusing on, based on gaze data.
[0582] "Means for retrieving information from a storage device" refers to methods for quickly obtaining information related to a specified object from a database or similar source.
[0583] "Means of generating suggestions" refers to the process of presenting the user with the most suitable actions or solutions based on the information acquired.
[0584] "Means for recording responses" refers to a function that saves the user's feedback on the suggestions presented, in order to improve the accuracy of the system's suggestions in the future.
[0585] To implement this invention, the user must first wear a device that detects their gaze. This device is equipped with sensors to collect the user's gaze information in real time. This data is transmitted to a server via a network.
[0586] The server receives and analyzes the gaze data, using advanced data processing software. From the gaze data, it identifies the object the user is fixated on and searches for related information within the storage device. Database query techniques and search algorithms are utilized in this process.
[0587] The device receives information sent from the server and generates action suggestions that are best suited to the user's current situation. This is done using a generative AI model, which takes eye-tracking data and related information as prompt text to generate specific suggestions.
[0588] For example, if a user fixates their gaze on a specific document while working at their desk, the server searches for the latest research and reference materials related to that document. The terminal analyzes this information and suggests the user's next course of action, such as "read the next chapter" or "review relevant meeting materials." The information is displayed on the goggles' screen according to the user's gaze, and user feedback can also be collected.
[0589] As an example of a prompt, you can input a message into the AI model in the format of, "Based on the user's eye-tracking data, generate optimal action suggestions to maximize work efficiency."
[0590] Thus, this invention enables efficient information provision and behavioral optimization by utilizing the user's gaze.
[0591] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0592] Step 1:
[0593] The user wears a device that detects their gaze. This is a preparatory step to enable the collection of gaze data. The gaze sensor detects the user's gaze direction in real time and generates data. The input is the direction of the user's gaze, and the output is the coordinates and movement data of the gaze.
[0594] Step 2:
[0595] The server receives gaze data. Here, the received gaze data is analyzed to identify the object the user is fixated on. In this process, the gaze data is converted into digital coordinates, and the database is queried to find matching information. The input is gaze coordinate data, and the output is the name of the identified object and related information.
[0596] Step 3:
[0597] The server searches the database for relevant information based on the identified target. The server efficiently accesses the database and quickly retrieves the relevant information. The data processing performed here is carried out using an information retrieval algorithm. The input is the name of the identified target, and the output is information related to that target.
[0598] Step 4:
[0599] The terminal receives information from the server and generates action suggestions tailored to the user's current situation. Using a generation AI model, it takes user eye-tracking data and related information as prompts and generates optimal suggestions based on those prompts. The input is related information and eye-tracking data, and the output is action suggestions.
[0600] Step 5:
[0601] The terminal displays generated action suggestions on the user's goggles display. The user uses their gaze to confirm the presented information and provide feedback. The input is the action suggestion, and the output is the information displayed to the user. In addition, the user's feedback is sent to the server as new data input.
[0602] (Application Example 1)
[0603] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0604] In modern factory environments, rapid and accurate information provision is essential to maintaining and improving worker efficiency and quality. However, traditional methods require manual work to access work instructions and documents, interrupting workers' work. Furthermore, as the amount of information increases, it becomes difficult for workers to quickly find the information they need. In such situations, work efficiency decreases, potentially negatively impacting productivity. In addition, there is a lack of systems for collecting feedback on given instructions and suggestions, which significantly limits the improvement of the accuracy of suggestions.
[0605] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0606] In this invention, the server includes means for tracking the direction in which the user is looking, means for identifying the object being looked at and retrieving related information, and means for suggesting an optimization of actions based on the identified information. This makes it possible to instantly obtain necessary information using eye movements in the work environment and optimize actions without using hands.
[0607] "Means for tracking the direction in which a user is looking" refers to a device or system that detects the direction of a user's gaze in real time and collects that information as digital data.
[0608] "Means for identifying objects under gaze and retrieving related information" refers to a device or system that uses sensors or other means to identify objects that fall within a user's line of sight and retrieves information related to those objects from a database or similar source.
[0609] A "means for proposing optimized behavior" refers to a device or system that has the function of generating and presenting the most effective action plan to the user based on acquired information and circumstances.
[0610] "Means for collecting feedback on provided proposals and improving the accuracy of subsequent proposals" refers to a device or system that has the function of collecting responses and opinions from users to previously received proposals and analyzing them to improve the accuracy of future proposals.
[0611] "Means for displaying procedural information and instructions related to an object of gaze in a work environment" refers to a device or system that has the function of displaying procedures and instructions related to an object or task identified by the user's gaze within the user's field of view in a workplace such as a factory.
[0612] "Means for instantly searching a data repository and presenting it in order of importance" refers to a device or system that has the function of quickly searching a database containing information related to an object being viewed and visually presenting that information to the user based on its importance.
[0613] The system for implementing this invention consists of smart glasses worn by the user and a server. The smart glasses are equipped with eye-tracking sensors that recognize the direction in which the user is looking in real time. This eye-tracking data is transmitted to the server via a network. The server analyzes the received eye-tracking data to identify the object being looked at. This analysis uses indicators such as gaze fixation time and concentration level.
[0614] The server searches its data repository for information related to the identified object. The information retrieval is performed quickly, and the results are ranked by importance. This ensures that the most relevant data is presented to the user preferentially. The presented information and suggestions are visually displayed on the smart glasses' screen, allowing the user to efficiently obtain information hands-free.
[0615] Furthermore, the server collects feedback from users on the information and suggestions presented. This feedback is used to improve the accuracy of future suggestions. The feedback includes information such as which information was helpful and which suggestions were appropriate.
[0616] As a concrete example, consider a scenario where a factory worker uses this system. When the worker looks at a new control panel, the smart glasses automatically display the assembly procedure. After completing the task, the worker can provide feedback on whether the displayed procedure was helpful, and this information is used to improve the accuracy of future procedure suggestions. In this way, eye-tracking interfaces significantly improve efficiency and accuracy in factory operations.
[0617] Example of a prompt:
[0618] "When workers are looking at new parts, please display the relevant manual on their glasses."
[0619] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0620] Step 1:
[0621] Sensors built into the smart glasses track the user's gaze and collect eye-tracking data. The direction of gaze, duration of fixed gaze, and level of concentration are measured as input. This data is converted into a digital format and sent to a server.
[0622] Step 2:
[0623] The server analyzes the received gaze data to identify the object being looked at. It receives gaze data as input, compares it with object information registered in the database, and outputs the specific object the gaze was directed at. This analysis uses a specific algorithm to process the data.
[0624] Step 3:
[0625] The server searches the data repository for information related to the identified object. It takes the object name as input and retrieves relevant information via a query. The search results are ranked based on importance, and the most relevant information is provided as output. This search and ranking process utilizes database queries and importance ranking algorithms.
[0626] Step 4:
[0627] Information transmitted from the server is displayed on the user's smart glasses. This displayed information includes relevant procedural data and important instructions. The user can intuitively confirm the information obtained through their gaze. Based on the relevant information received from the server as input, visual data is output to the smart glasses' display.
[0628] Step 5:
[0629] Users provide feedback on the displayed suggestions. Feedback is entered through the smart glasses interface, and the server collects it to improve the accuracy of future suggestions. Natural language processing technology is used to analyze the feedback, inputting the feedback data and performing specific data processing and calculations to identify areas for improvement.
[0630] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0631] This invention is a system that combines tracking the user's gaze direction and identifying the object of their gaze with an emotion engine that enables emotion recognition. The user wears a special pair of goggles, which contain an eye-tracking sensor and a biosensor that functions as the emotion engine.
[0632] The server acquires eye-tracking data and biometric data, and analyzes the user's gaze direction and emotional state. Eye-tracking analysis identifies the object of gaze, and the emotion engine uses biometric signals to estimate the user's emotional state, such as stress levels and concentration levels.
[0633] The device searches a database for information corresponding to the identified object of attention and the user's emotional state. The emotion engine then uses this information to generate information and suggestions best suited to the user's emotions. For example, if the user is stressed, it prioritizes displaying information that promotes relaxation; if the user is focused, it prioritizes displaying suggestions that enhance efficiency.
[0634] This system provides users with information and action suggestions through the goggles' display. The device collects user feedback and sends this data to a server to improve the accuracy of the suggestions. As a specific example, consider a scenario where a user uses the system during a meeting. The server detects that the user is intently looking at meeting materials and prioritizes displaying important information related to those materials. Furthermore, if the emotion engine detects a high level of concentration, it provides additional information and action plans to support further productivity improvements.
[0635] This technology allows users to instantly obtain necessary information using their gaze and emotions, enabling them to intuitively and efficiently optimize their actions.
[0636] The following describes the processing flow.
[0637] Step 1:
[0638] The server acquires the user's gaze data and biometric data in real time from the eye-tracking sensor and biosensor built into the goggles. This allows the server to observe the user's gaze direction and physiological responses.
[0639] Step 2:
[0640] The server identifies the direction and object of gaze from the eye-tracking data. This identification is achieved by analyzing the duration of gaze fixation and pupil movement.
[0641] Step 3:
[0642] The server analyzes the user's emotional state based on biometric data. The emotion engine analyzes biometric signals such as heart rate and skin temperature to estimate whether the user is excited, relaxed, focused, or otherwise in a particular state.
[0643] Step 4:
[0644] The device searches a database for information related to the gaze target identified by the eye-tracking data. During this process, high-priority information is selected and retrieved immediately.
[0645] Step 5:
[0646] The server creates appropriate action suggestions for the user based on the emotional state and related information analyzed by the emotion engine. For example, if the server determines that the user is stressed, it will suggest relaxation methods or encourage deep breathing.
[0647] Step 6:
[0648] Users view relevant information and action suggestions through the goggles' display. Information display based on eye gaze and emotion reduces user burden and allows for intuitive operation.
[0649] Step 7:
[0650] Users send feedback to the server regarding the information and suggestions provided. This feedback is recorded and analyzed by the server to improve accuracy in future instances.
[0651] (Example 2)
[0652] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0653] In modern society, users are required to quickly and accurately retrieve necessary information from a vast amount of data. However, conventional information retrieval methods have the problem of not being able to provide dynamic and optimal information that responds to changes in the user's gaze and emotions. To solve this problem, a system is needed that analyzes the user's gaze direction and emotional state in real time and provides information based on that analysis immediately.
[0654] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0655] In this invention, the server includes means for tracking the user's gaze and analyzing its direction, means for estimating the emotional state from biometric information, and means for retrieving information from a database based on the estimated gaze target and emotional state. This makes it possible to analyze the user's current gaze direction and emotional state in real time and quickly provide optimal information accordingly.
[0656] "Eye-tracking" is a technology that detects the movement of a user's eyeballs and understands the direction of their gaze in real time.
[0657] "Biometric information" refers to physiological data such as the user's heart rate and skin electrical responses, and is used to infer emotional states and other related information.
[0658] "Estimating emotional state" is a method of analyzing a user's emotional state from biometric information and evaluating factors such as stress levels and concentration levels.
[0659] "Searching for information in a database" is the process of retrieving relevant data from a collection of aggregated information based on specific criteria.
[0660] "Generating suggestions" is the process of creating optimal action instructions and information presentations for a user based on their analyzed state.
[0661] "Collecting feedback" is the activity of gathering user reactions and behaviors as data and analyzing it to improve the accuracy of future information provision.
[0662] This system starts operating when the user puts on special goggles. The goggles have built-in eye-tracking sensors and biosensors, which are used to acquire the user's gaze data and biometric information in real time. The eye-tracking sensors sense the movement of the user's eyeballs and provide technology to identify the direction of their gaze. The biosensors collect physiological data such as heart rate and skin electrical activity, and the server uses a generated AI model to estimate the user's emotional state.
[0663] The server analyzes the received data and evaluates what the user is looking at and their emotional state (e.g., stress level, concentration level). The analysis results are sent to the terminal, which then searches the database for relevant information based on the analyzed gaze and emotional state. This database search is performed quickly and efficiently, providing information that best suits the user's needs.
[0664] In this system, the generative AI model plays a crucial role in learning from the user's past behavior and feedback to generate personalized suggestions for each individual user. These suggestions and information are presented through the goggles' display.
[0665] As a concrete example, let's assume a user is attending a public lecture. In this scenario, the server tracks the user's gaze and detects when they are focusing on a particular slide. The sentiment engine estimates that the user is highly interested in the lecture's content, and the device provides relevant materials and additional background information. This allows the user to understand the lecture more deeply.
[0666] An example of a prompt is, "Design a system that generates and presents personalized information in real time based on user emotions and eye-tracking data." This example demonstrates how a system can improve the user experience.
[0667] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0668] Step 1:
[0669] The user puts on a special goggle and begins data collection. An eye-tracking sensor built into the goggles detects the user's eye movements in real time and acquires the direction as digital data. In addition, biosensors collect physiological data such as the user's heart rate and skin electrical activity. The input is the user's gaze direction and biometric information, and the output is provided as a data stream in its original form.
[0670] Step 2:
[0671] The device transmits gaze data and biometric information acquired from the goggles to the server via wireless communication (e.g., Bluetooth or Wi-Fi). The input data consists of gaze position data and biometric signal data, and the output is the result of transmitting that data to the server.
[0672] Step 3:
[0673] The server analyzes the received gaze data to determine the user's gaze direction. In this process, the server uses an algorithm to identify the specific object the user is fixated on based on the gaze data. The input is gaze data, and the output is the identified object of fixation.
[0674] Step 4:
[0675] The server simultaneously analyzes biometric information and uses a generative AI model to estimate the user's emotional state. This involves performing the necessary data calculations to determine emotional states such as stress levels and concentration levels. The input is biometric information, and the output is an evaluation of the user's emotional state.
[0676] Step 5:
[0677] The terminal searches the database for relevant information based on the object of attention and emotional state received from the server. At this stage, the terminal selects high-priority information and prepares it to be provided to the user. The input is the identified object of attention and emotional state, and the output is the relevant information retrieved from the database.
[0678] Step 6:
[0679] The device provides the user with generated suggestions and information through the goggles' display. The user can visually confirm this. The input is the searched information and suggestions, and the output appears as the displayed content.
[0680] Step 7:
[0681] The device records the user's reactions to the information received through the goggles and sends them to the server as feedback. This feedback is used to improve future suggestions. The input is the user's reactions and feedback, and the output is the feedback data sent to the server.
[0682] (Application Example 2)
[0683] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0684] The challenge is to provide a system that supports efficient and accurate purchasing decisions by eliminating the confusion caused by information overload and numerous choices when users select products in physical stores. Furthermore, it is necessary to improve the shopping experience by providing personalized information tailored to the user's interests and emotional state.
[0685] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0686] In this invention, the server includes means for tracking the direction in which the user is looking, means for identifying the object being looked at and retrieving related information, means for suggesting optimized actions based on the identified information, means for collecting feedback on the provided suggestions and improving the accuracy of subsequent suggestions, means for detecting the user's emotional state and displaying information corresponding to that emotion, and means for prioritizing and displaying multiple pieces of information by combining gaze and emotion. This enables intuitive and effective information provision based on the user's gaze and emotions, significantly improving the user's purchasing experience.
[0687] "Means of tracking the direction a user is looking" refers to devices or technologies that detect the direction a user is looking and acquire it as digital data.
[0688] "Means for identifying objects being watched and retrieving related information" refers to technology that identifies objects or areas in the user's line of sight and extracts and provides information about them from a digital database.
[0689] "Means of suggesting optimized actions based on identified information" refers to technologies that present recommendations and action plans to support user decision-making based on acquired information.
[0690] "Means of collecting feedback on provided proposals and improving the accuracy of future proposals" refers to technologies that collect user reactions and opinions, analyze them, and increase the success rate of future proposals.
[0691] "A means of detecting a user's emotional state and displaying information corresponding to that emotion" refers to a technology that analyzes a user's emotions from their biological responses and behavior and provides information that is optimal for that state.
[0692] "A method for prioritizing and displaying multiple pieces of information by combining eye movements and emotions" is a technology that evaluates the importance of information and prioritizes its display by simultaneously analyzing the user's eye movements and emotional data.
[0693] The system that realizes this application example uses a special goggle that incorporates means to track the direction the user is looking and means to detect the user's emotional state. The goggle is equipped with an eye-tracking sensor and a biosensor, which collect data in real time.
[0694] The server is responsible for analyzing gaze data and biometric data transmitted from the goggles. It uses gaze data to identify the object of gaze and an emotion engine to estimate the user's emotional state based on biometric signals. Because the server is equipped with advanced data analysis algorithms, it can accurately assess the user's emotional state, such as stress levels and concentration.
[0695] The device searches a vast database for relevant information based on analysis results sent from the server and presents content tailored to the user's emotions. For example, if a user is looking at something they are interested in, product information and discount campaign information related to that item will be displayed immediately.
[0696] As a concrete example, imagine a user wearing goggles in a physical store and browsing the electronics section. The server detects from eye-tracking data and biometric information that the user is focusing on a specific product and has a deep interest in it. The device then displays promotions and further product information related to this product on the goggles' display, supporting the user's purchasing decision. Through this entire process, the user can enjoy a more intuitive and personalized shopping experience.
[0697] An example of a prompt message is, "Identify where the user is focusing their attention on their smartphone, analyze their feelings towards that product, and generate potential promotions to display." Based on this prompt message, a generation AI model is used to provide optimal information suggestions tailored to the user's focus and emotional state.
[0698] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0699] Step 1:
[0700] The user puts on a special pair of goggles. The eye-tracking sensor and biosensor built into the goggles are activated, collecting the user's eye-tracking data and biometric information in real time. This data is transmitted to a server via wireless communication.
[0701] Step 2:
[0702] The server analyzes the received gaze data to identify the object the user is fixated on. Specifically, it calculates the gaze vector and evaluates its intensity and duration to identify a particular object or area. As a result of the analysis, it outputs the ID and location information of the object being fixed on.
[0703] Step 3:
[0704] The server evaluates the user's emotional state based on biometric data. By analyzing biosignals such as heart rate, skin temperature, and electrical activity, it estimates whether the user is relaxed or focused. The emotion engine performs this process and outputs the evaluation result as the user's emotional state.
[0705] Step 4:
[0706] The device searches a database for relevant information based on the ID and location information of the object being watched, obtained from the server, and the results of the emotional state evaluation. At this time, a generative AI model is used to select high-priority information and information most appropriate to the emotion. Based on the search results, display candidates are generated and formatted into a list of information to be presented.
[0707] Step 5:
[0708] The device displays information based on a formatted list, optimized for the user's goggles display. Prioritized information is highlighted for greater importance. The user is also provided with options to directly interact with the displayed information.
[0709] Step 6:
[0710] Users provide feedback on the information and suggestions presented. This feedback is sent from the device to the server and stored in a database to improve accuracy for future use. For example, the user's actions and reactions to each piece of information are recorded as logs.
[0711] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0712] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0713] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[0714] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0715] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[0716] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[0717] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[0718] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[0719] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[0720] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[0721] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[0722] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[0723] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0724] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0725] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[0726] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[0727] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[0728] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[0729] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0730] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0731] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.
[0732] The following is further disclosed regarding the embodiments described above.
[0733] (Claim 1)
[0734] A means of tracking the direction in which the user is looking,
[0735] A means for identifying the object being focused on and searching for related information,
[0736] A means of proposing the optimization of actions based on identified information,
[0737] A means of collecting feedback on the proposed suggestions and improving the accuracy of future suggestions,
[0738] A system that includes this.
[0739] (Claim 2)
[0740] The system according to claim 1, comprising means for identifying the object being gazed upon based on the duration of gaze fixation and the degree of concentration.
[0741] (Claim 3)
[0742] The system according to claim 1, comprising means for immediately retrieving information corresponding to the object being watched from a database and displaying it in order of priority.
[0743] "Example 1"
[0744] (Claim 1)
[0745] A means for providing a user interface equipped with a device that detects eye movements,
[0746] A means for identifying the object the user is focusing on and retrieving related information from a storage device,
[0747] A means for processing information acquired by a terminal and generating suggestions to improve the efficiency of user actions,
[0748] A means for recording user responses to generated suggestions and improving the accuracy of the generated products,
[0749] A system that includes this.
[0750] (Claim 2)
[0751] The system according to claim 1, comprising means for considering the duration and intensity of gaze in identifying the object being gazed upon.
[0752] (Claim 3)
[0753] The system according to claim 1, further comprising means for rapidly retrieving information about a subject being watched from an information storage device and presenting the information based on its importance.
[0754] "Application Example 1"
[0755] (Claim 1)
[0756] A means of tracking the direction in which the user is looking,
[0757] A means for identifying the object being focused on and retrieving related information,
[0758] A means of proposing the optimization of actions based on identified information,
[0759] A means of collecting feedback on the proposed suggestions and improving the accuracy of future suggestions,
[0760] A means of displaying procedural information and instructions related to the object of attention in the work environment,
[0761] A system that includes this.
[0762] (Claim 2)
[0763] The system according to claim 1, comprising means for identifying a gazed-on object based on the duration of gaze fixation and the degree of concentration, and further comprising means for making suggestions to improve work efficiency.
[0764] (Claim 3)
[0765] The system according to claim 1, further comprising means for instantly retrieving information corresponding to the object being focused on from a data repository and presenting it in order of importance, and further means for adjusting the next suggestion based on user feedback.
[0766] "Example 2 of combining an emotion engine"
[0767] (Claim 1)
[0768] A means of tracking the user's gaze and analyzing its direction,
[0769] A method for estimating emotional states from biometric information associated with eye movements,
[0770] A means of retrieving information from a database based on the estimated object of attention and emotional state,
[0771] A means of generating suggestions that are appropriate to the user's emotions based on the acquired information,
[0772] A means of collecting user feedback on the provided suggestions and improving the accuracy of future suggestions,
[0773] A system that includes this.
[0774] (Claim 2)
[0775] The system according to claim 1, comprising means for identifying the object being gazed upon based on the duration of gaze fixation and the degree of concentration analyzed by the emotion engine.
[0776] (Claim 3)
[0777] The system according to claim 1, further comprising means for instantly retrieving information corresponding to the object being focused on from a database and displaying it in order of priority according to the user's emotional state.
[0778] "Application example 2 when combining with an emotional engine"
[0779] (Claim 1)
[0780] A means of tracking the direction in which the user is looking,
[0781] A means for identifying the object being focused on and searching for related information,
[0782] A means of proposing the optimization of actions based on identified information,
[0783] A means of collecting feedback on the proposed suggestions and improving the accuracy of future suggestions,
[0784] A means for detecting the user's emotional state and displaying information corresponding to that emotion,
[0785] A method for prioritizing and displaying multiple pieces of information by combining gaze and emotion,
[0786] A system that includes this.
[0787] (Claim 2)
[0788] The system according to claim 1, comprising means for identifying the object being gazed upon based on the duration of gaze fixation, the degree of concentration, and the emotional state.
[0789] (Claim 3)
[0790] The system according to claim 1, comprising means for instantly retrieving information corresponding to the object being watched from a database and displaying it in order of priority based on emotion. [Explanation of symbols]
[0791] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. A means of tracking the direction in which the user is looking, A means for identifying the object being focused on and searching for related information, A means of proposing the optimization of actions based on identified information, A means of collecting feedback on the proposed suggestions and improving the accuracy of future suggestions, A system that includes this.
2. The system according to claim 1, comprising means for identifying the object being gazed upon based on the duration of gaze fixation and the degree of concentration.
3. The system according to claim 1, further comprising means for immediately retrieving information corresponding to the object being watched from a database and displaying it in order of priority.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A