System

The system addresses inefficiencies in recording and sharing researcher observations by using eye-tracking devices for real-time data acquisition, analysis, and verbalization, enhancing research and educational quality.

JP2026016162APending Publication Date: 2026-02-03SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024117252
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-22
Publication Date
2026-02-03

AI Technical Summary

Technical Problem

Conventional methods for recording and sharing researchers' observations using binoculars or microscopes are time-consuming and inefficient, limiting the quality of research and education, especially in fields requiring real-time data sharing.

Method used

A system that uses eye-tracking devices to acquire gaze data, transmit it to a server for analysis and verbalization, and store/share the data for immediate feedback to researchers.

Benefits of technology

Enables real-time recording and sharing of observations, improving research efficiency and educational quality by providing fast and accurate analysis and feedback.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026016162000001_ABST
    Figure 2026016162000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system comprising: means for acquiring viewpoint data of a researcher by an eye tracking device; means for transmitting the acquired viewpoint data to a server; means for parsing and verbalizing the viewpoint data at the server; means for storing the parsed verbalized data and the viewpoint data; and means for sharing the stored data.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] In conventional research methods, researchers must use binoculars or microscopes to quickly and accurately record their observations and share them with other researchers and educators, which requires a lot of time and effort. This not only reduces research efficiency but also limits the improvement of the quality of education. In particular, in fields where detailed observation results from specific perspectives are important, the lack of a means to share observation data in real time has been an issue. [Means for solving the problem]

[0005] The present invention provides a system that acquires researchers' gaze data using an eye-tracking device, transmits the data to a server, analyzes it, and verbalizes it. Furthermore, by providing a means for saving the analyzed data and gaze data and sharing it with relevant parties, the system aims to improve the efficiency and quality of research and education. Specifically, the system includes the following means:

[0006] 1. A method for obtaining researchers' gaze data using eye-tracking devices

[0007] 2. A method for transmitting acquired viewpoint data to the server

[0008] 3. A method for analyzing and verbalizing viewpoint data on the server

[0009] 4. Means for storing analyzed verbalization data and viewpoint data

[0010] 5. Means of sharing stored data

[0011] This will enable researchers to record their observations in real time and quickly share them with other researchers and educators, which is expected to improve research efficiency and the quality of education.

[0012] An "eye tracking device" is a device that tracks the gaze of researchers using binoculars or microscopes and collects that data.

[0013] "Gaze data" is information obtained by an eye-tracking device that indicates where the researcher is looking.

[0014] A "server" is a computer system for receiving, analyzing, and storing viewpoint data.

[0015] "Verbalization" is the process of converting observations into natural language based on viewpoint data.

[0016] "Storage" is the act of permanently recording data in a database or storage system.

[0017] "Sharing" refers to the act of providing stored data to stakeholders and educational institutions and making it accessible.

[0018] "Calibration" is the process of adjusting an eye tracking device so that it can accurately acquire gaze data.

[0019] "Real-time" refers to a state in which viewpoint data is acquired and analyzed almost immediately. [Brief explanation of the drawings]

[0020] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11]FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0021] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0022] First, the terms used in the following description will be explained.

[0023] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0024] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0025] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0026] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0027] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0028] [First embodiment]

[0029] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0030] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0031] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0032] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0033] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0034] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0035] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0036] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0037] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0038] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0039] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0040] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0041] This invention is a system that uses an AI-equipped eye-tracking device to collect viewpoint data observed by researchers through binoculars or microscopes, analyzes the data, and verbalizes it. Furthermore, by storing and sharing this data, it aims to improve the quality of research and education.

[0042] System configuration

[0043] 1. Eye tracking device:

[0044] Terminals: Eye-tracking devices are attached to binoculars and microscopes to track the researcher's gaze.

[0045] Terminal: The device has the ability to capture eye movement and retrieve that data in real time.

[0046] 2. Transmitting viewpoint data:

[0047] Terminal: The acquired viewpoint data is sent to the server periodically to ensure that the viewpoint data reaches the server promptly.

[0048] 3. Saving viewpoint data:

[0049] Server: The server temporarily stores the viewpoint data it receives. A backup function may also be implemented to ensure that data is not lost.

[0050] 4. Analysis and verbalization of viewpoint data:

[0051] Server: Uses a generative AI to analyze the viewpoint data. The AI ​​processes the viewpoint data and verbalizes the observations in natural language.

[0052] Server: The verbalized data includes details about the points the researcher was observing and what they were looking at.

[0053] 5. Displaying the analysis results:

[0054] Terminal: The analysis results sent from the server are displayed to the researcher in real time, allowing the researcher to immediately check their observations.

[0055] 6. Data Storage and Sharing:

[0056] Server: Permanently stores verbalized text data and viewpoint data, which are stored for later reuse and analysis.

[0057] Server: The stored data is provided with an endpoint for sharing with specific stakeholders and educational personnel.

[0058] 7. User Actions:

[0059] User (researcher): The researcher calibrates the eye tracking device and then begins their research through binoculars and a microscope.

[0060] User (researcher): Researchers can use feedback from the system to make adjustments to improve the quality of their observations.

[0061] Specific examples

[0062] Example 1: A researcher is observing a particular plant through binoculars. During this observation, an eye-tracking device captures the researcher's gaze data in real time and sends it to a server. The server analyzes the gaze data and verbalizes in text what the researcher was focusing on. This allows the researcher to easily save their observation records and share them with other researchers.

[0063] Example 2: Another researcher uses a microscope to observe cells. The eye-tracking device acquires gaze data and sends it to the server. The server then analyzes and verbalizes the observation in real time based on the gaze data. As a result, important observation points, such as abnormal cells, are quickly identified, and the analysis results are fed back to the researcher in real time. The researcher can immediately check the observation results and continue with more detailed analysis.

[0064] Thus, the present invention provides a system that significantly improves the efficiency of recording, analyzing, and sharing observation data in research settings.

[0065] The processing flow will be explained below.

[0066] Step 1:

[0067] Device: Initial setup of the eye tracking device, attaching it to the binoculars or microscope and calibrating it, ensuring that the device accurately tracks the researcher's gaze.

[0068] Step 2:

[0069] User: The researcher begins using binoculars or a microscope. After confirming that the object is in a state where detailed observation is possible, the researcher begins observation.

[0070] Step 3:

[0071] Device: Once observation begins, the eye-tracking device acquires gaze data in real time. The device captures gaze data every 0.1 seconds and buffers the information.

[0072] Step 4:

[0073] Terminal: The acquired viewpoint data is sent to the server at a predetermined interval (for example, every second) using an HTTP POST request.

[0074] Step 5:

[0075] Server: The server receives the viewpoint data, adds it to a queue for analysis, and temporarily stores it in a database.

[0076] Step 6:

[0077] Server: Analyzes the received viewpoint data. Based on the viewpoint data, the generation AI describes the objects and observations the researcher was looking at in natural language.

[0078] Step 7:

[0079] Server: Generates the analyzed data in text format, verbalizing in detail what the researchers were observing.

[0080] Step 8:

[0081] Server: Stores the verbalized data and corresponding viewpoint data in a permanent database, including metadata such as date and time and researcher ID.

[0082] Step 9:

[0083] Server: Sends analysis results and verbalized data back to the device in real time, allowing researchers to instantly check their observations.

[0084] Step 10:

[0085] Terminal: The analysis results sent back from the server are displayed in real time. Researchers can manage the progress of their observations based on the displayed information.

[0086] Step 11:

[0087] User: Researchers can view the displayed analysis results, record them if necessary, and generate a link to share the results with other researchers and educators.

[0088] Step 12:

[0089] Server: Generates a shared link for the data and makes it accessible to interested parties. The shared link contains credentials to access the specific dataset.

[0090] Step 13:

[0091] Users: Researchers share their observations by sending a shared link to stakeholders, making the data available for collaborative research and educational use.

[0092] Example 1

[0093] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0094] Conventional eye tracking systems only collect gaze data, but lack a mechanism for analyzing that data and providing immediate feedback to researchers. As a result, analyzing and sharing the data requires a great deal of time and effort, reducing the efficiency of research and education. In addition, gaze data analysis is often done manually, which poses problems with accuracy and speed.

[0095] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0096] In this invention, the server includes a means for using a generating AI to analyze viewpoint data, a means for displaying the viewpoint data acquired in real time as an analysis result, and a means for saving and sharing the viewpoint data and its analysis result, thereby enabling fast and accurate analysis and feedback of viewpoint data.

[0097] An "eye tracking device" is a device that detects a user's gaze and records that gaze movement as data.

[0098] "Gaze point data" refers to data that includes information such as the position and movement of the user's gaze, and the duration of gaze.

[0099] A "server" is a computer system that stores, analyzes, and sends and receives data over a network.

[0100] "Generative AI" is an artificial intelligence technology that analyzes and predicts based on large amounts of data.

[0101] "Calibration" is the initial adjustment operation that eye tracking devices perform to obtain accurate gaze data.

[0102] "Analysis" refers to processing acquired data using statistical and logical methods to clarify its meaning.

[0103] "Verbalization" is the conversion of data and information into natural language, i.e., words that humans can understand.

[0104] "Real-time" refers to a situation that corresponds to the current time and is reflected immediately with little delay.

[0105] "Storage" means recording acquired data in a computer's storage device so that it can be reused later.

[0106] "Sharing" means linking acquired data and analysis results with other users and systems and making them accessible.

[0107] This system collects and analyzes viewpoint data observed by researchers through binoculars or microscopes, and provides feedback on the results, thereby improving the efficiency of research and education.

[0108] Hardware and software used

[0109] Eye tracking device: A device that tracks a user's gaze as they observe using binoculars or a microscope.

[0110] Server: A computer system that stores viewpoint data and analyzes it using generative AI.

[0111] Generative AI: An artificial intelligence technology that analyzes viewpoint data and verbalizes it in natural language.

[0112] Examples of data processing and data calculation

[0113] User: The researcher attaches the eye tracking device to a pair of binoculars or a microscope and calibrates the device before beginning observations. Once calibrated, the device tracks the researcher's gaze and collects gaze data in real time.

[0114] Terminal: The eye tracking device transmits the gaze data acquired in real time to the server, where it is stored.

[0115] Server: The viewpoint data stored on the server is analyzed using a generative AI model. The analysis of viewpoint data involves the generative AI verbalizing the observations in natural language based on the viewpoint position, movement, attention time, etc.

[0116] Terminal: The verbalized analysis results are displayed in real time on the user's terminal, allowing researchers to immediately check the observation results.

[0117] Server: The text data and viewpoint data generated as a result of the analysis are stored for later reuse and analysis. In addition, an endpoint is provided for sharing the stored data with specific stakeholders and educational institutions.

[0118] Specific examples of operation

[0119] Example 1: When a researcher is observing a particular plant through binoculars, the eye-tracking device captures the researcher's gaze data in real time and sends it to the server. The server then analyzes the gaze data and verbalizes it in natural language, such as "The researcher focused on a specific part of the leaf for three minutes." The researcher can instantly view this information, easily save the observation record, and share it with other researchers.

[0120] Example 2: When another researcher observes cells under a microscope, the eye-tracking device captures gaze data and sends it to the server. The generative AI analyzes the gaze data to identify important observation points, such as abnormal cells, and generates a statement in natural language, such as "The researcher observed that a specific cell was abnormal." This information is displayed in real time on the researcher's device, allowing them to immediately check the observation results.

[0121] Prompt Sentence Examples

[0122] "Analyze the viewpoint data of the plant you observed through the binoculars and explain in detail which part you focused on."

[0123] "Identify the key points of the cells you observed under the microscope and describe them in words."

[0124] In this way, the system of the present invention contributes to improving the quality of research and education through rapid and accurate analysis and collaboration of viewpoint data.

[0125] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0126] Step 1:

[0127] Calibrating your eye tracking device

[0128] User: The researcher attaches the eye tracking device to a pair of binoculars or a microscope and calibrates it.

[0129] Input: Researcher's gaze position

[0130] Data processing: Researchers direct their gaze to multiple designated points, allowing the device to adjust for accurate gaze tracking.

[0131] Output: Calibration complete message

[0132] Specific actions: The researcher follows the instructions on the screen and completes the calibration by fixing their gaze on a target point in front of them for a few seconds.

[0133] Step 2:

[0134] Obtaining viewpoint data

[0135] Device: An eye-tracking device captures the researcher's gaze movements in real time.

[0136] Input: Researcher's gaze position and movement

[0137] Data processing: Sensors within the device record eye movements and convert them into gaze data.

[0138] Output: Obtained viewpoint data

[0139] How it works: As researchers observe through binoculars or a microscope, the device tracks their changes in gaze.

[0140] Step 3:

[0141] Viewpoint data transmission

[0142] Terminal: Sends the acquired viewpoint data to the server.

[0143] Input: Viewpoint data

[0144] Data processing: The viewpoint data is formatted into an analyzable format and sent to the server via the network.

[0145] Output: Viewpoint data sent to the server

[0146] Specific operation: The device packets viewpoint data and periodically sends it to the server via the network.

[0147] Step 4:

[0148] Saving viewpoint data

[0149] Server: Temporarily stores the received viewpoint data.

[0150] Input: Viewpoint data sent from the device

[0151] Data processing: The data is saved to a database and a backup is created at the same time.

[0152] Output: Saved viewpoint data

[0153] Specific operations: The server verifies the received data, records it in the database, and copies it to the backup system.

[0154] Step 5:

[0155] Analysis and verbalization of viewpoint data

[0156] Server: Analyzes viewpoint data and verbalizes it into natural language.

[0157] Input: Saved viewpoint data

[0158] Data processing: A generative AI model analyzes viewpoint data and expresses observations in natural language.

[0159] Output: Text data translated into natural language

[0160] Specific operation: The AI ​​model in the server analyzes the viewpoint data and generates text such as "The researcher focused on a specific part of the leaf for three minutes."

[0161] Step 6:

[0162] Displaying analysis results

[0163] Terminal: Displays the analysis results to the user in real time.

[0164] Input: Text data verbalized in natural language

[0165] Data processing: Convert into a display format and display on the user interface.

[0166] Output: Displayed analysis results

[0167] Specific operation: The analysis results pop up in real time on the user's device and the researcher can check them.

[0168] Step 7:

[0169] Data storage and sharing

[0170] Server: Stores and shares verbalized text data and viewpoint data.

[0171] Input: Text data expressed in natural language, viewpoint data

[0172] Data processing: Set access permissions for stored data and share it via API or portal.

[0173] Output: Saved data and its share link

[0174] Specific behavior: Record data in a database and set sharing permissions to allow access to specific parties.

[0175] The above are the processing steps of the system.

[0176] (Application example 1)

[0177] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0178] Maintenance of robots and machines is extremely important in factories and manufacturing sites, but the work is labor-intensive and time-consuming. Furthermore, collecting and analyzing viewpoint data and identifying problem areas during maintenance work are often difficult. This creates a need for improved efficiency and accuracy in maintenance work. Currently, it is difficult to identify abnormal areas and take prompt action, resulting in issues such as reduced productivity and increased operating costs.

[0179] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0180] In this invention, the server includes means for acquiring gaze data of the worker using an eye tracking device, means for analyzing and verbalizing the acquired gaze data, and means for analyzing the gaze data during maintenance work and identifying problem areas, thereby enabling accurate collection, analysis, and real-time feedback of gaze data during maintenance work.

[0181] An "eye tracking device" is a device for acquiring data on a worker's gaze.

[0182] "Gaze point data" is data that indicates the position and movement of the worker's line of sight.

[0183] A "server" is a computer system that receives, analyzes, and stores viewpoint data.

[0184] "Analysis" is the process of extracting information based on viewpoint data and verbalizing it.

[0185] "Verbalization" means converting the analyzed data into natural language.

[0186] "Maintenance work" refers to maintenance work on robots and machines in factories and manufacturing sites.

[0187] "Problem area" refers to an area that is abnormal or requires repair during maintenance work.

[0188] "Feedback" is the process of immediately notifying the operator of the analysis results.

[0189] MODE FOR CARRYING OUT THE INVENTION

[0190] This invention is a system for improving the efficiency and accuracy of maintenance work for robots and machines in factories and manufacturing sites. This system uses an eye-tracking device to acquire worker gaze data, analyzes and verbalizes the data on a server, and identifies problem areas. Specific embodiments of this system are described below.

[0191] Hardware and software used

[0192] Eye tracking device: A device for acquiring worker gaze data in real time (e.g., Tobii Eye Tracker).

[0193] Server: A computer system that receives, analyzes, and stores viewpoint data.

[0194] Smartphone or head-mounted display: A display device that allows workers to receive real-time feedback.

[0195] Eye tracking library: A software library used to collect gaze data.

[0196] Requests library: A Python library used to communicate with the server.

[0197] Generative AI (e.g., OpenAI GPT-3): An AI model for analyzing and verbalizing viewpoint data.

[0198] Process and Data Flow

[0199] 1. Obtaining viewpoint data

[0200] The eye tracking device collects the worker's gaze data in real time, which is then processed by software (Eye tracking library) within the device.

[0201] 2. Data transmission

[0202] The viewpoint data acquired by the device is periodically sent to the server using the Requests library, and the server receives and temporarily stores it.

[0203] 3. Data analysis and verbalization

[0204] The server passes the received viewpoint data to the generation AI, which analyzes the data and translates it into natural language. The generation AI then identifies abnormalities and points of interest based on the analysis of the viewpoint data.

[0205] 4. Feedback and Representation

[0206] The analysis results are displayed in real time on the worker's smartphone or head-mounted display, allowing the worker to immediately identify any abnormalities and take appropriate measures.

[0207] 5. Data storage and sharing

[0208] The verbalized text data and the original viewpoint data are stored on a server, where they can be reused and further analyzed later and shared with stakeholders.

[0209] Specific examples

[0210] Consider a scenario in which a maintenance worker in a factory is inspecting the joints of a robot with a magnifying glass. An eye-tracking device tracks the worker's gaze and sends the gaze data to a server. A generative AI analyzes this data and generates an analysis result, such as "the bolts in the joint are loose." The result is immediately displayed on the worker's smartphone, allowing the worker to quickly identify the problem area and take the necessary measures.

[0211] Prompt Sentence Examples

[0212] An example of a prompt to be input to a generative AI model would be something like this:

[0213] Verbalize the following observations from your user-perspective data:

[0214] Viewpoint Data:

[0215] Coordinates: (x1, y1), time: t1

[0216] Coordinates: (x2, y2), time: t2

[0217] ...

[0218] Coordinates: (xn, yn), Time: tn

[0219] Post-verbalization:

[0220] In this way, a system is constructed that improves the efficiency and accuracy of maintenance work for robots and machines in factories and manufacturing sites.

[0221] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0222] Step 1:

[0223] The user (worker) wears an eye-tracking device and begins maintenance work. The eye-tracking device acquires the worker's gaze data (position and movement of the gaze point) in real time. The input here is the worker's gaze point, and the output is the acquired raw gaze data.

[0224] Step 2:

[0225] The device (eye tracking device) processes the acquired gaze data in real time using the built-in Eye tracking library. The processed gaze data is temporarily buffered within the device. The input is the acquired gaze data, and the output is the processed gaze data.

[0226] Step 3:

[0227] The device periodically sends the processed viewpoint data to the server using the Requests library. An internet connection is required for transmission, and the data is packetized in JSON format. The input is the processed viewpoint data, and the output is the transmitted data packet.

[0228] Step 4:

[0229] The server receives the viewpoint data sent from the terminal and temporarily stores it in a database. The server checks the integrity of the data and verifies that there are no inconsistencies. The input is the sent data packet, and the output is the temporarily stored viewpoint data.

[0230] Step 5:

[0231] The server passes the saved viewpoint data to the generative AI model for analysis and verbalization. Based on the prompt, the model identifies abnormalities and points of interest from the viewpoint data and explains them in natural language. The input is the saved viewpoint data, and the output is the verbalized analysis results.

[0232] Step 6:

[0233] The server sends the generated analysis results in real time to the worker's smartphone or head-mounted display. The transmission is via an internet connection, and the data is displayed to the worker in real time. The input is the verbalized analysis results, and the output is the displayed analysis results.

[0234] Step 7:

[0235] The user (operator) checks the analysis results displayed in real time and takes the necessary measures promptly. The operator identifies the problem area and performs specific maintenance work based on that. The input is the displayed analysis results, and the output is the maintenance work that has been carried out.

[0236] Step 8:

[0237] The server permanently stores the verbalized text data and the original viewpoint data, allowing for later reuse and further analysis. The stored data may also be shared among stakeholders. The input is the verbalized analysis results and the original viewpoint data, and the output is the stored data.

[0238] In this way, through the specific operations, inputs and outputs at each step, maintenance work at factories and manufacturing sites can be made more efficient and more accurate.

[0239] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0240] This invention is a system that uses an AI-equipped eye-tracking device and emotion engine to collect viewpoint data observed by researchers through binoculars or microscopes, analyzes the data, and verbalizes it. Furthermore, by recognizing the user's emotions and adjusting the data analysis results based on that information, it aims to improve the quality of research and education.

[0241] System configuration

[0242] 1. Eye tracking device:

[0243] Device: Eye-tracking devices are attached to binoculars or microscopes to track the researcher's gaze, allowing for real-time acquisition of researcher gaze data.

[0244] 2. Emotion Engine:

[0245] Device: Equipped with an emotion engine, it recognizes the emotions felt by researchers during observation. It may use facial expression analysis and heart rate sensors.

[0246] 3. Transmitting viewpoint data:

[0247] Device: The acquired viewpoint data and emotion data are periodically sent to the server using HTTP POST requests.

[0248] 4. Temporary data storage:

[0249] Server: The server receives the viewpoint data and emotion data and temporarily stores them in a database.

[0250] 5. Data analysis and verbalization:

[0251] Server: Synchronizes and analyzes gaze data and emotion data, adjusting the analysis content based on what the researcher was looking at and changes in emotion.

[0252] Server: The generative AI verbalizes the observations in natural language based on viewpoint data and emotion data.

[0253] 6. Displaying the analysis results:

[0254] Terminal: The analysis results sent from the server are displayed to the researcher in real time, allowing the researcher to instantly check their observations and receive emotional feedback.

[0255] 7. Data Storage and Sharing:

[0256] Server: Stores the verbalized text data, viewpoint data, and emotion data in a permanent database. Data storage also includes metadata such as date and time and researcher ID.

[0257] Server: The stored data is provided with an endpoint for sharing with specific stakeholders and educational personnel.

[0258] 8. User Actions:

[0259] User (researcher): The researcher calibrates the eye tracking device and emotion engine, then begins the study through binoculars and a microscope.

[0260] User (researcher): Researchers can use feedback from the system to make adjustments to improve the quality of their observations.

[0261] Specific examples

[0262] Example 1: A researcher is observing a particular plant through binoculars. An eye-tracking device captures the researcher's gaze data in real time, and an emotion engine captures the researcher's emotion data. The server synchronizes these data and analyzes and verbalizes the observations. The analysis results reflect the gaze data and the corresponding researcher's emotions, allowing the researcher to conduct further analysis based on the emotional changes at specific viewpoints.

[0263] Example 2: Another researcher uses a microscope to observe cells. The eye-tracking device collects gaze data, and the emotion engine collects emotion data. The server analyzes this data and verbalizes the observations in real time. The data includes the researcher's emotional reactions when he or she discovers abnormal cells, and emotional changes such as stress and surprise are recorded along with the gaze data.

[0264] In this way, the present invention provides a system that significantly improves the efficiency of synchronously recording, analyzing, and sharing observation data and emotion data in research settings. Researchers can gain deeper insights by receiving feedback based on viewpoint and emotion data.

[0265] The processing flow will be explained below.

[0266] Step 1:

[0267] Device: Initial setup of the eye tracking device and emotion engine. Attach the device to the binoculars or microscope and calibrate it so that the device can accurately track the researcher's gaze and emotions.

[0268] Step 2:

[0269] User: The researcher begins using binoculars or a microscope. After confirming that the object is in a state where detailed observation is possible, the researcher begins observation.

[0270] Step 3:

[0271] Device: When observation begins, the eye-tracking device acquires gaze data in real time, and the emotion engine captures the user's emotional data (e.g., facial expressions and heart rate). The device captures data every 0.1 seconds and buffers the information.

[0272] Step 4:

[0273] Device: The acquired viewpoint data and emotion data are sent to the server at a predetermined interval (for example, every second). The data is sent using an HTTP POST request.

[0274] Step 5:

[0275] Server: The server receives the gaze data and emotion data, adds the received data to a queue for analysis, and temporarily stores it in a database.

[0276] Step 6:

[0277] Server: Analyzes the received viewpoint and emotion data. The generation AI uses the viewpoint data to describe the objects and observations the researcher was looking at in natural language, and adjusts the analysis results based on the emotion data.

[0278] Step 7:

[0279] Server: Generates verbalized text based on the viewpoint data and emotion data, and generates analysis results that include details of observations and emotional changes.

[0280] Step 8:

[0281] Server: The analysis results and verbalized data are sent from the server to the terminal and fed back to the researcher in real time.

[0282] Step 9:

[0283] Terminal: Displays the analysis results sent from the server in real time. Researchers can manage the progress of their observations based on the displayed information.

[0284] Step 10:

[0285] Server: Stores the data in a permanent database, including viewpoint data, emotion data, analysis results, and metadata such as date and time and researcher ID.

[0286] Step 11:

[0287] Server: Generates shared links based on stored data and provides endpoints for stakeholders to access them.

[0288] Step 12:

[0289] User: Researchers review the analysis results, record them as needed, and generate a link to share the results with other researchers and educators.

[0290] Step 13:

[0291] Users: Researchers can send a shared link to stakeholders to share the dataset containing observations and emotion data, allowing them to use the data for collaborative research and educational purposes.

[0292] Example 2

[0293] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0294] When researchers conduct observations using binoculars or microscopes, the current technological environment lacks the means to efficiently acquire and analyze viewpoint and emotion data, and then verbalize and present the results. Furthermore, systems for appropriately sharing acquired data with stakeholders and improving the quality of research results are limited. This calls for effective methods to improve the quality of research and education.

[0295] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0296] In this invention, the server includes a means for acquiring researcher's gaze data using an eye-tracking device, a means for acquiring researcher's emotion data using an emotion engine, a means for transmitting the acquired gaze data and emotion data to the server, a means for analyzing the gaze data and emotion data in the server and verbalizing the analysis results using a generative AI model, a means for saving the analyzed verbalized data, gaze data, and emotion data, and a means for sharing the saved data with relevant parties. This allows researchers to analyze the gaze data and emotion data obtained during observation in real time and provide the results as immediate feedback. Furthermore, by storing the data long-term and sharing it with various relevant parties, the quality of research and education can be improved.

[0297] An "eye tracking device" is a device that collects gaze data in real time when researchers observe through binoculars or a microscope.

[0298] The "Emotion Engine" is a system that recognizes the emotional state of researchers in real time, using facial expression analysis and heart rate sensors.

[0299] "Gaze data" is information that records the position and movement of the researcher's gaze, and is data that indicates what they are observing.

[0300] "Emotional data" refers to data on the researcher's emotional state, and is information obtained based on facial expression analysis, heart rate, etc.

[0301] A "generative AI model" is an artificial intelligence model that generates explanatory text and analysis results in natural language based on input data.

[0302] "Calibration" is a method of adjusting the performance of the eye tracking device and emotion engine to enable accurate data acquisition.

[0303] "Analysis results" are the results of analysis based on viewpoint data and emotion data, and include content verbalized through a generative AI model.

[0304] "Real-time display means" refers to devices or software that instantly display acquired and analyzed data and provide feedback to the user.

[0305] A "permanent database" is a database for long-term storage of data, allowing it to be accessed and analyzed at a later date.

[0306] "Stakeholders" refers to researchers, educators, and other relevant individuals and organizations.

[0307] This system allows researchers to efficiently collect viewpoint and emotion data during observations using binoculars and microscopes, and then analyze, verbalize, and provide feedback. This system can improve the quality of research and education.

[0308] 1. Use of eye tracking devices

[0309] The device is attached to a binocular or microscope and collects the researcher's gaze data in real time. The device tracks the position of the researcher's pupils and obtains gaze coordinates. The gaze data is acquired at several tens of frames per second and temporarily stored in the device's internal memory.

[0310] 2. Use of Emotion Engine

[0311] The device utilizes an emotion engine to recognize researchers' emotions in real time. Specifically, it captures the researchers' facial expressions with its built-in camera and runs an expression analysis algorithm. It also obtains data from a heart rate sensor to determine their emotional state.

[0312] 3. Transmission of viewpoint and emotion data

[0313] The device periodically sends the collected viewpoint data and emotion data to the server using HTTP POST requests. When sending, the viewpoint data and emotion data are packaged into packets in batch format.

[0314] 4. Temporary storage of data

[0315] The server temporarily stores the received viewpoint data and emotion data in a database. The server extracts the viewpoint data and emotion data from the payload of the HTTP request and inserts them into a database table for temporary storage.

[0316] 5. Data analysis and verbalization

[0317] The server synchronizes and analyzes the viewpoint data and emotion data, and then uses a generative AI model to verbalize the analysis results in natural language. The server links the viewpoint data and emotion data along a timeline, runs a data analysis algorithm, and identifies the researcher's observations. The server then inputs the analyzed data into the generative AI model, generating a natural language explanation.

[0318] 6. Displaying the analysis results

[0319] The terminal displays the analysis results sent from the server to the researcher in real time. The terminal receives the analysis results from the server, displays the received results on a display, and provides feedback to the researcher.

[0320] 7. Data Storage and Sharing

[0321] The server stores the resulting text, viewpoint, and emotion data in a permanent database, along with metadata such as the date and time and the researcher's ID. The server also provides an endpoint for sharing the stored data with interested parties.

[0322] 8. User Operations

[0323] The user (researcher) calibrates the eye tracking device and emotion engine, then begins observing through binoculars or a microscope. Based on feedback from the system, the user makes adjustments to improve the quality of their observations.

[0324] Specific examples

[0325] Example 1:

[0326] A researcher is observing a particular plant through binoculars. An eye-tracking device collects gaze data in real time, and an emotion engine collects the researcher's emotion data. The server synchronizes and analyzes this data, and verbalizes the observation in natural language. The analysis results reflect the gaze data and the corresponding emotion of the researcher.

[0327] Example 2:

[0328] Another researcher uses a microscope to observe the cells. The eye-tracking device collects gaze data, and the emotion engine collects emotion data. The server analyzes this data in real time and verbalizes the observations. The researcher's emotions upon discovering abnormal cells are recorded along with the gaze data.

[0329] Example prompts for generative AI models

[0330] "A user is observing a particular plant through binoculars. Please provide a detailed description of the observation in natural language based on the viewpoint and emotion data."

[0331] This system collects, analyzes, and verbalizes viewpoint and emotion data in real time, and provides feedback to researchers, enabling them to gain deeper insights based on the observations. Furthermore, by storing the data over the long term and sharing it with relevant parties, it is possible to improve the quality of research and education.

[0332] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0333] Step 1:

[0334] The terminal attaches an eye-tracking device to binoculars or microscopes and collects researchers' gaze data in real time.

[0335] Specific operation: The position of the researcher's pupils is tracked and the gaze coordinates are acquired. The collected gaze data is acquired at several tens of frames per second and temporarily stored in the device's internal memory.

[0336] Input: Researcher perspective.

[0337] Output: Real-time viewpoint data (coordinate information).

[0338] Step 2:

[0339] The device utilizes an emotion engine to recognize researchers' emotional data in real time.

[0340] How it works: It captures the researcher's facial expressions with a built-in camera and runs an expression analysis algorithm, while also obtaining data from a heart rate sensor to determine the researcher's emotional state.

[0341] Input: Researcher's facial expression and heart rate.

[0342] Output: Emotion data (emotional state identification information).

[0343] Step 3:

[0344] The device periodically sends the collected viewpoint data and emotion data to the server using HTTP POST requests.

[0345] Specific operation: Viewpoint data and emotion data are batched into packets and sent as an HTTP POST request.

[0346] Input: gaze data and emotion data.

[0347] Output: Data packets sent to the server.

[0348] Step 4:

[0349] The server temporarily stores the received viewpoint data and emotion data in a database.

[0350] Specific operation: Extract viewpoint data and emotion data from the HTTP request payload and insert them into a database table for temporary storage.

[0351] Input: Viewpoint and emotion data included in the HTTP request.

[0352] Output: Insertion of data into temporary storage database.

[0353] Step 5:

[0354] The server synchronizes and analyzes the viewpoint data and emotion data, and then verbalizes the analysis results in natural language using a generative AI model.

[0355] How it works: The viewpoint data and emotion data are linked over time, and a data analysis algorithm is run to identify the researcher's observations. The analyzed data is then fed into a generative AI model, which generates a natural language description.

[0356] Input: gaze data and emotion data.

[0357] Output: Analysis results written in natural language.

[0358] Step 6:

[0359] The terminal displays the analysis results sent from the server to the researcher in real time.

[0360] Specific operation: Receives analysis results from the server and displays the results on the screen.

[0361] Input: Parsed results from the server.

[0362] Output: Display analysis results on the terminal display.

[0363] Step 7:

[0364] The server stores the text data, viewpoint data, and emotion data of the analysis results in a permanent database.

[0365] Specific operation: The data containing the analysis results is inserted into a database for permanent storage. When saved, metadata such as the date and time and the researcher's ID are also included.

[0366] Input: Analysis result text data, viewpoint data, emotion data, and metadata.

[0367] Output: Saving data to a permanent database.

[0368] Step 8:

[0369] The server provides an endpoint for sharing the stored data with interested parties.

[0370] Specific behavior: Provide an API endpoint for accessing data so that relevant parties can obtain the data they need.

[0371] Input: Data access request from interested party.

[0372] Output: Providing data as requested.

[0373] Step 9:

[0374] The user (researcher) calibrates the eye tracking device and emotion engine, then begins observation through binoculars or a microscope.

[0375] Specific operation: Start the calibration procedure from the device settings screen, and perform viewpoint calibration and emotion recognition calibration. Once the settings are complete, begin observing through binoculars or a microscope.

[0376] Input: Data required for the calibration procedure (gaze information, emotion information).

[0377] Output: Accurate gaze and emotion data acquisition system after calibration is completed.

[0378] (Application example 2)

[0379] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0380] Modern factories are seeking to improve work efficiency and safety by using worker viewpoint and emotion data. However, there is no system that can analyze this data in real time and provide verbalized feedback, making it difficult to accurately grasp worker trends and improve work quality. Furthermore, there is a lack of feedback using analysis results that take workers' emotions into account, making it difficult to reduce work stress and improve safety.

[0381] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for acquiring viewpoint data and emotion data in real time, means for analyzing and verbalizing the viewpoint data and emotion data, and means for displaying the analysis results in real time. This makes it possible to accurately grasp the trends of workers, improve the quality of work, reduce work stress, and improve safety.

[0382] An "eye tracking device" is a device that tracks the movement of a user's eyes and acquires the data as gaze point data.

[0383] The "Emotion Engine" is a system that integrates sensors and software to analyze the emotional state of workers.

[0384] "Viewpoint data" is information that indicates the direction in which the worker is looking.

[0385] "Emotion data" is information that indicates the emotional state of the worker, and includes data on heart rate and facial expression analysis.

[0386] A "server" is a computer system that receives, analyzes, stores, and shares data over a network.

[0387] "Verbalization" is the process of describing analyzed data in natural language.

[0388] "Real-time display" is a function that instantly displays analyzed data to the user.

[0389] "Calibration" is the process of adjusting the accuracy of a device to ensure proper operation.

[0390] "Storage" refers to the act of recording data in a state that allows it to be referenced later.

[0391] "Sharing" is the act of making stored data available for handling with other users or systems.

[0392] The system that realizes this invention acquires, analyzes, and verbalizes the worker's viewpoint data and emotion data in real time, and provides feedback. The following describes the program content of this system, the hardware and software used, and specific examples.

[0393] First, a device (smartphone or head-mounted display) equipped with an eye-tracking device and emotion engine is used. These devices acquire the worker's gaze data and emotion data and send them to a server using an HTTP POST request. The server temporarily stores the received data in a database and begins analysis processing.

[0394] The server then uses the following software and modules to analyze the gaze data and emotion data:

[0395] Emotion Analysis Module: Analyzes emotion data and identifies the emotional state of the worker.

[0396] Viewpoint Analysis module: Analyzes viewpoint data and identifies worker eye movements.

[0397] Generative AI models (such as GPT-4): Generate analysis results in natural language based on viewpoint data and emotion data.

[0398] The server sends the analysis results to the terminal in real time and provides feedback, allowing workers to immediately receive advice on how to improve their work efficiency and safety.

[0399] Specific examples

[0400] For example, consider a situation in which a robot monitors the movements of workers in a factory. When a worker encounters a problem in a particular process, eye-gaze data and emotional data are immediately collected. The server analyzes this data and identifies the area in which the worker is feeling stressed. It then sends appropriate feedback to the worker's device in real time.

[0401] Prompt Sentence Examples

[0402] "Gaze data: 'The worker is looking at the screw point on the left side of the machine,' 'The worker is looking at the panel on the right side of the machine,' Emotion data: 'The worker's heart rate increases,' 'The worker's brow is furrowed,' Generate analysis results in natural language based on the relationship between the observation and the emotion."

[0403] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0404] Step 1:

[0405] The terminal (smartphone or head-mounted display) acquires the worker's viewpoint data and emotion data.

[0406] Specifically, the eye tracking device tracks the worker's gaze movements, and the emotion engine analyzes the worker's heart rate and facial expressions.

[0407] Input: Worker's viewpoint and emotion data

[0408] Output: Obtained gaze data and emotion data

[0409] Step 2:

[0410] The viewpoint data and emotion data acquired by the device are sent to the server via an HTTP POST request.

[0411] Specifically, the terminal converts the data into an appropriate format and transmits it to the specified endpoint.

[0412] Input: Obtained gaze data and emotion data

[0413] Output: A request to send data to the server

[0414] Step 3:

[0415] The server temporarily stores the received viewpoint data and emotion data in a database.

[0416] Specifically, the server checks the received data and inserts it into the database.

[0417] Input: Viewpoint data and emotion data sent from the device

[0418] Output: Viewpoint data and emotion data stored in a database

[0419] Step 4:

[0420] The server analyzes the viewpoint data and emotion data.

[0421] Specifically, the server performs data analysis using an Emotion Analysis module and a Viewpoint Analysis module.

[0422] Input: Viewpoint data and emotion data stored in a database

[0423] Output: Parsed gaze data and emotion data

[0424] Step 5:

[0425] Based on the analyzed data, the server uses a generative AI model (such as GPT-4) to translate the data into natural language.

[0426] Specifically, the server sends a prompt to GPT-4 and converts the analysis results into natural language.

[0427] Input: Parsed gaze data and emotion data

[0428] Output: Analysis results expressed in natural language

[0429] Step 6:

[0430] The server transmits the verbalized analysis results to the terminal.

[0431] As a specific operation, the server makes a request to transmit the generated natural language data to the terminal.

[0432] Input: Analysis results expressed in natural language

[0433] Output: Request to send analysis results to the device

[0434] Step 7:

[0435] The analysis results received by the device are displayed in real time.

[0436] Specifically, the device renders the analysis results in a display area, allowing the worker to see immediate feedback.

[0437] Input: Analysis results sent from the server in natural language

[0438] Output: Displaying the analysis results to the user

[0439] Through the above steps, this system is able to analyze the worker's viewpoint data and emotion data in real time and provide appropriate feedback.

[0440] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0441] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0442] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0443] [Second embodiment]

[0444] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0445] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0446] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0447] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0448] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0449] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0450] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0451] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0452] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0453] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0454] In the smart glasses 214, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0455] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0456] This invention is a system that uses an AI-equipped eye-tracking device to collect viewpoint data observed by researchers through binoculars or microscopes, analyzes the data, and verbalizes it. Furthermore, by storing and sharing this data, it aims to improve the quality of research and education.

[0457] System configuration

[0458] 1. Eye tracking device:

[0459] Terminals: Eye-tracking devices are attached to binoculars and microscopes to track the researcher's gaze.

[0460] Terminal: The device has the ability to capture eye movement and retrieve that data in real time.

[0461] 2. Transmitting viewpoint data:

[0462] Terminal: The acquired viewpoint data is sent to the server periodically to ensure that the viewpoint data reaches the server promptly.

[0463] 3. Saving viewpoint data:

[0464] Server: The server temporarily stores the viewpoint data it receives. A backup function may also be implemented to ensure that data is not lost.

[0465] 4. Analysis and verbalization of viewpoint data:

[0466] Server: Uses a generative AI to analyze the viewpoint data. The AI ​​processes the viewpoint data and verbalizes the observations in natural language.

[0467] Server: The verbalized data includes details about the points the researcher was observing and what they were looking at.

[0468] 5. Displaying the analysis results:

[0469] Terminal: The analysis results sent from the server are displayed to the researcher in real time, allowing the researcher to immediately check their observations.

[0470] 6. Data Storage and Sharing:

[0471] Server: Permanently stores verbalized text data and viewpoint data, which are stored for later reuse and analysis.

[0472] Server: The stored data is provided with an endpoint for sharing with specific stakeholders and educational personnel.

[0473] 7. User Actions:

[0474] User (researcher): The researcher calibrates the eye tracking device and then begins their research through binoculars and a microscope.

[0475] User (researcher): Researchers can use feedback from the system to make adjustments to improve the quality of their observations.

[0476] Specific examples

[0477] Example 1: A researcher is observing a particular plant through binoculars. During this observation, an eye-tracking device captures the researcher's gaze data in real time and sends it to a server. The server analyzes the gaze data and verbalizes in text what the researcher was focusing on. This allows the researcher to easily save their observation records and share them with other researchers.

[0478] Example 2: Another researcher uses a microscope to observe cells. The eye-tracking device acquires gaze data and sends it to the server. The server then analyzes and verbalizes the observation in real time based on the gaze data. As a result, important observation points, such as abnormal cells, are quickly identified, and the analysis results are fed back to the researcher in real time. The researcher can immediately check the observation results and continue with more detailed analysis.

[0479] Thus, the present invention provides a system that significantly improves the efficiency of recording, analyzing, and sharing observation data in research settings.

[0480] The processing flow will be explained below.

[0481] Step 1:

[0482] Device: Initial setup of the eye tracking device, attaching it to the binoculars or microscope and calibrating it, ensuring that the device accurately tracks the researcher's gaze.

[0483] Step 2:

[0484] User: The researcher begins using binoculars or a microscope. After confirming that the object is in a state where detailed observation is possible, the researcher begins observation.

[0485] Step 3:

[0486] Device: Once observation begins, the eye-tracking device acquires gaze data in real time. The device captures gaze data every 0.1 seconds and buffers the information.

[0487] Step 4:

[0488] Terminal: The acquired viewpoint data is sent to the server at a predetermined interval (for example, every second) using an HTTP POST request.

[0489] Step 5:

[0490] Server: The server receives the viewpoint data, adds it to a queue for analysis, and temporarily stores it in a database.

[0491] Step 6:

[0492] Server: Analyzes the received viewpoint data. Based on the viewpoint data, the generation AI describes the objects and observations the researcher was looking at in natural language.

[0493] Step 7:

[0494] Server: Generates the analyzed data in text format, verbalizing in detail what the researchers were observing.

[0495] Step 8:

[0496] Server: Stores the verbalized data and corresponding viewpoint data in a permanent database, including metadata such as date and time and researcher ID.

[0497] Step 9:

[0498] Server: Sends analysis results and verbalized data back to the device in real time, allowing researchers to instantly check their observations.

[0499] Step 10:

[0500] Terminal: The analysis results sent back from the server are displayed in real time. Researchers can manage the progress of their observations based on the displayed information.

[0501] Step 11:

[0502] User: Researchers can view the displayed analysis results, record them if necessary, and generate a link to share the results with other researchers and educators.

[0503] Step 12:

[0504] Server: Generates a shared link for the data and makes it accessible to interested parties. The shared link contains credentials to access the specific dataset.

[0505] Step 13:

[0506] Users: Researchers share their observations by sending a shared link to stakeholders, making the data available for collaborative research and educational use.

[0507] Example 1

[0508] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0509] Conventional eye tracking systems only collect gaze data, but lack a mechanism for analyzing that data and providing immediate feedback to researchers. As a result, analyzing and sharing the data requires a great deal of time and effort, reducing the efficiency of research and education. In addition, gaze data analysis is often done manually, which poses problems with accuracy and speed.

[0510] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0511] In this invention, the server includes a means for using a generating AI to analyze viewpoint data, a means for displaying the viewpoint data acquired in real time as an analysis result, and a means for saving and sharing the viewpoint data and its analysis result, thereby enabling fast and accurate analysis and feedback of viewpoint data.

[0512] An "eye tracking device" is a device that detects a user's gaze and records that gaze movement as data.

[0513] "Gaze point data" refers to data that includes information such as the position and movement of the user's gaze, and the duration of gaze.

[0514] A "server" is a computer system that stores, analyzes, and sends and receives data over a network.

[0515] "Generative AI" is an artificial intelligence technology that analyzes and predicts based on large amounts of data.

[0516] "Calibration" is the initial adjustment operation that eye tracking devices perform to obtain accurate gaze data.

[0517] "Analysis" refers to processing acquired data using statistical and logical methods to clarify its meaning.

[0518] "Verbalization" is the conversion of data and information into natural language, i.e., words that humans can understand.

[0519] "Real-time" refers to a situation that corresponds to the current time and is reflected immediately with little delay.

[0520] "Storage" means recording acquired data in a computer's storage device so that it can be reused later.

[0521] "Sharing" means linking acquired data and analysis results with other users and systems and making them accessible.

[0522] This system collects and analyzes viewpoint data observed by researchers through binoculars or microscopes, and provides feedback on the results, thereby improving the efficiency of research and education.

[0523] Hardware and software used

[0524] Eye tracking device: A device that tracks a user's gaze as they observe using binoculars or a microscope.

[0525] Server: A computer system that stores viewpoint data and analyzes it using generative AI.

[0526] Generative AI: An artificial intelligence technology that analyzes viewpoint data and verbalizes it in natural language.

[0527] Examples of data processing and data calculation

[0528] User: The researcher attaches the eye tracking device to a pair of binoculars or a microscope and calibrates the device before beginning observations. Once calibrated, the device tracks the researcher's gaze and collects gaze data in real time.

[0529] Terminal: The eye tracking device transmits the gaze data acquired in real time to the server, where it is stored.

[0530] Server: The viewpoint data stored on the server is analyzed using a generative AI model. The analysis of viewpoint data involves the generative AI verbalizing the observations in natural language based on the viewpoint position, movement, attention time, etc.

[0531] Terminal: The verbalized analysis results are displayed in real time on the user's terminal, allowing researchers to immediately check the observation results.

[0532] Server: The text data and viewpoint data generated as a result of the analysis are stored for later reuse and analysis. In addition, an endpoint is provided for sharing the stored data with specific stakeholders and educational institutions.

[0533] Specific examples of operation

[0534] Example 1: When a researcher is observing a particular plant through binoculars, the eye-tracking device captures the researcher's gaze data in real time and sends it to the server. The server then analyzes the gaze data and verbalizes it in natural language, such as "The researcher focused on a specific part of the leaf for three minutes." The researcher can instantly view this information, easily save the observation record, and share it with other researchers.

[0535] Example 2: When another researcher observes cells under a microscope, the eye-tracking device captures gaze data and sends it to the server. The generative AI analyzes the gaze data to identify important observation points, such as abnormal cells, and generates a statement in natural language, such as "The researcher observed that a specific cell was abnormal." This information is displayed in real time on the researcher's device, allowing them to immediately check the observation results.

[0536] Prompt Sentence Examples

[0537] "Analyze the viewpoint data of the plant you observed through the binoculars and explain in detail which part you focused on."

[0538] "Identify the key points of the cells you observed under the microscope and describe them in words."

[0539] In this way, the system of the present invention contributes to improving the quality of research and education through rapid and accurate analysis and collaboration of viewpoint data.

[0540] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0541] Step 1:

[0542] Calibrating your eye tracking device

[0543] User: The researcher attaches the eye tracking device to a pair of binoculars or a microscope and calibrates it.

[0544] Input: Researcher's gaze position

[0545] Data processing: Researchers direct their gaze to multiple designated points, allowing the device to adjust for accurate gaze tracking.

[0546] Output: Calibration complete message

[0547] Specific actions: The researcher follows the instructions on the screen and completes the calibration by fixing their gaze on a target point in front of them for a few seconds.

[0548] Step 2:

[0549] Obtaining viewpoint data

[0550] Device: An eye-tracking device captures the researcher's gaze movements in real time.

[0551] Input: Researcher's gaze position and movement

[0552] Data processing: Sensors within the device record eye movements and convert them into gaze data.

[0553] Output: Obtained viewpoint data

[0554] How it works: As researchers observe through binoculars or a microscope, the device tracks their changes in gaze.

[0555] Step 3:

[0556] Viewpoint data transmission

[0557] Terminal: Sends the acquired viewpoint data to the server.

[0558] Input: Viewpoint data

[0559] Data processing: The viewpoint data is formatted into an analyzable format and sent to the server via the network.

[0560] Output: Viewpoint data sent to the server

[0561] Specific operation: The device packets viewpoint data and periodically sends it to the server via the network.

[0562] Step 4:

[0563] Saving viewpoint data

[0564] Server: Temporarily stores the received viewpoint data.

[0565] Input: Viewpoint data sent from the device

[0566] Data processing: The data is saved to a database and a backup is created at the same time.

[0567] Output: Saved viewpoint data

[0568] Specific operations: The server verifies the received data, records it in the database, and copies it to the backup system.

[0569] Step 5:

[0570] Analysis and verbalization of viewpoint data

[0571] Server: Analyzes viewpoint data and verbalizes it into natural language.

[0572] Input: Saved viewpoint data

[0573] Data processing: A generative AI model analyzes viewpoint data and expresses observations in natural language.

[0574] Output: Text data translated into natural language

[0575] Specific operation: The AI ​​model in the server analyzes the viewpoint data and generates text such as "The researcher focused on a specific part of the leaf for three minutes."

[0576] Step 6:

[0577] Displaying analysis results

[0578] Terminal: Displays the analysis results to the user in real time.

[0579] Input: Text data verbalized in natural language

[0580] Data processing: Convert into a display format and display on the user interface.

[0581] Output: Displayed analysis results

[0582] Specific operation: The analysis results pop up in real time on the user's device and the researcher can check them.

[0583] Step 7:

[0584] Data storage and sharing

[0585] Server: Stores and shares verbalized text data and viewpoint data.

[0586] Input: Text data expressed in natural language, viewpoint data

[0587] Data processing: Set access permissions for stored data and share it via API or portal.

[0588] Output: Saved data and its share link

[0589] Specific behavior: Record data in a database and set sharing permissions to allow access to specific parties.

[0590] The above are the processing steps of the system.

[0591] (Application example 1)

[0592] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0593] Maintenance of robots and machines is extremely important in factories and manufacturing sites, but the work is labor-intensive and time-consuming. Furthermore, collecting and analyzing viewpoint data and identifying problem areas during maintenance work are often difficult. This creates a need for improved efficiency and accuracy in maintenance work. Currently, it is difficult to identify abnormal areas and take prompt action, resulting in issues such as reduced productivity and increased operating costs.

[0594] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0595] In this invention, the server includes means for acquiring gaze data of the worker using an eye tracking device, means for analyzing and verbalizing the acquired gaze data, and means for analyzing the gaze data during maintenance work and identifying problem areas, thereby enabling accurate collection, analysis, and real-time feedback of gaze data during maintenance work.

[0596] An "eye tracking device" is a device for acquiring data on a worker's gaze.

[0597] "Gaze point data" is data that indicates the position and movement of the worker's line of sight.

[0598] A "server" is a computer system that receives, analyzes, and stores viewpoint data.

[0599] "Analysis" is the process of extracting information based on viewpoint data and verbalizing it.

[0600] "Verbalization" means converting the analyzed data into natural language.

[0601] "Maintenance work" refers to maintenance work on robots and machines in factories and manufacturing sites.

[0602] "Problem area" refers to an area that is abnormal or requires repair during maintenance work.

[0603] "Feedback" is the process of immediately notifying the operator of the analysis results.

[0604] MODE FOR CARRYING OUT THE INVENTION

[0605] This invention is a system for improving the efficiency and accuracy of maintenance work for robots and machines in factories and manufacturing sites. This system uses an eye-tracking device to acquire worker gaze data, analyzes and verbalizes the data on a server, and identifies problem areas. Specific embodiments of this system are described below.

[0606] Hardware and software used

[0607] Eye tracking device: A device for acquiring worker gaze data in real time (e.g., Tobii Eye Tracker).

[0608] Server: A computer system that receives, analyzes, and stores viewpoint data.

[0609] Smartphone or head-mounted display: A display device that allows workers to receive real-time feedback.

[0610] Eye tracking library: A software library used to collect gaze data.

[0611] Requests library: A Python library used to communicate with the server.

[0612] Generative AI (e.g., OpenAI GPT-3): An AI model for analyzing and verbalizing viewpoint data.

[0613] Process and Data Flow

[0614] 1. Obtaining viewpoint data

[0615] The eye tracking device collects the worker's gaze data in real time, which is then processed by software (Eye tracking library) within the device.

[0616] 2. Data transmission

[0617] The viewpoint data acquired by the device is periodically sent to the server using the Requests library, and the server receives and temporarily stores it.

[0618] 3. Data analysis and verbalization

[0619] The server passes the received viewpoint data to the generation AI, which analyzes the data and translates it into natural language. The generation AI then identifies abnormalities and points of interest based on the analysis of the viewpoint data.

[0620] 4. Feedback and Representation

[0621] The analysis results are displayed in real time on the worker's smartphone or head-mounted display, allowing the worker to immediately identify any abnormalities and take appropriate measures.

[0622] 5. Data storage and sharing

[0623] The verbalized text data and the original viewpoint data are stored on a server, where they can be reused and further analyzed later and shared with stakeholders.

[0624] Specific examples

[0625] Consider a scenario in which a maintenance worker in a factory is inspecting the joints of a robot with a magnifying glass. An eye-tracking device tracks the worker's gaze and sends the gaze data to a server. A generative AI analyzes this data and generates an analysis result, such as "the bolts in the joint are loose." The result is immediately displayed on the worker's smartphone, allowing the worker to quickly identify the problem area and take the necessary measures.

[0626] Prompt Sentence Examples

[0627] An example of a prompt to be input to a generative AI model would be something like this:

[0628] Verbalize the following observations from your user-perspective data:

[0629] Viewpoint Data:

[0630] Coordinates: (x1, y1), time: t1

[0631] Coordinates: (x2, y2), time: t2

[0632] ...

[0633] Coordinates: (xn, yn), Time: tn

[0634] Post-verbalization:

[0635] In this way, a system is constructed that improves the efficiency and accuracy of maintenance work for robots and machines in factories and manufacturing sites.

[0636] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0637] Step 1:

[0638] The user (worker) wears an eye-tracking device and begins maintenance work. The eye-tracking device acquires the worker's gaze data (position and movement of the gaze point) in real time. The input here is the worker's gaze point, and the output is the acquired raw gaze data.

[0639] Step 2:

[0640] The device (eye tracking device) processes the acquired gaze data in real time using the built-in Eye tracking library. The processed gaze data is temporarily buffered within the device. The input is the acquired gaze data, and the output is the processed gaze data.

[0641] Step 3:

[0642] The device periodically sends the processed viewpoint data to the server using the Requests library. An internet connection is required for transmission, and the data is packetized in JSON format. The input is the processed viewpoint data, and the output is the transmitted data packet.

[0643] Step 4:

[0644] The server receives the viewpoint data sent from the terminal and temporarily stores it in a database. The server checks the integrity of the data and verifies that there are no inconsistencies. The input is the sent data packet, and the output is the temporarily stored viewpoint data.

[0645] Step 5:

[0646] The server passes the saved viewpoint data to the generative AI model for analysis and verbalization. Based on the prompt, the model identifies abnormalities and points of interest from the viewpoint data and explains them in natural language. The input is the saved viewpoint data, and the output is the verbalized analysis results.

[0647] Step 6:

[0648] The server sends the generated analysis results in real time to the worker's smartphone or head-mounted display. The transmission is via an internet connection, and the data is displayed to the worker in real time. The input is the verbalized analysis results, and the output is the displayed analysis results.

[0649] Step 7:

[0650] The user (operator) checks the analysis results displayed in real time and takes the necessary measures promptly. The operator identifies the problem area and performs specific maintenance work based on that. The input is the displayed analysis results, and the output is the maintenance work that has been carried out.

[0651] Step 8:

[0652] The server permanently stores the verbalized text data and the original viewpoint data, allowing for later reuse and further analysis. The stored data may also be shared among stakeholders. The input is the verbalized analysis results and the original viewpoint data, and the output is the stored data.

[0653] In this way, through the specific operations, inputs and outputs at each step, maintenance work at factories and manufacturing sites can be made more efficient and more accurate.

[0654] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0655] This invention is a system that uses an AI-equipped eye-tracking device and emotion engine to collect viewpoint data observed by researchers through binoculars or microscopes, analyzes the data, and verbalizes it. Furthermore, by recognizing the user's emotions and adjusting the data analysis results based on that information, it aims to improve the quality of research and education.

[0656] System configuration

[0657] 1. Eye tracking device:

[0658] Device: Eye-tracking devices are attached to binoculars or microscopes to track the researcher's gaze, allowing for real-time acquisition of researcher gaze data.

[0659] 2. Emotion Engine:

[0660] Device: Equipped with an emotion engine, it recognizes the emotions felt by researchers during observation. It may use facial expression analysis and heart rate sensors.

[0661] 3. Transmitting viewpoint data:

[0662] Device: The acquired viewpoint data and emotion data are periodically sent to the server using HTTP POST requests.

[0663] 4. Temporary data storage:

[0664] Server: The server receives the viewpoint data and emotion data and temporarily stores them in a database.

[0665] 5. Data analysis and verbalization:

[0666] Server: Synchronizes and analyzes gaze data and emotion data, adjusting the analysis content based on what the researcher was looking at and changes in emotion.

[0667] Server: The generative AI verbalizes the observations in natural language based on viewpoint data and emotion data.

[0668] 6. Displaying the analysis results:

[0669] Terminal: The analysis results sent from the server are displayed to the researcher in real time, allowing the researcher to instantly check their observations and receive emotional feedback.

[0670] 7. Data Storage and Sharing:

[0671] Server: Stores the verbalized text data, viewpoint data, and emotion data in a permanent database. Data storage also includes metadata such as date and time and researcher ID.

[0672] Server: The stored data is provided with an endpoint for sharing with specific stakeholders and educational personnel.

[0673] 8. User Actions:

[0674] User (researcher): The researcher calibrates the eye tracking device and emotion engine, then begins the study through binoculars and a microscope.

[0675] User (researcher): Researchers can use feedback from the system to make adjustments to improve the quality of their observations.

[0676] Specific examples

[0677] Example 1: A researcher is observing a particular plant through binoculars. An eye-tracking device captures the researcher's gaze data in real time, and an emotion engine captures the researcher's emotion data. The server synchronizes these data and analyzes and verbalizes the observations. The analysis results reflect the gaze data and the corresponding researcher's emotions, allowing the researcher to conduct further analysis based on the emotional changes at specific viewpoints.

[0678] Example 2: Another researcher uses a microscope to observe cells. The eye-tracking device collects gaze data, and the emotion engine collects emotion data. The server analyzes this data and verbalizes the observations in real time. The data includes the researcher's emotional reactions when he or she discovers abnormal cells, and emotional changes such as stress and surprise are recorded along with the gaze data.

[0679] In this way, the present invention provides a system that significantly improves the efficiency of synchronously recording, analyzing, and sharing observation data and emotion data in research settings. Researchers can gain deeper insights by receiving feedback based on viewpoint and emotion data.

[0680] The processing flow will be explained below.

[0681] Step 1:

[0682] Device: Initial setup of the eye tracking device and emotion engine. Attach the device to the binoculars or microscope and calibrate it so that the device can accurately track the researcher's gaze and emotions.

[0683] Step 2:

[0684] User: The researcher begins using binoculars or a microscope. After confirming that the object is in a state where detailed observation is possible, the researcher begins observation.

[0685] Step 3:

[0686] Device: When observation begins, the eye-tracking device acquires gaze data in real time, and the emotion engine captures the user's emotional data (e.g., facial expressions and heart rate). The device captures data every 0.1 seconds and buffers the information.

[0687] Step 4:

[0688] Device: The acquired viewpoint data and emotion data are sent to the server at a predetermined interval (for example, every second). The data is sent using an HTTP POST request.

[0689] Step 5:

[0690] Server: The server receives the gaze data and emotion data, adds the received data to a queue for analysis, and temporarily stores it in a database.

[0691] Step 6:

[0692] Server: Analyzes the received viewpoint and emotion data. The generation AI uses the viewpoint data to describe the objects and observations the researcher was looking at in natural language, and adjusts the analysis results based on the emotion data.

[0693] Step 7:

[0694] Server: Generates verbalized text based on the viewpoint data and emotion data, and generates analysis results that include details of observations and emotional changes.

[0695] Step 8:

[0696] Server: The analysis results and verbalized data are sent from the server to the terminal and fed back to the researcher in real time.

[0697] Step 9:

[0698] Terminal: Displays the analysis results sent from the server in real time. Researchers can manage the progress of their observations based on the displayed information.

[0699] Step 10:

[0700] Server: Stores the data in a permanent database, including viewpoint data, emotion data, analysis results, and metadata such as date and time and researcher ID.

[0701] Step 11:

[0702] Server: Generates shared links based on stored data and provides endpoints for stakeholders to access them.

[0703] Step 12:

[0704] User: Researchers review the analysis results, record them as needed, and generate a link to share the results with other researchers and educators.

[0705] Step 13:

[0706] Users: Researchers can send a shared link to stakeholders to share the dataset containing observations and emotion data, allowing them to use the data for collaborative research and educational purposes.

[0707] Example 2

[0708] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0709] When researchers conduct observations using binoculars or microscopes, the current technological environment lacks the means to efficiently acquire and analyze viewpoint and emotion data, and then verbalize and present the results. Furthermore, systems for appropriately sharing acquired data with stakeholders and improving the quality of research results are limited. This calls for effective methods to improve the quality of research and education.

[0710] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0711] In this invention, the server includes a means for acquiring researcher's gaze data using an eye-tracking device, a means for acquiring researcher's emotion data using an emotion engine, a means for transmitting the acquired gaze data and emotion data to the server, a means for analyzing the gaze data and emotion data in the server and verbalizing the analysis results using a generative AI model, a means for saving the analyzed verbalized data, gaze data, and emotion data, and a means for sharing the saved data with relevant parties. This allows researchers to analyze the gaze data and emotion data obtained during observation in real time and provide the results as immediate feedback. Furthermore, by storing the data long-term and sharing it with various relevant parties, the quality of research and education can be improved.

[0712] An "eye tracking device" is a device that collects gaze data in real time when researchers observe through binoculars or a microscope.

[0713] The "Emotion Engine" is a system that recognizes the emotional state of researchers in real time, using facial expression analysis and heart rate sensors.

[0714] "Gaze data" is information that records the position and movement of the researcher's gaze, and is data that indicates what they are observing.

[0715] "Emotional data" refers to data on the researcher's emotional state, and is information obtained based on facial expression analysis, heart rate, etc.

[0716] A "generative AI model" is an artificial intelligence model that generates explanatory text and analysis results in natural language based on input data.

[0717] "Calibration" is a method of adjusting the performance of the eye tracking device and emotion engine to enable accurate data acquisition.

[0718] "Analysis results" are the results of analysis based on viewpoint data and emotion data, and include content verbalized through a generative AI model.

[0719] "Real-time display means" refers to devices or software that instantly display acquired and analyzed data and provide feedback to the user.

[0720] A "permanent database" is a database for long-term storage of data, allowing it to be accessed and analyzed at a later date.

[0721] "Stakeholders" refers to researchers, educators, and other relevant individuals and organizations.

[0722] This system allows researchers to efficiently collect viewpoint and emotion data during observations using binoculars and microscopes, and then analyze, verbalize, and provide feedback. This system can improve the quality of research and education.

[0723] 1. Use of eye tracking devices

[0724] The device is attached to a binocular or microscope and collects the researcher's gaze data in real time. The device tracks the position of the researcher's pupils and obtains gaze coordinates. The gaze data is acquired at several tens of frames per second and temporarily stored in the device's internal memory.

[0725] 2. Use of Emotion Engine

[0726] The device utilizes an emotion engine to recognize researchers' emotions in real time. Specifically, it captures the researchers' facial expressions with its built-in camera and runs an expression analysis algorithm. It also obtains data from a heart rate sensor to determine their emotional state.

[0727] 3. Transmission of viewpoint and emotion data

[0728] The device periodically sends the collected viewpoint data and emotion data to the server using HTTP POST requests. When sending, the viewpoint data and emotion data are packaged into packets in batch format.

[0729] 4. Temporary storage of data

[0730] The server temporarily stores the received viewpoint data and emotion data in a database. The server extracts the viewpoint data and emotion data from the payload of the HTTP request and inserts them into a database table for temporary storage.

[0731] 5. Data analysis and verbalization

[0732] The server synchronizes and analyzes the viewpoint data and emotion data, and then uses a generative AI model to verbalize the analysis results in natural language. The server links the viewpoint data and emotion data along a timeline, runs a data analysis algorithm, and identifies the researcher's observations. The server then inputs the analyzed data into the generative AI model, generating a natural language explanation.

[0733] 6. Displaying the analysis results

[0734] The terminal displays the analysis results sent from the server to the researcher in real time. The terminal receives the analysis results from the server, displays the received results on a display, and provides feedback to the researcher.

[0735] 7. Data Storage and Sharing

[0736] The server stores the resulting text, viewpoint, and emotion data in a permanent database, along with metadata such as the date and time and the researcher's ID. The server also provides an endpoint for sharing the stored data with interested parties.

[0737] 8. User Operations

[0738] The user (researcher) calibrates the eye tracking device and emotion engine, then begins observing through binoculars or a microscope. Based on feedback from the system, the user makes adjustments to improve the quality of their observations.

[0739] Specific examples

[0740] Example 1:

[0741] A researcher is observing a particular plant through binoculars. An eye-tracking device collects gaze data in real time, and an emotion engine collects the researcher's emotion data. The server synchronizes and analyzes this data, and verbalizes the observation in natural language. The analysis results reflect the gaze data and the corresponding emotion of the researcher.

[0742] Example 2:

[0743] Another researcher uses a microscope to observe the cells. The eye-tracking device collects gaze data, and the emotion engine collects emotion data. The server analyzes this data in real time and verbalizes the observations. The researcher's emotions upon discovering abnormal cells are recorded along with the gaze data.

[0744] Example prompts for generative AI models

[0745] "A user is observing a particular plant through binoculars. Please provide a detailed description of the observation in natural language based on the viewpoint and emotion data."

[0746] This system collects, analyzes, and verbalizes viewpoint and emotion data in real time, and provides feedback to researchers, enabling them to gain deeper insights based on the observations. Furthermore, by storing the data over the long term and sharing it with relevant parties, it is possible to improve the quality of research and education.

[0747] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0748] Step 1:

[0749] The terminal attaches an eye-tracking device to binoculars or microscopes and collects researchers' gaze data in real time.

[0750] Specific operation: The position of the researcher's pupils is tracked and the gaze coordinates are acquired. The collected gaze data is acquired at several tens of frames per second and temporarily stored in the device's internal memory.

[0751] Input: Researcher perspective.

[0752] Output: Real-time viewpoint data (coordinate information).

[0753] Step 2:

[0754] The device utilizes an emotion engine to recognize researchers' emotional data in real time.

[0755] How it works: It captures the researcher's facial expressions with a built-in camera and runs an expression analysis algorithm, while also obtaining data from a heart rate sensor to determine the researcher's emotional state.

[0756] Input: Researcher's facial expression and heart rate.

[0757] Output: Emotion data (emotional state identification information).

[0758] Step 3:

[0759] The device periodically sends the collected viewpoint data and emotion data to the server using HTTP POST requests.

[0760] Specific operation: Viewpoint data and emotion data are batched into packets and sent as an HTTP POST request.

[0761] Input: gaze data and emotion data.

[0762] Output: Data packets sent to the server.

[0763] Step 4:

[0764] The server temporarily stores the received viewpoint data and emotion data in a database.

[0765] Specific operation: Extract viewpoint data and emotion data from the HTTP request payload and insert them into a database table for temporary storage.

[0766] Input: Viewpoint and emotion data included in the HTTP request.

[0767] Output: Insertion of data into temporary storage database.

[0768] Step 5:

[0769] The server synchronizes and analyzes the viewpoint data and emotion data, and then verbalizes the analysis results in natural language using a generative AI model.

[0770] How it works: The viewpoint data and emotion data are linked over time, and a data analysis algorithm is run to identify the researcher's observations. The analyzed data is then fed into a generative AI model, which generates a natural language description.

[0771] Input: gaze data and emotion data.

[0772] Output: Analysis results written in natural language.

[0773] Step 6:

[0774] The terminal displays the analysis results sent from the server to the researcher in real time.

[0775] Specific operation: Receives analysis results from the server and displays the results on the screen.

[0776] Input: Parsed results from the server.

[0777] Output: Display analysis results on the terminal display.

[0778] Step 7:

[0779] The server stores the text data, viewpoint data, and emotion data of the analysis results in a permanent database.

[0780] Specific operation: The data containing the analysis results is inserted into a database for permanent storage. When saved, metadata such as the date and time and the researcher's ID are also included.

[0781] Input: Analysis result text data, viewpoint data, emotion data, and metadata.

[0782] Output: Saving data to a permanent database.

[0783] Step 8:

[0784] The server provides an endpoint for sharing the stored data with interested parties.

[0785] Specific behavior: Provide an API endpoint for accessing data so that relevant parties can obtain the data they need.

[0786] Input: Data access request from interested party.

[0787] Output: Providing data as requested.

[0788] Step 9:

[0789] The user (researcher) calibrates the eye tracking device and emotion engine, then begins observation through binoculars or a microscope.

[0790] Specific operation: Start the calibration procedure from the device settings screen, and perform viewpoint calibration and emotion recognition calibration. Once the settings are complete, begin observing through binoculars or a microscope.

[0791] Input: Data required for the calibration procedure (gaze information, emotion information).

[0792] Output: Accurate gaze and emotion data acquisition system after calibration is completed.

[0793] (Application example 2)

[0794] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0795] Modern factories are seeking to improve work efficiency and safety by using worker viewpoint and emotion data. However, there is no system that can analyze this data in real time and provide verbalized feedback, making it difficult to accurately grasp worker trends and improve work quality. Furthermore, there is a lack of feedback using analysis results that take workers' emotions into account, making it difficult to reduce work stress and improve safety.

[0796] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for acquiring viewpoint data and emotion data in real time, means for analyzing and verbalizing the viewpoint data and emotion data, and means for displaying the analysis results in real time. This makes it possible to accurately grasp the trends of workers, improve the quality of work, reduce work stress, and improve safety.

[0797] An "eye tracking device" is a device that tracks the movement of a user's eyes and acquires the data as gaze point data.

[0798] The "Emotion Engine" is a system that integrates sensors and software to analyze the emotional state of workers.

[0799] "Viewpoint data" is information that indicates the direction in which the worker is looking.

[0800] "Emotion data" is information that indicates the emotional state of the worker, and includes data on heart rate and facial expression analysis.

[0801] A "server" is a computer system that receives, analyzes, stores, and shares data over a network.

[0802] "Verbalization" is the process of describing analyzed data in natural language.

[0803] "Real-time display" is a function that instantly displays analyzed data to the user.

[0804] "Calibration" is the process of adjusting the accuracy of a device to ensure proper operation.

[0805] "Storage" refers to the act of recording data in a state that allows it to be referenced later.

[0806] "Sharing" is the act of making stored data available for handling with other users or systems.

[0807] The system that realizes this invention acquires, analyzes, and verbalizes the worker's viewpoint data and emotion data in real time, and provides feedback. The following describes the program content of this system, the hardware and software used, and specific examples.

[0808] First, a device (smartphone or head-mounted display) equipped with an eye-tracking device and emotion engine is used. These devices acquire the worker's gaze data and emotion data and send them to a server using an HTTP POST request. The server temporarily stores the received data in a database and begins analysis processing.

[0809] The server then uses the following software and modules to analyze the gaze data and emotion data:

[0810] Emotion Analysis Module: Analyzes emotion data and identifies the emotional state of the worker.

[0811] Viewpoint Analysis module: Analyzes viewpoint data and identifies worker eye movements.

[0812] Generative AI models (such as GPT-4): Generate analysis results in natural language based on viewpoint data and emotion data.

[0813] The server sends the analysis results to the terminal in real time and provides feedback, allowing workers to immediately receive advice on how to improve their work efficiency and safety.

[0814] Specific examples

[0815] For example, consider a situation in which a robot monitors the movements of workers in a factory. When a worker encounters a problem in a particular process, eye-gaze data and emotional data are immediately collected. The server analyzes this data and identifies the area in which the worker is feeling stressed. It then sends appropriate feedback to the worker's device in real time.

[0816] Prompt Sentence Examples

[0817] "Gaze data: 'The worker is looking at the screw point on the left side of the machine,' 'The worker is looking at the panel on the right side of the machine,' Emotion data: 'The worker's heart rate increases,' 'The worker's brow is furrowed,' Generate analysis results in natural language based on the relationship between the observation and the emotion."

[0818] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0819] Step 1:

[0820] The terminal (smartphone or head-mounted display) acquires the worker's viewpoint data and emotion data.

[0821] Specifically, the eye tracking device tracks the worker's gaze movements, and the emotion engine analyzes the worker's heart rate and facial expressions.

[0822] Input: Worker's viewpoint and emotion data

[0823] Output: Obtained gaze data and emotion data

[0824] Step 2:

[0825] The viewpoint data and emotion data acquired by the device are sent to the server via an HTTP POST request.

[0826] Specifically, the terminal converts the data into an appropriate format and transmits it to the specified endpoint.

[0827] Input: Obtained gaze data and emotion data

[0828] Output: A request to send data to the server

[0829] Step 3:

[0830] The server temporarily stores the received viewpoint data and emotion data in a database.

[0831] Specifically, the server checks the received data and inserts it into the database.

[0832] Input: Viewpoint data and emotion data sent from the device

[0833] Output: Viewpoint data and emotion data stored in a database

[0834] Step 4:

[0835] The server analyzes the viewpoint data and emotion data.

[0836] Specifically, the server performs data analysis using an Emotion Analysis module and a Viewpoint Analysis module.

[0837] Input: Viewpoint data and emotion data stored in a database

[0838] Output: Parsed gaze data and emotion data

[0839] Step 5:

[0840] Based on the analyzed data, the server uses a generative AI model (such as GPT-4) to translate the data into natural language.

[0841] Specifically, the server sends a prompt to GPT-4 and converts the analysis results into natural language.

[0842] Input: Parsed gaze data and emotion data

[0843] Output: Analysis results expressed in natural language

[0844] Step 6:

[0845] The server transmits the verbalized analysis results to the terminal.

[0846] As a specific operation, the server makes a request to transmit the generated natural language data to the terminal.

[0847] Input: Analysis results expressed in natural language

[0848] Output: Request to send analysis results to the device

[0849] Step 7:

[0850] The analysis results received by the device are displayed in real time.

[0851] Specifically, the device renders the analysis results in a display area, allowing the worker to see immediate feedback.

[0852] Input: Analysis results sent from the server in natural language

[0853] Output: Displaying the analysis results to the user

[0854] Through the above steps, this system is able to analyze the worker's viewpoint data and emotion data in real time and provide appropriate feedback.

[0855] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0856] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0857] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0858] [Third embodiment]

[0859] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0860] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0861] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0862] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0863] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0864] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0865] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0866] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0867] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0868] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0869] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0870] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0871] This invention is a system that uses an AI-equipped eye-tracking device to collect viewpoint data observed by researchers through binoculars or microscopes, analyzes the data, and verbalizes it. Furthermore, by storing and sharing this data, it aims to improve the quality of research and education.

[0872] System configuration

[0873] 1. Eye tracking device:

[0874] Terminals: Eye-tracking devices are attached to binoculars and microscopes to track the researcher's gaze.

[0875] Terminal: The device has the ability to capture eye movement and retrieve that data in real time.

[0876] 2. Transmitting viewpoint data:

[0877] Terminal: The acquired viewpoint data is sent to the server periodically to ensure that the viewpoint data reaches the server promptly.

[0878] 3. Saving viewpoint data:

[0879] Server: The server temporarily stores the viewpoint data it receives. A backup function may also be implemented to ensure that data is not lost.

[0880] 4. Analysis and verbalization of viewpoint data:

[0881] Server: Uses a generative AI to analyze the viewpoint data. The AI ​​processes the viewpoint data and verbalizes the observations in natural language.

[0882] Server: The verbalized data includes details about the points the researcher was observing and what they were looking at.

[0883] 5. Displaying the analysis results:

[0884] Terminal: The analysis results sent from the server are displayed to the researcher in real time, allowing the researcher to immediately check their observations.

[0885] 6. Data Storage and Sharing:

[0886] Server: Permanently stores verbalized text data and viewpoint data, which are stored for later reuse and analysis.

[0887] Server: The stored data is provided with an endpoint for sharing with specific stakeholders and educational personnel.

[0888] 7. User Actions:

[0889] User (researcher): The researcher calibrates the eye tracking device and then begins their research through binoculars and a microscope.

[0890] User (researcher): Researchers can use feedback from the system to make adjustments to improve the quality of their observations.

[0891] Specific examples

[0892] Example 1: A researcher is observing a particular plant through binoculars. During this observation, an eye-tracking device captures the researcher's gaze data in real time and sends it to a server. The server analyzes the gaze data and verbalizes in text what the researcher was focusing on. This allows the researcher to easily save their observation records and share them with other researchers.

[0893] Example 2: Another researcher uses a microscope to observe cells. The eye-tracking device acquires gaze data and sends it to the server. The server then analyzes and verbalizes the observation in real time based on the gaze data. As a result, important observation points, such as abnormal cells, are quickly identified, and the analysis results are fed back to the researcher in real time. The researcher can immediately check the observation results and continue with more detailed analysis.

[0894] Thus, the present invention provides a system that significantly improves the efficiency of recording, analyzing, and sharing observation data in research settings.

[0895] The processing flow will be explained below.

[0896] Step 1:

[0897] Device: Initial setup of the eye tracking device, attaching it to the binoculars or microscope and calibrating it, ensuring that the device accurately tracks the researcher's gaze.

[0898] Step 2:

[0899] User: The researcher begins using binoculars or a microscope. After confirming that the object is in a state where detailed observation is possible, the researcher begins observation.

[0900] Step 3:

[0901] Device: Once observation begins, the eye-tracking device acquires gaze data in real time. The device captures gaze data every 0.1 seconds and buffers the information.

[0902] Step 4:

[0903] Terminal: The acquired viewpoint data is sent to the server at a predetermined interval (for example, every second) using an HTTP POST request.

[0904] Step 5:

[0905] Server: The server receives the viewpoint data, adds it to a queue for analysis, and temporarily stores it in a database.

[0906] Step 6:

[0907] Server: Analyzes the received viewpoint data. Based on the viewpoint data, the generation AI describes the objects and observations the researcher was looking at in natural language.

[0908] Step 7:

[0909] Server: Generates the analyzed data in text format, verbalizing in detail what the researchers were observing.

[0910] Step 8:

[0911] Server: Stores the verbalized data and corresponding viewpoint data in a permanent database, including metadata such as date and time and researcher ID.

[0912] Step 9:

[0913] Server: Sends analysis results and verbalized data back to the device in real time, allowing researchers to instantly check their observations.

[0914] Step 10:

[0915] Terminal: The analysis results sent back from the server are displayed in real time. Researchers can manage the progress of their observations based on the displayed information.

[0916] Step 11:

[0917] User: Researchers can view the displayed analysis results, record them if necessary, and generate a link to share the results with other researchers and educators.

[0918] Step 12:

[0919] Server: Generates a shared link for the data and makes it accessible to interested parties. The shared link contains credentials to access the specific dataset.

[0920] Step 13:

[0921] Users: Researchers share their observations by sending a shared link to stakeholders, making the data available for collaborative research and educational use.

[0922] Example 1

[0923] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0924] Conventional eye tracking systems only collect gaze data, but lack a mechanism for analyzing that data and providing immediate feedback to researchers. As a result, analyzing and sharing the data requires a great deal of time and effort, reducing the efficiency of research and education. In addition, gaze data analysis is often done manually, which poses problems with accuracy and speed.

[0925] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0926] In this invention, the server includes a means for using a generating AI to analyze viewpoint data, a means for displaying the viewpoint data acquired in real time as an analysis result, and a means for saving and sharing the viewpoint data and its analysis result, thereby enabling fast and accurate analysis and feedback of viewpoint data.

[0927] An "eye tracking device" is a device that detects a user's gaze and records that gaze movement as data.

[0928] "Gaze point data" refers to data that includes information such as the position and movement of the user's gaze, and the duration of gaze.

[0929] A "server" is a computer system that stores, analyzes, and sends and receives data over a network.

[0930] "Generative AI" is an artificial intelligence technology that analyzes and predicts based on large amounts of data.

[0931] "Calibration" is the initial adjustment operation that eye tracking devices perform to obtain accurate gaze data.

[0932] "Analysis" refers to processing acquired data using statistical and logical methods to clarify its meaning.

[0933] "Verbalization" is the conversion of data and information into natural language, i.e., words that humans can understand.

[0934] "Real-time" refers to a situation that corresponds to the current time and is reflected immediately with little delay.

[0935] "Storage" means recording acquired data in a computer's storage device so that it can be reused later.

[0936] "Sharing" means linking acquired data and analysis results with other users and systems and making them accessible.

[0937] This system collects and analyzes viewpoint data observed by researchers through binoculars or microscopes, and provides feedback on the results, thereby improving the efficiency of research and education.

[0938] Hardware and software used

[0939] Eye tracking device: A device that tracks a user's gaze as they observe using binoculars or a microscope.

[0940] Server: A computer system that stores viewpoint data and analyzes it using generative AI.

[0941] Generative AI: An artificial intelligence technology that analyzes viewpoint data and verbalizes it in natural language.

[0942] Examples of data processing and data calculation

[0943] User: The researcher attaches the eye tracking device to a pair of binoculars or a microscope and calibrates the device before beginning observations. Once calibrated, the device tracks the researcher's gaze and collects gaze data in real time.

[0944] Terminal: The eye tracking device transmits the gaze data acquired in real time to the server, where it is stored.

[0945] Server: The viewpoint data stored on the server is analyzed using a generative AI model. The analysis of viewpoint data involves the generative AI verbalizing the observations in natural language based on the viewpoint position, movement, attention time, etc.

[0946] Terminal: The verbalized analysis results are displayed in real time on the user's terminal, allowing researchers to immediately check the observation results.

[0947] Server: The text data and viewpoint data generated as a result of the analysis are stored for later reuse and analysis. In addition, an endpoint is provided for sharing the stored data with specific stakeholders and educational institutions.

[0948] Specific examples of operation

[0949] Example 1: When a researcher is observing a particular plant through binoculars, the eye-tracking device captures the researcher's gaze data in real time and sends it to the server. The server then analyzes the gaze data and verbalizes it in natural language, such as "The researcher focused on a specific part of the leaf for three minutes." The researcher can instantly view this information, easily save the observation record, and share it with other researchers.

[0950] Example 2: When another researcher observes cells under a microscope, the eye-tracking device captures gaze data and sends it to the server. The generative AI analyzes the gaze data to identify important observation points, such as abnormal cells, and generates a statement in natural language, such as "The researcher observed that a specific cell was abnormal." This information is displayed in real time on the researcher's device, allowing them to immediately check the observation results.

[0951] Prompt Sentence Examples

[0952] "Analyze the viewpoint data of the plant you observed through the binoculars and explain in detail which part you focused on."

[0953] "Identify the key points of the cells you observed under the microscope and describe them in words."

[0954] In this way, the system of the present invention contributes to improving the quality of research and education through rapid and accurate analysis and collaboration of viewpoint data.

[0955] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0956] Step 1:

[0957] Calibrating your eye tracking device

[0958] User: The researcher attaches the eye tracking device to a pair of binoculars or a microscope and calibrates it.

[0959] Input: Researcher's gaze position

[0960] Data processing: Researchers direct their gaze to multiple designated points, allowing the device to adjust for accurate gaze tracking.

[0961] Output: Calibration complete message

[0962] Specific actions: The researcher follows the instructions on the screen and completes the calibration by fixing their gaze on a target point in front of them for a few seconds.

[0963] Step 2:

[0964] Obtaining viewpoint data

[0965] Device: An eye-tracking device captures the researcher's gaze movements in real time.

[0966] Input: Researcher's gaze position and movement

[0967] Data processing: Sensors within the device record eye movements and convert them into gaze data.

[0968] Output: Obtained viewpoint data

[0969] How it works: As researchers observe through binoculars or a microscope, the device tracks their changes in gaze.

[0970] Step 3:

[0971] Viewpoint data transmission

[0972] Terminal: Sends the acquired viewpoint data to the server.

[0973] Input: Viewpoint data

[0974] Data processing: The viewpoint data is formatted into an analyzable format and sent to the server via the network.

[0975] Output: Viewpoint data sent to the server

[0976] Specific operation: The device packets viewpoint data and periodically sends it to the server via the network.

[0977] Step 4:

[0978] Saving viewpoint data

[0979] Server: Temporarily stores the received viewpoint data.

[0980] Input: Viewpoint data sent from the device

[0981] Data processing: The data is saved to a database and a backup is created at the same time.

[0982] Output: Saved viewpoint data

[0983] Specific operations: The server verifies the received data, records it in the database, and copies it to the backup system.

[0984] Step 5:

[0985] Analysis and verbalization of viewpoint data

[0986] Server: Analyzes viewpoint data and verbalizes it into natural language.

[0987] Input: Saved viewpoint data

[0988] Data processing: A generative AI model analyzes viewpoint data and expresses observations in natural language.

[0989] Output: Text data translated into natural language

[0990] Specific operation: The AI ​​model in the server analyzes the viewpoint data and generates text such as "The researcher focused on a specific part of the leaf for three minutes."

[0991] Step 6:

[0992] Displaying analysis results

[0993] Terminal: Displays the analysis results to the user in real time.

[0994] Input: Text data verbalized in natural language

[0995] Data processing: Convert into a display format and display on the user interface.

[0996] Output: Displayed analysis results

[0997] Specific operation: The analysis results pop up in real time on the user's device and the researcher can check them.

[0998] Step 7:

[0999] Data storage and sharing

[1000] Server: Stores and shares verbalized text data and viewpoint data.

[1001] Input: Text data expressed in natural language, viewpoint data

[1002] Data processing: Set access permissions for stored data and share it via API or portal.

[1003] Output: Saved data and its share link

[1004] Specific behavior: Record data in a database and set sharing permissions to allow access to specific parties.

[1005] The above are the processing steps of the system.

[1006] (Application example 1)

[1007] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1008] Maintenance of robots and machines is extremely important in factories and manufacturing sites, but the work is labor-intensive and time-consuming. Furthermore, collecting and analyzing viewpoint data and identifying problem areas during maintenance work are often difficult. This creates a need for improved efficiency and accuracy in maintenance work. Currently, it is difficult to identify abnormal areas and take prompt action, resulting in issues such as reduced productivity and increased operating costs.

[1009] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1010] In this invention, the server includes means for acquiring gaze data of the worker using an eye tracking device, means for analyzing and verbalizing the acquired gaze data, and means for analyzing the gaze data during maintenance work and identifying problem areas, thereby enabling accurate collection, analysis, and real-time feedback of gaze data during maintenance work.

[1011] An "eye tracking device" is a device for acquiring data on a worker's gaze.

[1012] "Gaze point data" is data that indicates the position and movement of the worker's line of sight.

[1013] A "server" is a computer system that receives, analyzes, and stores viewpoint data.

[1014] "Analysis" is the process of extracting information based on viewpoint data and verbalizing it.

[1015] "Verbalization" means converting the analyzed data into natural language.

[1016] "Maintenance work" refers to maintenance work on robots and machines in factories and manufacturing sites.

[1017] "Problem area" refers to an area that is abnormal or requires repair during maintenance work.

[1018] "Feedback" is the process of immediately notifying the operator of the analysis results.

[1019] MODE FOR CARRYING OUT THE INVENTION

[1020] This invention is a system for improving the efficiency and accuracy of maintenance work for robots and machines in factories and manufacturing sites. This system uses an eye-tracking device to acquire worker gaze data, analyzes and verbalizes the data on a server, and identifies problem areas. Specific embodiments of this system are described below.

[1021] Hardware and software used

[1022] Eye tracking device: A device for acquiring worker gaze data in real time (e.g., Tobii Eye Tracker).

[1023] Server: A computer system that receives, analyzes, and stores viewpoint data.

[1024] Smartphone or head-mounted display: A display device that allows workers to receive real-time feedback.

[1025] Eye tracking library: A software library used to collect gaze data.

[1026] Requests library: A Python library used to communicate with the server.

[1027] Generative AI (e.g., OpenAI GPT-3): An AI model for analyzing and verbalizing viewpoint data.

[1028] Process and Data Flow

[1029] 1. Obtaining viewpoint data

[1030] The eye tracking device collects the worker's gaze data in real time, which is then processed by software (Eye tracking library) within the device.

[1031] 2. Data transmission

[1032] The viewpoint data acquired by the device is periodically sent to the server using the Requests library, and the server receives and temporarily stores it.

[1033] 3. Data analysis and verbalization

[1034] The server passes the received viewpoint data to the generation AI, which analyzes the data and translates it into natural language. The generation AI then identifies abnormalities and points of interest based on the analysis of the viewpoint data.

[1035] 4. Feedback and Representation

[1036] The analysis results are displayed in real time on the worker's smartphone or head-mounted display, allowing the worker to immediately identify any abnormalities and take appropriate measures.

[1037] 5. Data storage and sharing

[1038] The verbalized text data and the original viewpoint data are stored on a server, where they can be reused and further analyzed later and shared with stakeholders.

[1039] Specific examples

[1040] Consider a scenario in which a maintenance worker in a factory is inspecting the joints of a robot with a magnifying glass. An eye-tracking device tracks the worker's gaze and sends the gaze data to a server. A generative AI analyzes this data and generates an analysis result, such as "the bolts in the joint are loose." The result is immediately displayed on the worker's smartphone, allowing the worker to quickly identify the problem area and take the necessary measures.

[1041] Prompt Sentence Examples

[1042] An example of a prompt to be input to a generative AI model would be something like this:

[1043] Verbalize the following observations from your user-perspective data:

[1044] Viewpoint Data:

[1045] Coordinates: (x1, y1), time: t1

[1046] Coordinates: (x2, y2), time: t2

[1047] ...

[1048] Coordinates: (xn, yn), Time: tn

[1049] Post-verbalization:

[1050] In this way, a system is constructed that improves the efficiency and accuracy of maintenance work for robots and machines in factories and manufacturing sites.

[1051] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1052] Step 1:

[1053] The user (worker) wears an eye-tracking device and begins maintenance work. The eye-tracking device acquires the worker's gaze data (position and movement of the gaze point) in real time. The input here is the worker's gaze point, and the output is the acquired raw gaze data.

[1054] Step 2:

[1055] The device (eye tracking device) processes the acquired gaze data in real time using the built-in Eye tracking library. The processed gaze data is temporarily buffered within the device. The input is the acquired gaze data, and the output is the processed gaze data.

[1056] Step 3:

[1057] The device periodically sends the processed viewpoint data to the server using the Requests library. An internet connection is required for transmission, and the data is packetized in JSON format. The input is the processed viewpoint data, and the output is the transmitted data packet.

[1058] Step 4:

[1059] The server receives the viewpoint data sent from the terminal and temporarily stores it in a database. The server checks the integrity of the data and verifies that there are no inconsistencies. The input is the sent data packet, and the output is the temporarily stored viewpoint data.

[1060] Step 5:

[1061] The server passes the saved viewpoint data to the generative AI model for analysis and verbalization. Based on the prompt, the model identifies abnormalities and points of interest from the viewpoint data and explains them in natural language. The input is the saved viewpoint data, and the output is the verbalized analysis results.

[1062] Step 6:

[1063] The server sends the generated analysis results in real time to the worker's smartphone or head-mounted display. The transmission is via an internet connection, and the data is displayed to the worker in real time. The input is the verbalized analysis results, and the output is the displayed analysis results.

[1064] Step 7:

[1065] The user (worker) checks the analysis results displayed in real time and takes the necessary measures promptly. The worker identifies the problem area and performs specific maintenance work based on that. The input is the displayed analysis results, and the output is the maintenance work that has been carried out.

[1066] Step 8:

[1067] The server permanently stores the verbalized text data and the original viewpoint data, allowing for later reuse and further analysis. The saved data may also be shared among stakeholders. The input is the verbalized analysis results and the original viewpoint data, and the output is the saved data.

[1068] In this way, through the specific operations, inputs and outputs at each step, maintenance work at factories and manufacturing sites can be made more efficient and more accurate.

[1069] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1070] This invention is a system that uses an AI-equipped eye-tracking device and emotion engine to collect viewpoint data observed by researchers through binoculars or microscopes, analyzes the data, and verbalizes it. Furthermore, by recognizing the user's emotions and adjusting the data analysis results based on that information, it aims to improve the quality of research and education.

[1071] System configuration

[1072] 1. Eye tracking device:

[1073] Device: Eye-tracking devices are attached to binoculars or microscopes to track the researcher's gaze, allowing for real-time acquisition of researcher gaze data.

[1074] 2. Emotion Engine:

[1075] Device: Equipped with an emotion engine, it recognizes the emotions felt by researchers during observation. It may use facial expression analysis and heart rate sensors.

[1076] 3. Transmitting viewpoint data:

[1077] Device: The acquired viewpoint data and emotion data are periodically sent to the server using HTTP POST requests.

[1078] 4. Temporary data storage:

[1079] Server: The server receives the viewpoint data and emotion data and temporarily stores them in a database.

[1080] 5. Data analysis and verbalization:

[1081] Server: Synchronizes and analyzes gaze data and emotion data, adjusting the analysis content based on what the researcher was looking at and changes in emotion.

[1082] Server: The generative AI verbalizes the observations in natural language based on viewpoint data and emotion data.

[1083] 6. Displaying the analysis results:

[1084] Terminal: The analysis results sent from the server are displayed to the researcher in real time, allowing the researcher to instantly check their observations and receive emotional feedback.

[1085] 7. Data Storage and Sharing:

[1086] Server: Stores the verbalized text data, viewpoint data, and emotion data in a permanent database. Data storage also includes metadata such as date and time and researcher ID.

[1087] Server: The stored data is provided with an endpoint for sharing with specific stakeholders and educational personnel.

[1088] 8. User Actions:

[1089] User (researcher): The researcher calibrates the eye tracking device and emotion engine, then begins the study through binoculars and a microscope.

[1090] User (researcher): Researchers can use feedback from the system to make adjustments to improve the quality of their observations.

[1091] Specific examples

[1092] Example 1: A researcher is observing a particular plant through binoculars. An eye-tracking device captures the researcher's gaze data in real time, and an emotion engine captures the researcher's emotion data. The server synchronizes these data and analyzes and verbalizes the observations. The analysis results reflect the gaze data and the corresponding researcher's emotions, allowing the researcher to conduct further analysis based on the emotional changes at specific viewpoints.

[1093] Example 2: Another researcher uses a microscope to observe cells. The eye-tracking device collects gaze data, and the emotion engine collects emotion data. The server analyzes this data and verbalizes the observations in real time. The data includes the researcher's emotional reactions when he or she discovers abnormal cells, and emotional changes such as stress and surprise are recorded along with the gaze data.

[1094] In this way, the present invention provides a system that significantly improves the efficiency of synchronously recording, analyzing, and sharing observation data and emotion data in research settings. Researchers can gain deeper insights by receiving feedback based on viewpoint and emotion data.

[1095] The processing flow will be explained below.

[1096] Step 1:

[1097] Device: Initial setup of the eye tracking device and emotion engine. Attach the device to the binoculars or microscope and calibrate it so that the device can accurately track the researcher's gaze and emotions.

[1098] Step 2:

[1099] User: The researcher begins using binoculars or a microscope. After confirming that the object is in a state where detailed observation is possible, the researcher begins observation.

[1100] Step 3:

[1101] Device: When observation begins, the eye-tracking device acquires gaze data in real time, and the emotion engine captures the user's emotional data (e.g., facial expressions and heart rate). The device captures data every 0.1 seconds and buffers the information.

[1102] Step 4:

[1103] Device: The acquired viewpoint data and emotion data are sent to the server at a predetermined interval (for example, every second). The data is sent using an HTTP POST request.

[1104] Step 5:

[1105] Server: The server receives the gaze data and emotion data, adds the received data to a queue for analysis, and temporarily stores it in a database.

[1106] Step 6:

[1107] Server: Analyzes the received viewpoint and emotion data. The generation AI uses the viewpoint data to describe the objects and observations the researcher was looking at in natural language, and adjusts the analysis results based on the emotion data.

[1108] Step 7:

[1109] Server: Generates verbalized text based on the viewpoint data and emotion data, and generates analysis results that include details of observations and emotional changes.

[1110] Step 8:

[1111] Server: The analysis results and verbalized data are sent from the server to the terminal and fed back to the researcher in real time.

[1112] Step 9:

[1113] Terminal: Displays the analysis results sent from the server in real time. Researchers can manage the progress of their observations based on the displayed information.

[1114] Step 10:

[1115] Server: Stores the data in a permanent database, including viewpoint data, emotion data, analysis results, and metadata such as date and time and researcher ID.

[1116] Step 11:

[1117] Server: Generates shared links based on stored data and provides endpoints for stakeholders to access them.

[1118] Step 12:

[1119] User: Researchers review the analysis results, record them as needed, and generate a link to share the results with other researchers and educators.

[1120] Step 13:

[1121] Users: Researchers can send a shared link to stakeholders to share the dataset containing observations and emotion data, allowing them to use the data for collaborative research and educational purposes.

[1122] Example 2

[1123] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1124] When researchers conduct observations using binoculars or microscopes, the current technological environment lacks the means to efficiently acquire and analyze viewpoint and emotion data, and then verbalize and present the results. Furthermore, systems for appropriately sharing acquired data with stakeholders and improving the quality of research results are limited. This calls for effective methods to improve the quality of research and education.

[1125] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1126] In this invention, the server includes a means for acquiring researcher's gaze data using an eye-tracking device, a means for acquiring researcher's emotion data using an emotion engine, a means for transmitting the acquired gaze data and emotion data to the server, a means for analyzing the gaze data and emotion data in the server and verbalizing the analysis results using a generative AI model, a means for saving the analyzed verbalized data, gaze data, and emotion data, and a means for sharing the saved data with relevant parties. This allows researchers to analyze the gaze data and emotion data obtained during observation in real time and provide the results as immediate feedback. Furthermore, by storing the data long-term and sharing it with various relevant parties, the quality of research and education can be improved.

[1127] An "eye tracking device" is a device that collects gaze data in real time when researchers observe through binoculars or a microscope.

[1128] The "Emotion Engine" is a system that recognizes the emotional state of researchers in real time, using facial expression analysis and heart rate sensors.

[1129] "Gaze data" is information that records the position and movement of the researcher's gaze, and is data that indicates what they are observing.

[1130] "Emotional data" refers to data on the researcher's emotional state, and is information obtained based on facial expression analysis, heart rate, etc.

[1131] A "generative AI model" is an artificial intelligence model that generates explanatory text and analysis results in natural language based on input data.

[1132] "Calibration" is a method of adjusting the performance of the eye tracking device and emotion engine to enable accurate data acquisition.

[1133] "Analysis results" are the results of analysis based on viewpoint data and emotion data, and include content verbalized through a generative AI model.

[1134] "Real-time display means" refers to devices or software that instantly display acquired and analyzed data and provide feedback to the user.

[1135] A "permanent database" is a database for long-term storage of data, allowing it to be accessed and analyzed at a later date.

[1136] "Stakeholders" refers to researchers, educators, and other relevant individuals and organizations.

[1137] This system allows researchers to efficiently collect viewpoint and emotion data during observations using binoculars and microscopes, and then analyze, verbalize, and provide feedback. This system can improve the quality of research and education.

[1138] 1. Use of eye tracking devices

[1139] The device is attached to a binocular or microscope and collects the researcher's gaze data in real time. The device tracks the position of the researcher's pupils and obtains gaze coordinates. The gaze data is acquired at several tens of frames per second and temporarily stored in the device's internal memory.

[1140] 2. Use of Emotion Engine

[1141] The device utilizes an emotion engine to recognize researchers' emotions in real time. Specifically, it captures the researchers' facial expressions with its built-in camera and runs an expression analysis algorithm. It also obtains data from a heart rate sensor to determine their emotional state.

[1142] 3. Transmission of viewpoint and emotion data

[1143] The device periodically sends the collected viewpoint data and emotion data to the server using HTTP POST requests. When sending, the viewpoint data and emotion data are packaged into packets in batch format.

[1144] 4. Temporary storage of data

[1145] The server temporarily stores the received viewpoint data and emotion data in a database. The server extracts the viewpoint data and emotion data from the payload of the HTTP request and inserts them into a database table for temporary storage.

[1146] 5. Data analysis and verbalization

[1147] The server synchronizes and analyzes the viewpoint data and emotion data, and then uses a generative AI model to verbalize the analysis results in natural language. The server links the viewpoint data and emotion data along a timeline, runs a data analysis algorithm, and identifies the researcher's observations. The server then inputs the analyzed data into the generative AI model, generating a natural language explanation.

[1148] 6. Displaying the analysis results

[1149] The terminal displays the analysis results sent from the server to the researcher in real time. The terminal receives the analysis results from the server, displays the received results on a display, and provides feedback to the researcher.

[1150] 7. Data Storage and Sharing

[1151] The server stores the resulting text, viewpoint, and emotion data in a permanent database, along with metadata such as the date and time and the researcher's ID. The server also provides an endpoint for sharing the stored data with interested parties.

[1152] 8. User Operations

[1153] The user (researcher) calibrates the eye tracking device and emotion engine, then begins observing through binoculars or a microscope. Based on feedback from the system, the user makes adjustments to improve the quality of their observations.

[1154] Specific examples

[1155] Example 1:

[1156] A researcher is observing a particular plant through binoculars. An eye-tracking device collects gaze data in real time, and an emotion engine collects the researcher's emotion data. The server synchronizes and analyzes this data, and verbalizes the observation in natural language. The analysis results reflect the gaze data and the corresponding emotion of the researcher.

[1157] Example 2:

[1158] Another researcher uses a microscope to observe the cells. The eye-tracking device collects gaze data, and the emotion engine collects emotion data. The server analyzes this data in real time and verbalizes the observations. The researcher's emotions upon discovering abnormal cells are recorded along with the gaze data.

[1159] Example prompts for generative AI models

[1160] "A user is observing a particular plant through binoculars. Please provide a detailed description of the observation in natural language based on the viewpoint and emotion data."

[1161] This system collects, analyzes, and verbalizes viewpoint and emotion data in real time, and provides feedback to researchers, enabling them to gain deeper insights based on the observations. Furthermore, by storing the data over the long term and sharing it with relevant parties, it is possible to improve the quality of research and education.

[1162] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1163] Step 1:

[1164] The terminal attaches an eye-tracking device to binoculars or microscopes and collects researchers' gaze data in real time.

[1165] Specific operation: The position of the researcher's pupils is tracked and the gaze coordinates are acquired. The collected gaze data is acquired at several tens of frames per second and temporarily stored in the device's internal memory.

[1166] Input: Researcher perspective.

[1167] Output: Real-time viewpoint data (coordinate information).

[1168] Step 2:

[1169] The device utilizes an emotion engine to recognize researchers' emotional data in real time.

[1170] How it works: It captures the researcher's facial expressions with a built-in camera and runs an expression analysis algorithm, while also obtaining data from a heart rate sensor to determine the researcher's emotional state.

[1171] Input: Researcher's facial expression and heart rate.

[1172] Output: Emotion data (emotional state identification information).

[1173] Step 3:

[1174] The device periodically sends the collected viewpoint data and emotion data to the server using HTTP POST requests.

[1175] Specific operation: Viewpoint data and emotion data are batched into packets and sent as an HTTP POST request.

[1176] Input: gaze data and emotion data.

[1177] Output: Data packets sent to the server.

[1178] Step 4:

[1179] The server temporarily stores the received viewpoint data and emotion data in a database.

[1180] Specific operation: Extract viewpoint data and emotion data from the HTTP request payload and insert them into a database table for temporary storage.

[1181] Input: Viewpoint and emotion data included in the HTTP request.

[1182] Output: Insertion of data into temporary storage database.

[1183] Step 5:

[1184] The server synchronizes and analyzes the viewpoint data and emotion data, and then verbalizes the analysis results in natural language using a generative AI model.

[1185] How it works: The viewpoint data and emotion data are linked over time, and a data analysis algorithm is run to identify the researcher's observations. The analyzed data is then fed into a generative AI model, which generates a natural language description.

[1186] Input: gaze data and emotion data.

[1187] Output: Analysis results written in natural language.

[1188] Step 6:

[1189] The terminal displays the analysis results sent from the server to the researcher in real time.

[1190] Specific operation: Receives analysis results from the server and displays the results on the screen.

[1191] Input: Parsed results from the server.

[1192] Output: Display analysis results on the terminal display.

[1193] Step 7:

[1194] The server stores the text data, viewpoint data, and emotion data of the analysis results in a permanent database.

[1195] Specific operation: The data containing the analysis results is inserted into a database for permanent storage. When saved, metadata such as the date and time and the researcher's ID are also included.

[1196] Input: Analysis result text data, viewpoint data, emotion data, and metadata.

[1197] Output: Saving data to a permanent database.

[1198] Step 8:

[1199] The server provides an endpoint for sharing the stored data with interested parties.

[1200] Specific behavior: Provide an API endpoint for accessing data so that relevant parties can obtain the data they need.

[1201] Input: Data access request from interested party.

[1202] Output: Providing data as requested.

[1203] Step 9:

[1204] The user (researcher) calibrates the eye tracking device and emotion engine, then begins observation through binoculars or a microscope.

[1205] Specific operation: Start the calibration procedure from the device settings screen, and perform viewpoint calibration and emotion recognition calibration. Once the settings are complete, begin observing through binoculars or a microscope.

[1206] Input: Data required for the calibration procedure (gaze information, emotion information).

[1207] Output: Accurate gaze and emotion data acquisition system after calibration is completed.

[1208] (Application example 2)

[1209] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1210] Modern factories are seeking to improve work efficiency and safety by using worker viewpoint and emotion data. However, there is no system that can analyze this data in real time and provide verbalized feedback, making it difficult to accurately grasp worker trends and improve work quality. Furthermore, there is a lack of feedback using analysis results that take workers' emotions into account, making it difficult to reduce work stress and improve safety.

[1211] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for acquiring viewpoint data and emotion data in real time, means for analyzing and verbalizing the viewpoint data and emotion data, and means for displaying the analysis results in real time. This makes it possible to accurately grasp the trends of workers, improve the quality of work, reduce work stress, and improve safety.

[1212] An "eye tracking device" is a device that tracks the movement of a user's eyes and acquires the data as gaze point data.

[1213] The "Emotion Engine" is a system that integrates sensors and software to analyze the emotional state of workers.

[1214] "Viewpoint data" is information that indicates the direction in which the worker is looking.

[1215] "Emotion data" is information that indicates the emotional state of the worker, and includes data on heart rate and facial expression analysis.

[1216] A "server" is a computer system that receives, analyzes, stores, and shares data over a network.

[1217] "Verbalization" is the process of describing analyzed data in natural language.

[1218] "Real-time display" is a function that instantly displays analyzed data to the user.

[1219] "Calibration" is the process of adjusting the accuracy of a device to ensure proper operation.

[1220] "Storage" refers to the act of recording data in a state that allows it to be referenced later.

[1221] "Sharing" is the act of making stored data available for handling with other users or systems.

[1222] The system that realizes this invention acquires, analyzes, and verbalizes the worker's viewpoint data and emotion data in real time, and provides feedback. The following describes the program content of this system, the hardware and software used, and specific examples.

[1223] First, a device (smartphone or head-mounted display) equipped with an eye-tracking device and emotion engine is used. These devices acquire the worker's gaze data and emotion data and send them to a server using an HTTP POST request. The server temporarily stores the received data in a database and begins analysis processing.

[1224] The server then uses the following software and modules to analyze the gaze data and emotion data:

[1225] Emotion Analysis Module: Analyzes emotion data and identifies the emotional state of the worker.

[1226] Viewpoint Analysis module: Analyzes viewpoint data and identifies worker eye movements.

[1227] Generative AI models (such as GPT-4): Generate analysis results in natural language based on viewpoint data and emotion data.

[1228] The server sends the analysis results to the device in real time and provides feedback, allowing workers to immediately receive advice on how to improve their work efficiency and safety.

[1229] Specific examples

[1230] For example, consider a situation in which a robot monitors the movements of workers in a factory. When a worker encounters a problem in a particular process, eye-gaze data and emotional data are immediately collected. The server analyzes this data and identifies the area in which the worker is feeling stressed. It then sends appropriate feedback to the worker's device in real time.

[1231] Prompt Sentence Examples

[1232] "Gaze data: 'The worker is looking at the screw point on the left side of the machine,' 'The worker is looking at the panel on the right side of the machine,' Emotion data: 'The worker's heart rate increases,' 'The worker's brow is furrowed,' Generate analysis results in natural language based on the relationship between the observation and the emotion."

[1233] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1234] Step 1:

[1235] The terminal (smartphone or head-mounted display) acquires the worker's viewpoint data and emotion data.

[1236] Specifically, the eye tracking device tracks the worker's gaze movements, and the emotion engine analyzes the worker's heart rate and facial expressions.

[1237] Input: Worker's viewpoint and emotion data

[1238] Output: Obtained gaze data and emotion data

[1239] Step 2:

[1240] The viewpoint data and emotion data acquired by the device are sent to the server via an HTTP POST request.

[1241] Specifically, the terminal converts the data into an appropriate format and transmits it to the specified endpoint.

[1242] Input: Obtained gaze data and emotion data

[1243] Output: A request to send data to the server

[1244] Step 3:

[1245] The server temporarily stores the received viewpoint data and emotion data in a database.

[1246] Specifically, the server checks the received data and inserts it into the database.

[1247] Input: Viewpoint data and emotion data sent from the device

[1248] Output: Viewpoint data and emotion data stored in a database

[1249] Step 4:

[1250] The server analyzes the viewpoint data and emotion data.

[1251] Specifically, the server performs data analysis using an Emotion Analysis module and a Viewpoint Analysis module.

[1252] Input: Viewpoint data and emotion data stored in a database

[1253] Output: Parsed gaze data and emotion data

[1254] Step 5:

[1255] Based on the analyzed data, the server uses a generative AI model (such as GPT-4) to translate the data into natural language.

[1256] Specifically, the server sends a prompt to GPT-4 and converts the analysis results into natural language.

[1257] Input: Parsed gaze data and emotion data

[1258] Output: Analysis results expressed in natural language

[1259] Step 6:

[1260] The server transmits the verbalized analysis results to the terminal.

[1261] As a specific operation, the server makes a request to transmit the generated natural language data to the terminal.

[1262] Input: Analysis results expressed in natural language

[1263] Output: Request to send analysis results to the device

[1264] Step 7:

[1265] The analysis results received by the device are displayed in real time.

[1266] Specifically, the device renders the analysis results in a display area, allowing the worker to see immediate feedback.

[1267] Input: Analysis results sent from the server in natural language

[1268] Output: Displaying the analysis results to the user

[1269] Through the above steps, this system is able to analyze the worker's viewpoint data and emotion data in real time and provide appropriate feedback.

[1270] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1271] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1272] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1273] [Fourth embodiment]

[1274] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1275] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1276] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1277] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1278] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1279] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1280] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1281] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1282] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1283] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1284] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1285] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1286] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1287] This invention is a system that uses an AI-equipped eye-tracking device to collect viewpoint data observed by researchers through binoculars or microscopes, analyzes the data, and verbalizes it. Furthermore, by storing and sharing this data, it aims to improve the quality of research and education.

[1288] System configuration

[1289] 1. Eye tracking device:

[1290] Terminals: Eye-tracking devices are attached to binoculars and microscopes to track the researcher's gaze.

[1291] Terminal: The device has the ability to capture eye movement and retrieve that data in real time.

[1292] 2. Transmitting viewpoint data:

[1293] Terminal: The acquired viewpoint data is sent to the server periodically to ensure that the viewpoint data reaches the server promptly.

[1294] 3. Saving viewpoint data:

[1295] Server: The server temporarily stores the viewpoint data it receives. A backup function may also be implemented to ensure that data is not lost.

[1296] 4. Analysis and verbalization of viewpoint data:

[1297] Server: Uses a generative AI to analyze the viewpoint data. The AI ​​processes the viewpoint data and verbalizes the observations in natural language.

[1298] Server: The verbalized data includes details about the points the researcher was observing and what they were looking at.

[1299] 5. Displaying the analysis results:

[1300] Terminal: The analysis results sent from the server are displayed to the researcher in real time, allowing the researcher to immediately check their observations.

[1301] 6. Data Storage and Sharing:

[1302] Server: Permanently stores verbalized text data and viewpoint data, which are stored for later reuse and analysis.

[1303] Server: The stored data is provided with an endpoint for sharing with specific stakeholders and educational personnel.

[1304] 7. User Actions:

[1305] User (researcher): The researcher calibrates the eye tracking device and then begins their research through binoculars and a microscope.

[1306] User (researcher): Researchers can use feedback from the system to make adjustments to improve the quality of their observations.

[1307] Specific examples

[1308] Example 1: A researcher is observing a particular plant through binoculars. During this observation, an eye-tracking device captures the researcher's gaze data in real time and sends it to a server. The server analyzes the gaze data and verbalizes in text what the researcher was focusing on. This allows the researcher to easily save their observation records and share them with other researchers.

[1309] Example 2: Another researcher uses a microscope to observe cells. The eye-tracking device acquires gaze data and sends it to the server. The server then analyzes and verbalizes the observation in real time based on the gaze data. As a result, important observation points, such as abnormal cells, are quickly identified, and the analysis results are fed back to the researcher in real time. The researcher can immediately check the observation results and continue with more detailed analysis.

[1310] Thus, the present invention provides a system that significantly improves the efficiency of recording, analyzing, and sharing observation data in research settings.

[1311] The processing flow will be explained below.

[1312] Step 1:

[1313] Device: Initial setup of the eye tracking device, attaching it to the binoculars or microscope and calibrating it, ensuring that the device accurately tracks the researcher's gaze.

[1314] Step 2:

[1315] User: The researcher begins using binoculars or a microscope. After confirming that the object is in a state where detailed observation is possible, the researcher begins observation.

[1316] Step 3:

[1317] Device: Once observation begins, the eye-tracking device acquires gaze data in real time. The device captures gaze data every 0.1 seconds and buffers the information.

[1318] Step 4:

[1319] Terminal: The acquired viewpoint data is sent to the server at a predetermined interval (for example, every second) using an HTTP POST request.

[1320] Step 5:

[1321] Server: The server receives the viewpoint data, adds it to a queue for analysis, and temporarily stores it in a database.

[1322] Step 6:

[1323] Server: Analyzes the received viewpoint data. Based on the viewpoint data, the generation AI describes the objects and observations the researcher was looking at in natural language.

[1324] Step 7:

[1325] Server: Generates the analyzed data in text format, verbalizing in detail what the researchers were observing.

[1326] Step 8:

[1327] Server: Stores the verbalized data and corresponding viewpoint data in a permanent database, including metadata such as date and time and researcher ID.

[1328] Step 9:

[1329] Server: Sends analysis results and verbalized data back to the device in real time, allowing researchers to instantly check their observations.

[1330] Step 10:

[1331] Terminal: The analysis results sent back from the server are displayed in real time. Researchers can manage the progress of their observations based on the displayed information.

[1332] Step 11:

[1333] User: Researchers can view the displayed analysis results, record them if necessary, and generate a link to share the results with other researchers and educators.

[1334] Step 12:

[1335] Server: Generates a shared link for the data and makes it accessible to interested parties. The shared link contains credentials to access the specific dataset.

[1336] Step 13:

[1337] Users: Researchers share their observations by sending a shared link to stakeholders, making the data available for collaborative research and educational use.

[1338] Example 1

[1339] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1340] Conventional eye tracking systems only collect gaze data, but lack a mechanism for analyzing that data and providing immediate feedback to researchers. As a result, analyzing and sharing the data requires a great deal of time and effort, reducing the efficiency of research and education. In addition, gaze data analysis is often done manually, which poses problems with accuracy and speed.

[1341] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1342] In this invention, the server includes a means for using a generating AI to analyze viewpoint data, a means for displaying the viewpoint data acquired in real time as an analysis result, and a means for saving and sharing the viewpoint data and its analysis result, thereby enabling fast and accurate analysis and feedback of viewpoint data.

[1343] An "eye tracking device" is a device that detects a user's gaze and records that gaze movement as data.

[1344] "Gaze point data" refers to data that includes information such as the position and movement of the user's gaze, and the duration of gaze.

[1345] A "server" is a computer system that stores, analyzes, and sends and receives data over a network.

[1346] "Generative AI" is an artificial intelligence technology that analyzes and predicts based on large amounts of data.

[1347] "Calibration" is the initial adjustment operation that eye tracking devices perform to obtain accurate gaze data.

[1348] "Analysis" refers to processing acquired data using statistical and logical methods to clarify its meaning.

[1349] "Verbalization" is the conversion of data and information into natural language, i.e., words that humans can understand.

[1350] "Real-time" refers to a situation that corresponds to the current time and is reflected immediately with little delay.

[1351] "Storage" means recording acquired data in a computer's storage device so that it can be reused later.

[1352] "Sharing" means linking acquired data and analysis results with other users and systems and making them accessible.

[1353] This system collects and analyzes viewpoint data observed by researchers through binoculars or microscopes, and provides feedback on the results, thereby improving the efficiency of research and education.

[1354] Hardware and software used

[1355] Eye tracking device: A device that tracks a user's gaze as they observe using binoculars or a microscope.

[1356] Server: A computer system that stores viewpoint data and analyzes it using generative AI.

[1357] Generative AI: An artificial intelligence technology that analyzes viewpoint data and verbalizes it in natural language.

[1358] Examples of data processing and data calculation

[1359] User: The researcher attaches the eye tracking device to a pair of binoculars or a microscope and calibrates the device before beginning observations. Once calibrated, the device tracks the researcher's gaze and collects gaze data in real time.

[1360] Terminal: The eye tracking device transmits the gaze data acquired in real time to the server, where it is stored.

[1361] Server: The viewpoint data stored on the server is analyzed using a generative AI model. The analysis of viewpoint data involves the generative AI verbalizing the observations in natural language based on the viewpoint position, movement, attention time, etc.

[1362] Terminal: The verbalized analysis results are displayed in real time on the user's terminal, allowing researchers to immediately check the observation results.

[1363] Server: The text data and viewpoint data generated as a result of the analysis are stored for later reuse and analysis. In addition, an endpoint is provided for sharing the stored data with specific stakeholders and educational institutions.

[1364] Specific examples of operation

[1365] Example 1: When a researcher is observing a particular plant through binoculars, the eye-tracking device captures the researcher's gaze data in real time and sends it to the server. The server then analyzes the gaze data and verbalizes it in natural language, such as "The researcher focused on a specific part of the leaf for three minutes." The researcher can instantly view this information, easily save the observation record, and share it with other researchers.

[1366] Example 2: When another researcher observes cells under a microscope, the eye-tracking device captures gaze data and sends it to the server. The generative AI analyzes the gaze data to identify important observation points, such as abnormal cells, and generates a statement in natural language, such as "The researcher observed that a specific cell was abnormal." This information is displayed in real time on the researcher's device, allowing them to immediately check the observation results.

[1367] Prompt Sentence Examples

[1368] "Analyze the viewpoint data of the plant you observed through the binoculars and explain in detail which part you focused on."

[1369] "Identify the key points of the cells you observed under the microscope and describe them in words."

[1370] In this way, the system of the present invention contributes to improving the quality of research and education through rapid and accurate analysis and collaboration of viewpoint data.

[1371] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1372] Step 1:

[1373] Calibrating your eye tracking device

[1374] User: The researcher attaches the eye tracking device to a pair of binoculars or a microscope and calibrates it.

[1375] Input: Researcher's gaze position

[1376] Data processing: Researchers direct their gaze to multiple designated points, allowing the device to adjust for accurate gaze tracking.

[1377] Output: Calibration complete message

[1378] Specific actions: The researcher follows the instructions on the screen and completes the calibration by fixing their gaze on a target point in front of them for a few seconds.

[1379] Step 2:

[1380] Obtaining viewpoint data

[1381] Device: An eye-tracking device captures the researcher's gaze movements in real time.

[1382] Input: Researcher's gaze position and movement

[1383] Data processing: Sensors within the device record eye movements and convert them into gaze data.

[1384] Output: Obtained viewpoint data

[1385] How it works: As researchers observe through binoculars or a microscope, the device tracks their changes in gaze.

[1386] Step 3:

[1387] Viewpoint data transmission

[1388] Terminal: Sends the acquired viewpoint data to the server.

[1389] Input: Viewpoint data

[1390] Data processing: The viewpoint data is formatted into an analyzable format and sent to the server via the network.

[1391] Output: Viewpoint data sent to the server

[1392] Specific operation: The device packets viewpoint data and periodically sends it to the server via the network.

[1393] Step 4:

[1394] Saving viewpoint data

[1395] Server: Temporarily stores the received viewpoint data.

[1396] Input: Viewpoint data sent from the device

[1397] Data processing: The data is saved to a database and a backup is created at the same time.

[1398] Output: Saved viewpoint data

[1399] Specific operations: The server verifies the received data, records it in the database, and copies it to the backup system.

[1400] Step 5:

[1401] Analysis and verbalization of viewpoint data

[1402] Server: Analyzes viewpoint data and verbalizes it into natural language.

[1403] Input: Saved viewpoint data

[1404] Data processing: A generative AI model analyzes viewpoint data and expresses observations in natural language.

[1405] Output: Text data translated into natural language

[1406] Specific operation: The AI ​​model in the server analyzes the viewpoint data and generates text such as "The researcher focused on a specific part of the leaf for three minutes."

[1407] Step 6:

[1408] Displaying analysis results

[1409] Terminal: Displays the analysis results to the user in real time.

[1410] Input: Text data verbalized in natural language

[1411] Data processing: Convert into a display format and display on the user interface.

[1412] Output: Displayed analysis results

[1413] Specific operation: The analysis results pop up in real time on the user's device and the researcher can check them.

[1414] Step 7:

[1415] Data storage and sharing

[1416] Server: Stores and shares verbalized text data and viewpoint data.

[1417] Input: Text data expressed in natural language, viewpoint data

[1418] Data processing: Set access permissions for stored data and share it via API or portal.

[1419] Output: Saved data and its share link

[1420] Specific behavior: Record data in a database and set sharing permissions to allow access to specific parties.

[1421] The above are the processing steps of the system.

[1422] (Application example 1)

[1423] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1424] Maintenance of robots and machines is extremely important in factories and manufacturing sites, but the work is labor-intensive and time-consuming. Furthermore, collecting and analyzing viewpoint data and identifying problem areas during maintenance work are often difficult. This creates a need for improved efficiency and accuracy in maintenance work. Currently, it is difficult to identify abnormal areas and take prompt action, resulting in issues such as reduced productivity and increased operating costs.

[1425] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1426] In this invention, the server includes means for acquiring gaze data of the worker using an eye tracking device, means for analyzing and verbalizing the acquired gaze data, and means for analyzing the gaze data during maintenance work and identifying problem areas, thereby enabling accurate collection, analysis, and real-time feedback of gaze data during maintenance work.

[1427] An "eye tracking device" is a device for acquiring data on a worker's gaze.

[1428] "Gaze point data" is data that indicates the position and movement of the worker's line of sight.

[1429] A "server" is a computer system that receives, analyzes, and stores viewpoint data.

[1430] "Analysis" is the process of extracting information based on viewpoint data and verbalizing it.

[1431] "Verbalization" means converting the analyzed data into natural language.

[1432] "Maintenance work" refers to maintenance work on robots and machines in factories and manufacturing sites.

[1433] "Problem area" refers to an area that is abnormal or requires repair during maintenance work.

[1434] "Feedback" is the process of immediately notifying the operator of the analysis results.

[1435] MODE FOR CARRYING OUT THE INVENTION

[1436] This invention is a system for improving the efficiency and accuracy of maintenance work for robots and machines in factories and manufacturing sites. This system uses an eye-tracking device to acquire worker gaze data, analyzes and verbalizes the data on a server, and identifies problem areas. Specific embodiments of this system are described below.

[1437] Hardware and software used

[1438] Eye tracking device: A device for acquiring worker gaze data in real time (e.g., Tobii Eye Tracker).

[1439] Server: A computer system that receives, analyzes, and stores viewpoint data.

[1440] Smartphone or head-mounted display: A display device that allows workers to receive real-time feedback.

[1441] Eye tracking library: A software library used to collect gaze data.

[1442] Requests library: A Python library used to communicate with the server.

[1443] Generative AI (e.g., OpenAI GPT-3): An AI model for analyzing and verbalizing viewpoint data.

[1444] Process and Data Flow

[1445] 1. Obtaining viewpoint data

[1446] The eye tracking device collects the worker's gaze data in real time, which is then processed by software (Eye tracking library) within the device.

[1447] 2. Data transmission

[1448] The viewpoint data acquired by the device is periodically sent to the server using the Requests library, and the server receives and temporarily stores it.

[1449] 3. Data analysis and verbalization

[1450] The server passes the received viewpoint data to the generation AI, which analyzes the data and translates it into natural language. The generation AI then identifies abnormalities and points of interest based on the analysis of the viewpoint data.

[1451] 4. Feedback and Representation

[1452] The analysis results are displayed in real time on the worker's smartphone or head-mounted display, allowing the worker to immediately identify any abnormalities and take appropriate measures.

[1453] 5. Data storage and sharing

[1454] The verbalized text data and the original viewpoint data are stored on a server, where they can be reused and further analyzed later and shared with stakeholders.

[1455] Specific examples

[1456] Consider a scenario in which a maintenance worker in a factory is inspecting the joints of a robot with a magnifying glass. An eye-tracking device tracks the worker's gaze and sends the gaze data to a server. A generative AI analyzes this data and generates an analysis result, such as "the bolts in the joint are loose." The result is immediately displayed on the worker's smartphone, allowing the worker to quickly identify the problem area and take the necessary measures.

[1457] Prompt Sentence Examples

[1458] An example of a prompt to be input to a generative AI model would be something like this:

[1459] Verbalize the following observations from your user-perspective data:

[1460] Viewpoint Data:

[1461] Coordinates: (x1, y1), time: t1

[1462] Coordinates: (x2, y2), time: t2

[1463] ...

[1464] Coordinates: (xn, yn), Time: tn

[1465] Post-verbalization:

[1466] In this way, a system is constructed that improves the efficiency and accuracy of maintenance work for robots and machines in factories and manufacturing sites.

[1467] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1468] Step 1:

[1469] The user (worker) wears an eye-tracking device and begins maintenance work. The eye-tracking device acquires the worker's gaze data (position and movement of the gaze point) in real time. The input here is the worker's gaze point, and the output is the acquired raw gaze data.

[1470] Step 2:

[1471] The device (eye tracking device) processes the acquired gaze data in real time using the built-in Eye tracking library. The processed gaze data is temporarily buffered within the device. The input is the acquired gaze data, and the output is the processed gaze data.

[1472] Step 3:

[1473] The device periodically sends the processed viewpoint data to the server using the Requests library. An internet connection is required for transmission, and the data is packetized in JSON format. The input is the processed viewpoint data, and the output is the transmitted data packet.

[1474] Step 4:

[1475] The server receives the viewpoint data sent from the terminal and temporarily stores it in a database. The server checks the integrity of the data and verifies that there are no inconsistencies. The input is the sent data packet, and the output is the temporarily stored viewpoint data.

[1476] Step 5:

[1477] The server passes the saved viewpoint data to the generative AI model for analysis and verbalization. Based on the prompt, the model identifies abnormalities and points of interest from the viewpoint data and explains them in natural language. The input is the saved viewpoint data, and the output is the verbalized analysis results.

[1478] Step 6:

[1479] The server sends the generated analysis results in real time to the worker's smartphone or head-mounted display. The transmission is via an internet connection, and the data is displayed to the worker in real time. The input is the verbalized analysis results, and the output is the displayed analysis results.

[1480] Step 7:

[1481] The user (operator) checks the analysis results displayed in real time and takes the necessary measures promptly. The operator identifies the problem area and performs specific maintenance work based on that. The input is the displayed analysis results, and the output is the maintenance work that has been carried out.

[1482] Step 8:

[1483] The server permanently stores the verbalized text data and the original viewpoint data, allowing for later reuse and further analysis. The stored data may also be shared among stakeholders. The input is the verbalized analysis results and the original viewpoint data, and the output is the stored data.

[1484] In this way, through the specific operations, inputs and outputs at each step, maintenance work at factories and manufacturing sites can be made more efficient and more accurate.

[1485] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1486] This invention is a system that uses an AI-equipped eye-tracking device and emotion engine to collect viewpoint data observed by researchers through binoculars or microscopes, analyzes the data, and verbalizes it. Furthermore, by recognizing the user's emotions and adjusting the data analysis results based on that information, it aims to improve the quality of research and education.

[1487] System configuration

[1488] 1. Eye tracking device:

[1489] Device: Eye-tracking devices are attached to binoculars or microscopes to track the researcher's gaze, allowing for real-time acquisition of researcher gaze data.

[1490] 2. Emotion Engine:

[1491] Device: Equipped with an emotion engine, it recognizes the emotions felt by researchers during observation. It may use facial expression analysis and heart rate sensors.

[1492] 3. Transmitting viewpoint data:

[1493] Device: The acquired viewpoint data and emotion data are periodically sent to the server using HTTP POST requests.

[1494] 4. Temporary data storage:

[1495] Server: The server receives the viewpoint data and emotion data and temporarily stores them in a database.

[1496] 5. Data analysis and verbalization:

[1497] Server: Synchronizes and analyzes gaze data and emotion data, adjusting the analysis content based on what the researcher was looking at and changes in emotion.

[1498] Server: The generative AI verbalizes the observations in natural language based on viewpoint data and emotion data.

[1499] 6. Displaying the analysis results:

[1500] Terminal: The analysis results sent from the server are displayed to the researcher in real time, allowing the researcher to instantly check their observations and receive emotional feedback.

[1501] 7. Data Storage and Sharing:

[1502] Server: Stores the verbalized text data, viewpoint data, and emotion data in a permanent database. Data storage also includes metadata such as date and time and researcher ID.

[1503] Server: The stored data is provided with an endpoint for sharing with specific stakeholders and educational personnel.

[1504] 8. User Actions:

[1505] User (researcher): The researcher calibrates the eye tracking device and emotion engine, then begins the study through binoculars and a microscope.

[1506] User (researcher): Researchers can use feedback from the system to make adjustments to improve the quality of their observations.

[1507] Specific examples

[1508] Example 1: A researcher is observing a particular plant through binoculars. An eye-tracking device captures the researcher's gaze data in real time, and an emotion engine captures the researcher's emotion data. The server synchronizes these data and analyzes and verbalizes the observations. The analysis results reflect the gaze data and the corresponding researcher's emotions, allowing the researcher to conduct further analysis based on the emotional changes at specific viewpoints.

[1509] Example 2: Another researcher uses a microscope to observe cells. The eye-tracking device collects gaze data, and the emotion engine collects emotion data. The server analyzes this data and verbalizes the observations in real time. The data includes the researcher's emotional reactions when he or she discovers abnormal cells, and emotional changes such as stress and surprise are recorded along with the gaze data.

[1510] In this way, the present invention provides a system that significantly improves the efficiency of synchronously recording, analyzing, and sharing observation data and emotion data in research settings. Researchers can gain deeper insights by receiving feedback based on viewpoint and emotion data.

[1511] The processing flow will be explained below.

[1512] Step 1:

[1513] Device: Initial setup of the eye tracking device and emotion engine. Attach the device to the binoculars or microscope and calibrate it so that the device can accurately track the researcher's gaze and emotions.

[1514] Step 2:

[1515] User: The researcher begins using binoculars or a microscope. After confirming that the object is in a state where detailed observation is possible, the researcher begins observation.

[1516] Step 3:

[1517] Device: When observation begins, the eye-tracking device acquires gaze data in real time, and the emotion engine captures the user's emotional data (e.g., facial expressions and heart rate). The device captures data every 0.1 seconds and buffers the information.

[1518] Step 4:

[1519] Device: The acquired viewpoint data and emotion data are sent to the server at a predetermined interval (for example, every second). The data is sent using an HTTP POST request.

[1520] Step 5:

[1521] Server: The server receives the gaze data and emotion data, adds the received data to a queue for analysis, and temporarily stores it in a database.

[1522] Step 6:

[1523] Server: Analyzes the received viewpoint and emotion data. The generation AI uses the viewpoint data to describe the objects and observations the researcher was looking at in natural language, and adjusts the analysis results based on the emotion data.

[1524] Step 7:

[1525] Server: Generates verbalized text based on the viewpoint data and emotion data, and generates analysis results that include details of observations and emotional changes.

[1526] Step 8:

[1527] Server: The analysis results and verbalized data are sent from the server to the terminal and fed back to the researcher in real time.

[1528] Step 9:

[1529] Terminal: Displays the analysis results sent from the server in real time. Researchers can manage the progress of their observations based on the displayed information.

[1530] Step 10:

[1531] Server: Stores the data in a permanent database, including viewpoint data, emotion data, analysis results, and metadata such as date and time and researcher ID.

[1532] Step 11:

[1533] Server: Generates shared links based on stored data and provides endpoints for stakeholders to access them.

[1534] Step 12:

[1535] User: Researchers review the analysis results, record them as needed, and generate a link to share the results with other researchers and educators.

[1536] Step 13:

[1537] Users: Researchers can send a shared link to stakeholders to share the dataset containing observations and emotion data, allowing them to use the data for collaborative research and educational purposes.

[1538] Example 2

[1539] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1540] When researchers conduct observations using binoculars or microscopes, the current technological environment lacks the means to efficiently acquire and analyze viewpoint and emotion data, and then verbalize and present the results. Furthermore, systems for appropriately sharing acquired data with stakeholders and improving the quality of research results are limited. This calls for effective methods to improve the quality of research and education.

[1541] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1542] In this invention, the server includes a means for acquiring researcher's gaze data using an eye-tracking device, a means for acquiring researcher's emotion data using an emotion engine, a means for transmitting the acquired gaze data and emotion data to the server, a means for analyzing the gaze data and emotion data in the server and verbalizing the analysis results using a generative AI model, a means for saving the analyzed verbalized data, gaze data, and emotion data, and a means for sharing the saved data with relevant parties. This allows researchers to analyze the gaze data and emotion data obtained during observation in real time and provide the results as immediate feedback. Furthermore, by storing the data long-term and sharing it with various relevant parties, the quality of research and education can be improved.

[1543] An "eye tracking device" is a device that collects gaze data in real time when researchers observe through binoculars or a microscope.

[1544] The "Emotion Engine" is a system that recognizes the emotional state of researchers in real time, using facial expression analysis and heart rate sensors.

[1545] "Gaze data" is information that records the position and movement of the researcher's gaze, and is data that indicates what they are observing.

[1546] "Emotional data" refers to data on the researcher's emotional state, and is information obtained based on facial expression analysis, heart rate, etc.

[1547] A "generative AI model" is an artificial intelligence model that generates explanatory text and analysis results in natural language based on input data.

[1548] "Calibration" is a method of adjusting the performance of the eye tracking device and emotion engine to enable accurate data acquisition.

[1549] "Analysis results" are the results of analysis based on viewpoint data and emotion data, and include content verbalized through a generative AI model.

[1550] "Real-time display means" refers to devices or software that instantly display acquired and analyzed data and provide feedback to the user.

[1551] A "permanent database" is a database for long-term storage of data, allowing it to be accessed and analyzed at a later date.

[1552] "Stakeholders" refers to researchers, educators, and other relevant individuals and organizations.

[1553] This system allows researchers to efficiently collect viewpoint and emotion data during observations using binoculars and microscopes, and then analyze, verbalize, and provide feedback. This system can improve the quality of research and education.

[1554] 1. Use of eye tracking devices

[1555] The device is attached to a binocular or microscope and collects the researcher's gaze data in real time. The device tracks the position of the researcher's pupils and obtains gaze coordinates. The gaze data is acquired at several tens of frames per second and temporarily stored in the device's internal memory.

[1556] 2. Use of Emotion Engine

[1557] The device utilizes an emotion engine to recognize researchers' emotions in real time. Specifically, it captures the researchers' facial expressions with its built-in camera and runs an expression analysis algorithm. It also obtains data from a heart rate sensor to determine their emotional state.

[1558] 3. Transmission of viewpoint and emotion data

[1559] The device periodically sends the collected viewpoint data and emotion data to the server using HTTP POST requests. When sending, the viewpoint data and emotion data are packaged into packets in batch format.

[1560] 4. Temporary storage of data

[1561] The server temporarily stores the received viewpoint data and emotion data in a database. The server extracts the viewpoint data and emotion data from the payload of the HTTP request and inserts them into a database table for temporary storage.

[1562] 5. Data analysis and verbalization

[1563] The server synchronizes and analyzes the viewpoint data and emotion data, and then uses a generative AI model to verbalize the analysis results in natural language. The server links the viewpoint data and emotion data along a timeline, runs a data analysis algorithm, and identifies the researcher's observations. The server then inputs the analyzed data into the generative AI model, generating a natural language explanation.

[1564] 6. Displaying the analysis results

[1565] The terminal displays the analysis results sent from the server to the researcher in real time. The terminal receives the analysis results from the server, displays the received results on a display, and provides feedback to the researcher.

[1566] 7. Data Storage and Sharing

[1567] The server stores the resulting text, viewpoint, and emotion data in a permanent database, along with metadata such as the date and time and the researcher's ID. The server also provides an endpoint for sharing the stored data with interested parties.

[1568] 8. User Operations

[1569] The user (researcher) calibrates the eye tracking device and emotion engine, then begins observing through binoculars or a microscope. Based on feedback from the system, the user makes adjustments to improve the quality of their observations.

[1570] Specific examples

[1571] Example 1:

[1572] A researcher is observing a particular plant through binoculars. An eye-tracking device collects gaze data in real time, and an emotion engine collects the researcher's emotion data. The server synchronizes and analyzes this data, and verbalizes the observation in natural language. The analysis results reflect the gaze data and the corresponding emotion of the researcher.

[1573] Example 2:

[1574] Another researcher uses a microscope to observe the cells. The eye-tracking device collects gaze data, and the emotion engine collects emotion data. The server analyzes this data in real time and verbalizes the observations. The researcher's emotions upon discovering abnormal cells are recorded along with the gaze data.

[1575] Example prompts for generative AI models

[1576] "A user is observing a particular plant through binoculars. Please provide a detailed description of the observation in natural language based on the viewpoint and emotion data."

[1577] This system collects, analyzes, and verbalizes viewpoint and emotion data in real time, and provides feedback to researchers, enabling them to gain deeper insights based on the observations. Furthermore, by storing the data over the long term and sharing it with relevant parties, it is possible to improve the quality of research and education.

[1578] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1579] Step 1:

[1580] The terminal attaches an eye-tracking device to binoculars or microscopes and collects researchers' gaze data in real time.

[1581] Specific operation: The position of the researcher's pupils is tracked and the gaze coordinates are acquired. The collected gaze data is acquired at several tens of frames per second and temporarily stored in the device's internal memory.

[1582] Input: Researcher perspective.

[1583] Output: Real-time viewpoint data (coordinate information).

[1584] Step 2:

[1585] The device utilizes an emotion engine to recognize researchers' emotional data in real time.

[1586] How it works: It captures the researcher's facial expressions with a built-in camera and runs an expression analysis algorithm, while also obtaining data from a heart rate sensor to determine the researcher's emotional state.

[1587] Input: Researcher's facial expression and heart rate.

[1588] Output: Emotion data (emotional state identification information).

[1589] Step 3:

[1590] The device periodically sends the collected viewpoint data and emotion data to the server using HTTP POST requests.

[1591] Specific operation: Viewpoint data and emotion data are batched into packets and sent as an HTTP POST request.

[1592] Input: gaze data and emotion data.

[1593] Output: Data packets sent to the server.

[1594] Step 4:

[1595] The server temporarily stores the received viewpoint data and emotion data in a database.

[1596] Specific operation: Extract viewpoint data and emotion data from the HTTP request payload and insert them into a database table for temporary storage.

[1597] Input: Viewpoint and emotion data included in the HTTP request.

[1598] Output: Insertion of data into temporary storage database.

[1599] Step 5:

[1600] The server synchronizes and analyzes the viewpoint data and emotion data, and then verbalizes the analysis results in natural language using a generative AI model.

[1601] How it works: The viewpoint data and emotion data are linked over time, and a data analysis algorithm is run to identify the researcher's observations. The analyzed data is then fed into a generative AI model, which generates a natural language description.

[1602] Input: gaze data and emotion data.

[1603] Output: Analysis results written in natural language.

[1604] Step 6:

[1605] The terminal displays the analysis results sent from the server to the researcher in real time.

[1606] Specific operation: Receives analysis results from the server and displays the results on the screen.

[1607] Input: Parsed results from the server.

[1608] Output: Display analysis results on the terminal display.

[1609] Step 7:

[1610] The server stores the text data, viewpoint data, and emotion data of the analysis results in a permanent database.

[1611] Specific operation: The data containing the analysis results is inserted into a database for permanent storage. When saved, metadata such as the date and time and the researcher's ID are also included.

[1612] Input: Analysis result text data, viewpoint data, emotion data, and metadata.

[1613] Output: Saving data to a permanent database.

[1614] Step 8:

[1615] The server provides an endpoint for sharing the stored data with interested parties.

[1616] Specific behavior: Provide an API endpoint for accessing data so that relevant parties can obtain the data they need.

[1617] Input: Data access request from interested party.

[1618] Output: Providing data as requested.

[1619] Step 9:

[1620] The user (researcher) calibrates the eye tracking device and emotion engine, then begins observation through binoculars or a microscope.

[1621] Specific operation: Start the calibration procedure from the device settings screen, and perform viewpoint calibration and emotion recognition calibration. Once the settings are complete, begin observing through binoculars or a microscope.

[1622] Input: Data required for the calibration procedure (gaze information, emotion information).

[1623] Output: Accurate gaze and emotion data acquisition system after calibration is completed.

[1624] (Application example 2)

[1625] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1626] Modern factories are seeking to improve work efficiency and safety by using worker viewpoint and emotion data. However, there is no system that can analyze this data in real time and provide verbalized feedback, making it difficult to accurately grasp worker trends and improve work quality. Furthermore, there is a lack of feedback using analysis results that take workers' emotions into account, making it difficult to reduce work stress and improve safety.

[1627] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for acquiring viewpoint data and emotion data in real time, means for analyzing and verbalizing the viewpoint data and emotion data, and means for displaying the analysis results in real time. This makes it possible to accurately grasp the trends of workers, improve the quality of work, reduce work stress, and improve safety.

[1628] An "eye tracking device" is a device that tracks the movement of a user's eyes and acquires the data as gaze point data.

[1629] The "Emotion Engine" is a system that integrates sensors and software to analyze the emotional state of workers.

[1630] "Viewpoint data" is information that indicates the direction in which the worker is looking.

[1631] "Emotion data" is information that indicates the emotional state of the worker, and includes data on heart rate and facial expression analysis.

[1632] A "server" is a computer system that receives, analyzes, stores, and shares data over a network.

[1633] "Verbalization" is the process of describing analyzed data in natural language.

[1634] "Real-time display" is a function that instantly displays analyzed data to the user.

[1635] "Calibration" is the process of adjusting the accuracy of a device to ensure proper operation.

[1636] "Storage" refers to the act of recording data in a state that allows it to be referenced later.

[1637] "Sharing" is the act of making stored data available for handling with other users or systems.

[1638] The system that realizes this invention acquires, analyzes, and verbalizes the worker's viewpoint data and emotion data in real time, and provides feedback. The following describes the program content of this system, the hardware and software used, and specific examples.

[1639] First, a device (smartphone or head-mounted display) equipped with an eye-tracking device and emotion engine is used. These devices acquire the worker's gaze data and emotion data and send them to a server using an HTTP POST request. The server temporarily stores the received data in a database and begins analysis processing.

[1640] The server then uses the following software and modules to analyze the gaze data and emotion data:

[1641] Emotion Analysis Module: Analyzes emotion data and identifies the emotional state of the worker.

[1642] Viewpoint Analysis module: Analyzes viewpoint data and identifies worker eye movements.

[1643] Generative AI models (such as GPT-4): Generate analysis results in natural language based on viewpoint data and emotion data.

[1644] The server sends the analysis results to the terminal in real time and provides feedback, allowing workers to immediately receive advice on how to improve their work efficiency and safety.

[1645] Specific examples

[1646] For example, consider a situation in which a robot monitors the movements of workers in a factory. When a worker encounters a problem in a particular process, eye-gaze data and emotional data are immediately collected. The server analyzes this data and identifies the area in which the worker is feeling stressed. It then sends appropriate feedback to the worker's device in real time.

[1647] Prompt Sentence Examples

[1648] "Gaze data: 'The worker is looking at the screw point on the left side of the machine,' 'The worker is looking at the panel on the right side of the machine,' Emotion data: 'The worker's heart rate increases,' 'The worker's brow is furrowed,' Generate analysis results in natural language based on the relationship between the observation and the emotion."

[1649] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1650] Step 1:

[1651] The terminal (smartphone or head-mounted display) acquires the worker's viewpoint data and emotion data.

[1652] Specifically, the eye tracking device tracks the worker's gaze movements, and the emotion engine analyzes the worker's heart rate and facial expressions.

[1653] Input: Worker's viewpoint and emotion data

[1654] Output: Obtained gaze data and emotion data

[1655] Step 2:

[1656] The viewpoint data and emotion data acquired by the device are sent to the server via an HTTP POST request.

[1657] Specifically, the terminal converts the data into an appropriate format and transmits it to the specified endpoint.

[1658] Input: Obtained gaze data and emotion data

[1659] Output: A request to send data to the server

[1660] Step 3:

[1661] The server temporarily stores the received viewpoint data and emotion data in a database.

[1662] Specifically, the server checks the received data and inserts it into the database.

[1663] Input: Viewpoint data and emotion data sent from the device

[1664] Output: Viewpoint data and emotion data stored in a database

[1665] Step 4:

[1666] The server analyzes the viewpoint data and emotion data.

[1667] Specifically, the server performs data analysis using an Emotion Analysis module and a Viewpoint Analysis module.

[1668] Input: Viewpoint data and emotion data stored in a database

[1669] Output: Parsed gaze data and emotion data

[1670] Step 5:

[1671] Based on the analyzed data, the server uses a generative AI model (such as GPT-4) to translate the data into natural language.

[1672] Specifically, the server sends a prompt to GPT-4 and converts the analysis results into natural language.

[1673] Input: Parsed gaze data and emotion data

[1674] Output: Analysis results expressed in natural language

[1675] Step 6:

[1676] The server transmits the verbalized analysis results to the terminal.

[1677] As a specific operation, the server makes a request to transmit the generated natural language data to the terminal.

[1678] Input: Analysis results expressed in natural language

[1679] Output: Request to send analysis results to the device

[1680] Step 7:

[1681] The analysis results received by the device are displayed in real time.

[1682] Specifically, the device renders the analysis results in a display area, allowing the worker to see immediate feedback.

[1683] Input: Analysis results sent from the server in natural language

[1684] Output: Displaying the analysis results to the user

[1685] Through the above steps, this system is able to analyze the worker's viewpoint data and emotion data in real time and provide appropriate feedback.

[1686] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1687] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1688] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1689] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1690] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1691] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1692] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1693] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1694] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1695] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1696] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1697] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1698] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1699] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1700] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1701] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1702] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1703] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1704] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1705] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, in order to avoid confusion and to facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1706] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1707] The following is further disclosed regarding the above embodiment.

[1708] (Claim 1)

[1709] A means of acquiring researcher gaze data through an eye-tracking device;

[1710] means for transmitting the acquired viewpoint data to a server;

[1711] A means for analyzing and verbalizing viewpoint data in a server;

[1712] a means for storing the analyzed verbalization data and viewpoint data;

[1713] a means for sharing the stored data;

[1714] A system including:

[1715] (Claim 2)

[1716] 10. The system of claim 1, further comprising means for calibrating the eye tracking device.

[1717] (Claim 3)

[1718] 10. The system of claim 1, further comprising means for acquiring viewpoint data in real time and displaying analysis results in real time.

[1719] "Example 1"

[1720] (Claim 1)

[1721] A means of acquiring researcher gaze data through an eye-tracking device;

[1722] means for transmitting the acquired viewpoint data to a server;

[1723] A means for analyzing and verbalizing viewpoint data in a server;

[1724] a means for storing the analyzed verbalization data and viewpoint data;

[1725] a means for sharing the stored data;

[1726] A means of using generative AI to analyze viewpoint data; and

[1727] A means for displaying the viewpoint data acquired in real time as an analysis result;

[1728] A system including:

[1729] (Claim 2)

[1730] 10. The system of claim 1, further comprising means for calibrating the eye tracking device.

[1731] (Claim 3)

[1732] 10. The system of claim 1, further comprising means for acquiring viewpoint data in real time and displaying analysis results in real time.

[1733] "Application Example 1"

[1734] (Claim 1)

[1735] A means for acquiring gaze data of a worker by an eye tracking device;

[1736] means for transmitting the acquired viewpoint data to a server;

[1737] A means for analyzing and verbalizing viewpoint data in a server;

[1738] a means for storing the analyzed verbalization data and viewpoint data;

[1739] a means for sharing the stored data;

[1740] A means for analyzing viewpoint data and identifying problem areas during maintenance work;

[1741] A means of providing real-time feedback of analysis results to the operator,

[1742] A system including:

[1743] (Claim 2)

[1744] 10. The system of claim 1, further comprising means for calibrating the eye tracking device.

[1745] (Claim 3)

[1746] 10. The system of claim 1, further comprising means for acquiring viewpoint data in real time and displaying analysis results in real time.

[1747] "Example 2: Combining Emotion Engines"

[1748] (Claim 1)

[1749] A means of acquiring researcher gaze data through an eye-tracking device;

[1750] a means for acquiring researcher emotion data using an emotion engine;

[1751] means for transmitting the acquired viewpoint data and emotion data to a server;

[1752] A means for analyzing viewpoint data and emotion data in a server and verbalizing the analysis results using a generative AI model;

[1753] A means for storing the analyzed verbalization data, viewpoint data, and emotion data;

[1754] A means of sharing the stored data with interested parties;

[1755] A system including:

[1756] (Claim 2)

[1757] 10. The system of claim 1, further comprising means for calibrating the eye tracking device and the emotion engine.

[1758] (Claim 3)

[1759] 10. The system of claim 1, further comprising means for acquiring gaze data and emotion data in real time and displaying analysis results in real time.

[1760] "Application example 2 when combining emotion engines"

[1761] (Claim 1)

[1762] A means for acquiring gaze data of a worker by an eye tracking device;

[1763] A means for acquiring emotion data of a worker by an emotion engine;

[1764] means for transmitting the acquired viewpoint data and emotion data to a server;

[1765] a means for analyzing and verbalizing the viewpoint data and emotion data in the server;

[1766] a means for storing the analyzed verbalization data, viewpoint data, and emotion data;

[1767] a means for sharing the stored data;

[1768] A system including:

[1769] (Claim 2)

[1770] 10. The system of claim 1, further comprising means for calibrating the eye tracking device and the emotion engine.

[1771] (Claim 3)

[1772] 10. The system of claim 1, further comprising means for acquiring gaze data and emotion data in real time and displaying analysis results in real time. [Explanation of symbols]

[1773] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. A means of acquiring researcher gaze data through an eye-tracking device; means for transmitting the acquired viewpoint data to a server; A means for analyzing and verbalizing viewpoint data in a server; a means for storing the analyzed verbalization data and viewpoint data; a means for sharing the stored data; A system including:

2. The system of claim 1 , further comprising means for calibrating the eye tracking device.

3. The system according to claim 1 , further comprising means for acquiring viewpoint data in real time and displaying analysis results in real time.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A