System
The system addresses confusion in online meetings by using face recognition and real-time explanation generation to improve communication efficiency and participant safety.
Patent Information
- Application Number
- JP2024116356
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-19
- Publication Date
- 2026-01-29
AI Technical Summary
In online meetings, participants often fail to understand technical terms and concepts, leading to inefficiencies and reduced psychological safety due to unresolved doubts and confusion.
A system utilizing face recognition, facial expression analysis, confusion detection, and real-time explanation generation to identify and address participants' confusion by displaying explanations on their screens.
Enables immediate resolution of confusion during meetings, ensuring smooth communication and improved participant understanding and satisfaction.
Smart Images

Figure 2026014882000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] In online meetings, if participants do not understand a particular word or concept, the meeting often proceeds without resolving their doubts, which can lead to problems that undermine the efficiency of the meeting and the psychological safety of participants. This problem is particularly pronounced in meetings that involve a lot of technical terms and new concepts. The present invention aims to improve the efficiency of online meetings and ensure the psychological safety of participants by detecting participants' doubts and confusion in real time and providing appropriate explanations. [Means for solving the problem]
[0005] The present invention provides a system including a face recognition means, a facial expression analysis means, a confusion state detection means, a generation means, and a display means. The face recognition means detects the faces of participants from a video feed of an online conference, and the facial expression analysis means analyzes their facial expressions. The confusion state detection means identifies facial expressions that indicate participants are confused or skeptical, and the generation means automatically generates explanations for words or concepts that cause confusion. Finally, the display means displays the generated explanations on the participants' screens in real time. Through this series of processes, participants' questions can be immediately resolved while the conference is in progress, enabling smooth communication.
[0006] "Facial Recognition Means" means technological means for detecting a participant's face from a video feed.
[0007] The "expression analysis means" is a technical means for analyzing the detected facial expression and determining the type of expression.
[0008] The "confusion detection means" is a technical means for identifying whether a participant is feeling confused or confused based on the results of facial expression analysis.
[0009] "Generative means" are technical means for automatically generating explanations for confusing words or concepts.
[0010] "Display means" refers to the technical means for displaying the generated explanation on the participant's screen in real time.
[0011] "Video Feed Acquisition Means" means the technical means for acquiring the video feeds of participants from an online conference.
[0012] A "real-time processing means" is a technical means that performs each processing step in real time for each frame of the video feed. [Brief explanation of the drawings]
[0013] [Figure 1]1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0014] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0015] First, the terms used in the following description will be explained.
[0016] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0017] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0018] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0019] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0020] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0021] [First embodiment]
[0022] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0023] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0024] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0025] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0026] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0027] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0028] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0029] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0030] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0031] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0032] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0033] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0034] The present invention describes a system for quickly identifying words or concepts that confuse online meeting participants and providing explanations for them. A specific embodiment of this system is described below.
[0035] System initialization
[0036] The server starts the system, which includes a facial recognition unit, a facial expression analysis unit, a confusion state detection unit, a generation unit, and a display unit, and loads the models and necessary libraries for each unit. The server also works with an online conference tool to prepare for capturing the user's video feed in real time.
[0037] Capturing a video feed
[0038] The server captures each user's video feed and converts it into an analyzable format, providing the data necessary for face detection and facial expression analysis of participants.
[0039] Face detection and facial expression analysis
[0040] The server analyzes each frame of the captured video feed, detects the faces of participants using a facial recognition means, and acquires facial expression data of the detected faces using an expression analysis means to determine whether the user is confused or confused.
[0041] Detecting confusion
[0042] If a participant is found to have a confused expression, the server records the user's face ID and the current timestamp, allowing it to determine the exact moment of confusion.
[0043] Identifying confusing words
[0044] The server analyzes the audio data from the online meeting based on the recorded timestamps and identifies the words or concepts that are causing confusion. This analysis uses speech recognition and natural language processing technologies.
[0045] Generate word descriptions
[0046] The server sends the identified words and concepts to a generator (e.g., a generative AI) that generates a concise and accurate explanation for them, presented in a format that is easy for the user to understand.
[0047] Display Description
[0048] Finally, the server displays the generated explanation on the confused user's device in a pop-up format for immediate review, and can provide the same explanation to other conference participants if necessary.
[0049] Specific examples
[0050] For example, if the term "ecosystem" is used during a meeting and User A finds it difficult to understand, the following will work:
[0051] 1. The server detects that User A has a confused expression.
[0052] 2. Analyze the conversation based on timestamps and identify that the word "ecosystem" is the source of confusion.
[0053] 3. Use the generative tool to generate the statement, "An ecosystem is a collection of interconnected organisms and their surrounding environment."
[0054] 4. This explanation will be displayed in real time on User A's screen.
[0055] This allows the question of User A to be resolved immediately, allowing the flow of the meeting to proceed without being interrupted. This system also works effectively when other participants have the same problem.
[0056] The processing flow will be explained below.
[0057] Step 1:
[0058] The server loads models and libraries for face recognition, facial expression analysis, confusion detection, generation, and display. It also establishes a connection with the online conference tool and prepares to acquire the video feed.
[0059] Step 2:
[0060] The server captures each user's video feed and converts it into a format that can be analyzed in real time, thereby continuously receiving video data from participants.
[0061] Step 3:
[0062] The server analyzes each frame of the captured video feed and detects participants' faces using facial recognition techniques, and the detected face information is recorded for each frame.
[0063] Step 4:
[0064] The server uses an expression analysis means to obtain facial expression data of the detected face, which includes detailed facial expression information such as eyebrow movements, eye opening and closing, and mouth movements.
[0065] Step 5:
[0066] The server analyzes the facial expression data acquired by the server using a confusion detection method to identify the participant's confused expression. The server records the timestamp and face ID of the identified confused expression.
[0067] Step 6:
[0068] The server analyzes the audio data of the conversation based on the recorded timestamps to identify words or concepts that may be confusing, and uses speech recognition technology to convert the audio data into text.
[0069] Step 7:
[0070] The server uses a generator to automatically generate explanations for the identified words and concepts in a concise and easy-to-understand format.
[0071] Step 8:
[0072] The server sends the generated explanation to the terminal of the confused participant via a display means. The explanation is displayed in a pop-up format so that the participant can immediately check it.
[0073] Step 9:
[0074] Users can review explanations displayed on their devices to gain a better understanding of words or concepts they were confused about, allowing them to keep up with the progress of the meeting.
[0075] Step 10:
[0076] The server continuously performs this process for each frame of the video feed, detecting and addressing confusion in real time, allowing for immediate response if new questions arise during the meeting.
[0077] Example 1
[0078] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0079] In online meetings, participants may be confused by the technical terms and concepts used, which can cause communication to be disrupted. This confusion reduces the effectiveness of the meeting and hinders participants' understanding. Therefore, a system that can instantly provide easy-to-understand explanations for content that participants are confused about during a meeting is desirable.
[0080] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0081] In this invention, the server includes a face recognition means, an expression analysis means, a confusion detection means, a generation means, a display means, a video feed acquisition means, an audio data analysis means, a natural language processing means, and a pop-up display means, which enable the server to detect when a participant shows a confused expression, identify target words or concepts based on that timing, and provide explanations to resolve the confusion in real time.
[0082] "Facial Recognition Method" refers to technology for detecting participants' faces within a video feed and analyzing their features.
[0083] The "facial expression analysis means" is a technology for analyzing detected facial expression data and determining the emotional state of the participant.
[0084] The "confusion state detection means" is a technology that uses facial expression analysis means to determine whether a participant is confused and detects that state.
[0085] "Generative means" refers to technology for generating explanations for identified words or concepts, including generative AI models.
[0086] "Display means" refers to a technique for presenting the generated explanation to participants.
[0087] The "video feed acquisition means" is a technology that acquires a video feed from an online conference tool and converts it into an analyzable format.
[0088] "Audio data analysis means" is a technology that analyzes the audio data of a conference and converts it into text.
[0089] "Natural language processing means" is a technology that analyzes acquired text data and identifies words and concepts that cause confusion.
[0090] "Pop-up display means" is a technique for visually presenting the generated explanation to participants in real time.
[0091] The present invention is a system that instantly identifies words or concepts that participants in an online meeting find confusing and provides explanations for them. This system is built using a combination of specific hardware and software.
[0092] Hardware and software used
[0093] The server runs the system using the following hardware and software:
[0094] Face recognition method (e.g. OpenCV)
[0095] Facial expression analysis means (e.g. Facial Expression Recognition)
[0096] Confusion detection means
[0097] Generation method (e.g. GPT-3)
[0098] Display means
[0099] Video feed acquisition method (e.g. Zoom, Teams)
[0100] Voice data analysis method (e.g., Google Speech-to-Text API)
[0101] Natural language processing tools (e.g., NLTK, Transformers)
[0102] Pop-up display method
[0103] Program processing overview
[0104] The server starts the various components of the system and loads the necessary models and libraries, including the facial recognition model, facial expression analysis model, confusion detection model, natural language processing library, and generative AI model. The server also connects with online meeting tools to obtain user video feeds.
[0105] The server captures each user's video feed in real time via the online conferencing tool, converts the video feed into an analyzable format, and uses it as data for facial recognition and facial expression analysis.
[0106] The server sequentially analyzes each frame of the captured video feed and detects the participant's face using a facial recognition model, then uses an expression analysis model to obtain facial expression data from the detected face and determine whether the user is confused or not.
[0107] If the user shows a confused expression, the server records the confused user's face ID and the current timestamp, allowing the specific moment of confusion to be identified.
[0108] The server analyzes the audio data of online meetings based on the recorded timestamps, converts the audio data into text using speech recognition technology, and then uses natural language processing technology to extract words or concepts that are causing confusion.
[0109] The server uses a generative AI model to generate an explanation for the identified word or concept: for example, for "ecosystem," it generates the explanation "an ecosystem is a collection of interconnected organisms and their surrounding environment."
[0110] Finally, the server sends the generated explanation to the confused user's terminal, where it displays the explanation in a pop-up format and can provide the same explanation to other conference participants.
[0111] Specific examples
[0112] For example, if the term "ecosystem" is used during a meeting and it is difficult for User A to understand, the following will work:
[0113] 1. The server detects that User A has a confused expression.
[0114] 2. Analyze the conversation based on timestamps and identify that the word "ecosystem" is the source of confusion.
[0115] 3. The server uses a generative AI model to generate the explanation that "an ecosystem is a collection of interconnected organisms and their surrounding environments."
[0116] 4. This explanation will be displayed in real time on User A's screen.
[0117] Prompt Sentence Examples
[0118] "The term 'ecosystem' comes up during a meeting and the user looks confused. Generate a concise explanation of this term and display it to the user in a popup."
[0119] This system makes it possible to instantly resolve any confusion participants may have during online meetings and ensure smooth communication.
[0120] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0121] Step 1: Initialize the system
[0122] The server initiates the system's various functions, including preparing to load models and libraries for facial recognition, facial expression analysis, confusion detection, generative AI models, video feed acquisition, audio data analysis, and display, such as OpenCV, Facial Expression Recognition, NLTK, Transformers, and GPT-3.
[0123] Input: Models and libraries of various methods
[0124] Output: Initialized system
[0125] Step 2: Capture and convert your video feed
[0126] The server captures the user's video feed in real time through the online conferencing tool and converts it into a format that can be analyzed: the video feed is converted into a series of JPEG frames.
[0127] Input: Video feed from online meeting tool
[0128] Output: Video frames converted to a parsable format
[0129] Step 3: Face detection
[0130] The server analyzes each video frame and detects the participants' faces using a facial recognition method, such as OpenCV.
[0131] Input: Transformed video frames
[0132] Output: Detected face data
[0133] Step 4: Facial Expression Analysis
[0134] The server uses the facial expression analysis model to analyze the facial expressions detected by the facial recognition means, thereby determining the emotional state of the participant.
[0135] Input: Detected face data
[0136] Output: Analyzed facial expression data
[0137] Step 5: Detecting confusion
[0138] If a participant shows a confused expression, the server records the user's face ID and the current timestamp, allowing the specific moment of confusion to be identified.
[0139] Input: Analyzed facial expression data
[0140] Output: Confusion detection result (face ID and timestamp)
[0141] Step 6: Analyze the audio data
[0142] The server analyzes the audio data of the online meeting based on the recorded timestamps, and the audio data is converted into text using speech recognition technology.
[0143] Input: Online meeting audio data and timestamps
[0144] Output: Audio data converted to text
[0145] Step 7: Identify the confusing word
[0146] The server analyzes the text of the speech data and identifies confusing words and concepts using natural language processing means.
[0147] Input: Audio data converted to text
[0148] Output: Words or concepts that cause confusion
[0149] Step 8: Generate word descriptions
[0150] The server sends the identified words and concepts to a generative AI model to generate a concise and accurate description of them. For example, a generative AI model can be used to generate a description of an "ecosystem."
[0151] Input: Identified words or concepts
[0152] Output: Generated description
[0153] Step 9: View the description
[0154] The server sends the generated explanation to the confused user's device, which displays it in a pop-up format. If necessary, the same explanation can be provided to other conference participants.
[0155] Input: Generated Description
[0156] Output: Description displayed on the user's terminal
[0157] This series of processes makes it possible to provide an appropriate explanation immediately and support smooth communication even if a user becomes confused during a meeting.
[0158] (Application example 1)
[0159] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0160] In modern factories, workers are required to quickly adapt to new machines and technologies to maintain or increase productivity. However, if workers are confused by technical terms and operating procedures, time is wasted and work efficiency is reduced. To address this issue, a system that allows workers to instantly resolve their questions is needed. Therefore, a system and method that allows workers to obtain explanations of technical terms and operating procedures in real time is needed.
[0161] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0162] In this invention, the server includes a face recognition means, an expression analysis means, a confusion state detection means, a generation means, a display means, a real-time video feed acquisition means, an audio analysis means, and a visual display means for displaying the generated explanation, thereby enabling visual explanations of technical terms and operating procedures to be provided in real time when a worker is confused.
[0163] "Facial recognition tools" are technologies for detecting faces in video feeds and identifying individual people.
[0164] "Facial expression analysis means" is a technology for analyzing recognized facial expressions and inferring the emotional state of the person.
[0165] The "confusion state detection means" is a technology that determines whether the user is confused or not based on the results of facial expression analysis.
[0166] A "generative means" is a technology for generating specific information or explanations, which in this case is a generative AI model.
[0167] "Display means" refers to a technique for visually displaying the generated information and explanations to the user.
[0168] "Real-time video feed acquisition means" refers to technology for acquiring and processing video of the current user in real time.
[0169] "Speech analysis means" refers to technology that analyzes speech data in real time and recognizes specific words and phrases.
[0170] "Visual display means" refers to the display of smart glasses or other devices and the technology used to display information thereon.
[0171] A "generative AI model" is an artificial intelligence model that generates appropriate content based on input information.
[0172] The present invention is a system for providing appropriate explanations in real time to factory workers so that they can continue working without being confused by technical terms and operating procedures. The system includes the following means.
[0173] The server includes a facial recognition means, an expression analysis means, a confusion detection means, a generation means, a display means, a real-time video feed acquisition means, an audio analysis means and a visual display means.
[0174] Hardware and software used
[0175] For facial recognition and facial expression analysis, Google MediaPipe and Azure Face API are used, for example. Real-time video feeds are acquired using the smart glasses' camera. Voice data analysis is performed using Google Cloud Speech-to-Text and Amazon Transcribe. Explanation generation utilizes generative AI models such as OpenAI's GPT-4, which generate appropriate explanations in real time. The generated explanations are then provided to the user using visual display means such as the smart glasses' display.
[0176] Data processing and data calculation
[0177] The server processes various data in real time. When a video feed is sent from the smart glasses, the face of the worker is identified using a facial recognition means, and confusion is detected using an expression analysis means. If a confusion state is detected, the voice data is analyzed based on the timestamp to identify the technical term or operation procedure that is causing the confusion. The technical term or operation procedure is then sent to a generative AI model, which generates a concise explanation for it. This explanation is displayed in real time on the visual display means (the display of the smart glasses).
[0178] Examples of concrete examples and prompts
[0179] As a concrete example, if a factory worker is operating a new machine and is confused by the technical term "scan speed," the system works as follows:
[0180] 1. Detect when a worker has a confused expression.
[0181] 2. Analyze the conversation based on timestamps and identify that the technical term "scanning rate" is the source of confusion.
[0182] 3. Use a generative AI model to generate the explanation, "Scanning speed is the speed at which the scan head scans the surface. It is typically expressed in meters per second (m / s)."
[0183] 4. This explanation will be displayed in real time on the smart glasses display.
[0184] Example prompt sentence:
[0185] "Scanning speed is the speed at which the scan head scans the surface."
[0186] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0187] Step 1:
[0188] Getting a Video Feed
[0189] The device (smart glasses) captures a real-time video feed of the user and sends it to the server. The input is the video feed, which serves as the data source for facial recognition and facial expression analysis. The output is image data for analysis.
[0190] Step 2:
[0191] Facial Recognition and Expression Analysis
[0192] The server uses a facial recognition mechanism to identify faces in the video feed and a facial expression analysis mechanism to analyze the user's facial expressions. The input is image data from the video feed, and the output is the user's face ID and its facial expression data. Processing is performed using Google MediaPipe and Azure Face API.
[0193] Step 3:
[0194] Detecting confusion
[0195] The server uses a confusion detection means to determine whether the user is confused based on the results of the facial expression analysis. The input is facial expression data, and the output is the user's confusion state and its timestamp. If a confusion state is detected, the timestamp is recorded.
[0196] Step 4:
[0197] Analysis of audio data
[0198] The server uses voice analysis tools to analyze the audio data of online meetings based on the recorded timestamps and identify confusing technical terms and operating procedures. The input is the audio data and timestamps, and the output is text data of the technical terms and operating procedures. Google Cloud Speech-to-Text and Amazon Transcribe are used.
[0199] Step 5:
[0200] Generate Description
[0201] The server sends the identified technical terms and operating procedures to a generative AI model, which generates a concise explanation for them. The input is text data of the technical terms and operating procedures, and the output is explanatory text. The explanation is generated using OpenAI's GPT-4 or similar software.
[0202] Step 6:
[0203] Display Description
[0204] The server sends the generated explanation to the display of the terminal (smart glasses) and displays it to the user using a visual display means. The input is the explanation text, and the output is the information displayed in the user's field of view. In concrete terms, the smart glasses display displays "Scanning speed is the speed at which the scan head scans the surface. It is typically expressed in meters per second (m / s)."
[0205] By repeating this series of steps, workers can proceed with their work without confusion.
[0206] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0207] The present invention aims to further improve the efficiency of online meetings and the satisfaction of participants by incorporating an emotion engine that recognizes the emotions of users in addition to a system that quickly detects confusion among participants in online meetings and provides appropriate explanations. Specific embodiments of the present invention will be described below.
[0208] System initialization
[0209] The server loads models and libraries for face recognition, facial expression analysis, confusion detection, generation, display, and emotion recognition, performs necessary initialization, establishes a connection with the online conferencing tool, and prepares to capture each user's video feed in real time.
[0210] Capturing a video feed
[0211] The server captures each user's video feed, converts it into an analyzable format, and stores it. This video feed is the key data used for face detection and facial expression and emotion analysis.
[0212] Face detection and facial expression analysis
[0213] The server analyzes the captured video feed and detects each participant's face using a facial recognition means, and acquires facial expression data of the detected faces using an expression analysis means, including detailed facial expression information such as eyebrow movements, eye opening and closing, and mouth movements.
[0214] Detecting embarrassment and emotions
[0215] The server uses both facial expression analysis and emotion recognition to determine whether the participant is confused or not. In addition, it simultaneously analyzes the user's emotions (happiness, sadness, anger, etc.) and obtains detailed emotional data.
[0216] Identifying confusing words
[0217] The server analyzes the conversation based on the identified confusion state and its timestamp, and identifies words and concepts that are the subject of confusion. This analysis process uses speech recognition and natural language processing technologies.
[0218] Generate word descriptions
[0219] The server uses a generator to automatically generate explanations for the identified words and concepts, and the generated explanations are provided to the participants in a concise and easy-to-understand format.
[0220] Displaying explanations and emotional feedback
[0221] The server sends the generated explanation to the confused user's device via a display means, and may provide similar feedback to other participants in some cases. It may also provide feedback regarding the progress of the meeting based on the emotion recognition results.
[0222] Specific examples
[0223] For example, if User B looks confused when the term "ecosystem" is used during a meeting:
[0224] 1. The server detects User B's confusion state and timestamp.
[0225] 2. Using facial expression analysis and emotion recognition means, we recognize that User B is not only confused but also a little nervous.
[0226] 3. Analyze the conversation and identify that the word "ecosystem" is the source of the question.
[0227] 4. Use the generative tools to generate the statement, "An ecosystem is a collection of interconnected organisms and their surrounding environment."
[0228] 5. This description may pop up on User B's screen and be shared with other participants.
[0229] 6. Based on emotional feedback, advice and tips on how to proceed with the meeting are also provided.
[0230] This series of processes quickly resolves any questions or confusion that meeting participants may have, allowing everyone to participate comfortably in the meeting. Feedback based on emotional data can also be used to improve the atmosphere of the meeting.
[0231] The processing flow will be explained below.
[0232] Step 1:
[0233] The server loads models and libraries for face recognition, facial expression analysis, confusion state detection, generation, display, and emotion recognition, performs necessary initial settings, and prepares to establish a connection with the online conference tool.
[0234] Step 2:
[0235] The server captures each participant's video feed in real time through the online conferencing tool, converts this data into a parsable format, and prepares it for processing.
[0236] Step 3:
[0237] The server analyzes each frame of the captured video feed and uses facial recognition to detect participants' faces, which are then stored for use in the next step.
[0238] Step 4:
[0239] The server acquires the facial expression data detected using the facial expression analysis means, and analyzes detailed facial expression information such as eyebrow movements, eye opening and closing, and mouth movements.
[0240] Step 5:
[0241] The server uses emotion recognition means to analyze facial expression data and other biometric signals to identify the user's emotions (happiness, sadness, anger, confusion, etc.), thereby determining with high accuracy whether the user is confused.
[0242] Step 6:
[0243] If the server detects a confused state, it records the timestamp and the face ID of the user. Based on this information, the next step is to analyze the conversation.
[0244] Step 7:
[0245] The server analyzes the audio data of the meeting based on the recorded timestamps, identifies words or concepts that are confusing, converts the audio data into text using speech recognition technology, and extracts keywords using natural language processing technology.
[0246] Step 8:
[0247] The server uses a generator to automatically generate explanations for the identified words and concepts, which are then summarized in a concise and easy-to-understand format.
[0248] Step 9:
[0249] The server sends the generated explanation to the confused user's terminal and displays it in a pop-up format through a display means, and provides similar feedback to other participants as needed.
[0250] Step 10:
[0251] Users can review explanations displayed on their devices to gain a better understanding of words or concepts they were confused about, allowing them to keep up with the progress of the meeting.
[0252] Step 11:
[0253] The server uses the emotion data to provide feedback and advice on how the meeting should proceed. For example, if the user is nervous, it will encourage them to relax, and if they are tired, it will suggest taking a break.
[0254] Step 12:
[0255] The server runs this process continuously for each frame of the video feed, detecting and addressing confusion and emotions in real time, allowing for immediate response as new questions or emotional states arise.
[0256] Example 2
[0257] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0258] In online meetings, participants often become confused by technical terms and complex concepts they do not understand. This confusion can hinder the progress of the meeting and reduce participant satisfaction. It is also problematic when the meeting proceeds without noticing that some participants are confused. Furthermore, if participants' emotional states cannot be addressed, the meeting atmosphere becomes awkward and productivity decreases. An efficient solution to this problem is needed.
[0259] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0260] In this invention, the server includes face recognition means, facial expression analysis means, confusion detection means, generation means, display means, emotion recognition means, and conversation content analysis means, which enable the server to quickly detect participants' confusion, provide appropriate explanations, and analyze participants' emotional states in real time to provide feedback.
[0261] "Facial recognition means" is a technology for identifying the faces of each user participating in an online conference.
[0262] The "facial expression analysis means" is a technology for analyzing the facial expressions of each detected user and acquiring facial expression data.
[0263] The "confusion state detection means" is a technique for determining whether the user is confused or not based on the facial expression data obtained by the facial expression analysis means.
[0264] A "generator" is a technology that automatically generates explanations for identified words or concepts.
[0265] The "display means" is a technique for displaying the generated explanation on the user's terminal.
[0266] The "emotion recognition means" is a technology that analyzes the user's emotional state (joy, sadness, anger, etc.) based on the facial expression data obtained by the facial expression analysis means.
[0267] The "conversation content analysis means" is a technology that analyzes the conversation content before and after the time when a confused state is detected, and identifies the specific words or phrases that cause the confusion.
[0268] The present invention provides a system for quickly detecting participant confusion in an online conference and providing appropriate explanations. This system further incorporates an emotion engine that recognizes user emotions, thereby improving the efficiency of the conference and participant satisfaction. A specific embodiment of the system is described below.
[0269] The server first loads models and libraries for face recognition, facial expression analysis, confusion detection, generation, display, and emotion recognition. This process uses machine learning libraries such as TensorFlow and Keras. It also establishes a connection with online conferencing tools (e.g., Zoom, Microsoft Teams) and prepares to capture each user's video feed in real time. At this time, it obtains the necessary API keys and access tokens.
[0270] The server then captures each user's video feed and converts it into a format that can be analyzed using a conversion tool such as FFmpeg. This video data is crucial for facial recognition and facial expression analysis.
[0271] The server analyzes the captured video feed frame by frame and detects each participant's face using the OpenCV library. It then uses the dlib library to extract facial feature points and perform a detailed analysis of the movements of each part of the face. Based on this data, the facial expression analyzer performs a detailed analysis of the user's facial expressions.
[0272] Furthermore, the server uses the Emotion Recognition API to analyze the user's emotional state based on the acquired facial expression data, for example, determining whether the participant is confused or expressing emotions such as joy, sadness, or anger.
[0273] If a confusion state is detected, the server analyzes the conversation content based on the timestamp. The conversation content is converted into text using online conferencing tools or speech recognition technology (e.g., Google Cloud Speech-to-Text API), and analyzed with a natural language processing engine (e.g., the BERT model) to identify words and phrases that cause confusion.
[0274] The server then uses a generative method to automatically generate explanations for the identified words and phrases. Specifically, it inputs the prompt "What is an ecosystem?" into a generative AI model such as GPT-3 and uses the generated answer to create an explanation. For example, it generates an explanation such as "An ecosystem is a collection of interconnected organisms and their surrounding environment."
[0275] Finally, the server sends the generated explanation to the user's device and displays it as a pop-up on the display. In some cases, the same explanation can be shared with other participants. The server also provides feedback on the progress of the meeting based on the emotion recognition results. For example, it may give advice such as, "The participants seem nervous, so it would be more effective if you spoke a little more slowly."
[0276] This series of processes quickly resolves any questions or confusion that meeting participants may have, allowing everyone to participate comfortably in the meeting. Feedback based on emotional data can also be used to improve the atmosphere of the meeting.
[0277] Prompt Sentence Examples
[0278] "What is an ecosystem?"
[0279] "Please provide additional explanation for the confusion."
[0280] "Please analyze participants' facial expression data and perform emotion recognition."
[0281] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0282] Step 1:
[0283] The server loads models and libraries for face recognition, facial expression analysis, confusion detection, generation, display, and emotion recognition. Specifically, it loads the necessary models into memory using machine learning libraries such as TensorFlow and Keras. It also establishes a connection using the online conference tool's API and prepares to obtain each user's video feed. The input is a request to load the models and libraries, and the output is the results of those loads.
[0284] Step 2:
[0285] The server captures each user's video feed in real time. Specifically, it obtains the video stream through the API of a conferencing tool such as Zoom or Microsoft Teams, and converts it into an analyzable format (e.g., MP4) using FFmpeg. This video data is used for subsequent processing. The input is each user's video feed, and the output is the video data converted into an analyzable format.
[0286] Step 3:
[0287] The server analyzes the captured video feed frame by frame and detects each participant's face using the OpenCV library. It then uses the dlib library to extract facial feature points and perform detailed analysis of eyebrow, eye, and mouth movements. The input is each frame of video data, and the output is facial data with extracted feature points.
[0288] Step 4:
[0289] The server uses facial expression analysis means to obtain facial expression data based on the acquired facial data. Specifically, it uses the Emotion Recognition API to analyze the user's emotional state based on information such as eyebrow movements, eye opening and closing, and mouth movements. The input is facial data, and the output is detailed data on each facial expression and the results of that emotion analysis.
[0290] Step 5:
[0291] The server uses the confusion detection means to determine the user's confusion state. Based on the data obtained by the facial expression analysis means, it applies a specific algorithm (e.g., random forest or support vector machine) to determine whether the user is confused. The input is facial expression data and the emotion analysis result, and the output is the confusion state determination result.
[0292] Step 6:
[0293] The server analyzes the conversation content based on the timestamp of the user who was detected as confused. It acquires the audio data using the conferencing tool's API and converts the audio into text using the Google Cloud Speech-to-Text API. It then analyzes the text data using a natural language processing engine such as the BERT model to identify words and phrases that cause confusion. The input is the audio data, and the output is the text data and the identified words and phrases.
[0294] Step 7:
[0295] The server automatically generates explanations for the identified words and phrases using a generation method. Specifically, a prompt such as "What is an ecosystem?" is input to a generative AI model such as GPT-3, and an explanation is generated based on the response. The input is the prompt, and the output is the generated explanation.
[0296] Step 8:
[0297] The server sends the generated explanation to the confused user's device and displays it as a pop-up using a display device. It also shares the same explanation with other participants as needed. It also provides feedback on the progress of the meeting based on the emotion recognition results. The inputs are the generated explanation and the emotion analysis results, and the outputs are the display of the explanation and the provision of feedback.
[0298] (Application example 2)
[0299] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0300] Conventional online conference systems and factory work support systems have the problem of making it difficult to quickly detect situations and provide appropriate feedback when workers or meeting participants are confused or emotionally troubled. In particular, in factories, confusion can lead to reduced work efficiency and incorrect work procedures, which can also affect safety. To solve this problem, the present invention provides a system that detects a worker's emotions and confusion in real time and provides appropriate support.
[0301] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes a face recognition means, a facial expression analysis means, an emotion recognition means, a confusion state detection means, a generation means, a display means, a video feed acquisition means, a real-time processing means, a means for generating an explanation using a generative AI model, and a means for generating an explanation using a prompt sentence. This makes it possible to detect in real time that a worker is confused, and to instantly generate and display appropriate explanations regarding specific work procedures. This makes it possible to improve work efficiency and ensure safety.
[0302] "Facial recognition tools" are technologies and devices used to detect an individual's face from a video feed and identify that person.
[0303] "Facial expression analysis means" refers to technology and devices for analyzing subtle changes in facial expressions from a detected face and identifying the person's emotional state.
[0304] "Emotion recognition means" refers to technology and devices that use facial expression analysis data to determine a person's emotions (joy, sadness, anger, confusion, etc.).
[0305] The "confusion state detection means" refers to a technique and device for identifying whether a person is confused or not based on emotion recognition and facial expression analysis.
[0306] The "generation means" refers to technology and devices for automatically generating necessary explanations and information based on the identified confusion state and emotions.
[0307] "Display means" refers to the technology and devices used to visually present the generated explanations and information to the user.
[0308] "Video Feed Acquisition Measures" means the technology and equipment used to capture video footage in real time and convert and store that data in an analyzable format.
[0309] "Real-time processing means" refers to techniques and devices that rapidly process captured video feeds and analytical data to provide immediate feedback.
[0310] "Means for generating explanations using generative AI models" refers to technologies and devices that use artificial intelligence models to automatically generate explanations based on specific situations and contexts.
[0311] The "means for generating an explanation using a prompt sentence" refers to a technique and device that allows an artificial intelligence model to generate an appropriate explanation based on a specific prompt sentence given as input.
[0312] The present invention is a system for detecting the emotions and confusion of factory workers in real time and providing appropriate support. Specific embodiments of the system will be described below.
[0313] System Configuration
[0314] The server includes a facial recognition means, a facial expression analysis means, an emotion recognition means, a confusion state detection means, a generation means, a display means, a video feed acquisition means, a real-time processing means, a means for generating an explanation using a generative AI model, and a means for generating an explanation using a prompt sentence.
[0315] Hardware and software used
[0316] This system uses the following hardware and software:
[0317] Camera: Mounted on the robot, it captures a video feed of the worker in real time.
[0318] Server: Processes the video feed and performs facial recognition, facial expression analysis, and emotion recognition.
[0319] dlib: A library for face detection.
[0320] EmotionEngine: A proprietary module for facial expression analysis and emotion recognition.
[0321] TextGenerator: A module for generating explanations using generative AI models.
[0322] Data processing and calculation
[0323] The server processes and calculates the data as follows:
[0324] 1. Capturing and saving video feed:
[0325] The robot's camera captures a real-time video feed of the worker, converts it into an analyzable format, and stores it.
[0326] 2. Facial Recognition and Expression Analysis:
[0327] The dlib library is used to detect the worker's face from the video feed, and the EmotionEngine is used to obtain detailed facial expression data.
[0328] 3. Detecting embarrassment and emotions:
[0329] The EmotionEngine analyzes facial expression data to determine whether the worker is confused and obtains emotional data (e.g., tension, anxiety, joy, etc.).
[0330] 4. Identify the task steps that cause confusion:
[0331] Based on the identified confusion state and its timestamp, the work content is analyzed to identify the part that causes the confusion.
[0332] 5. Generating explanations using generative means:
[0333] Using TextGenerator, a prompt sentence is input into the generative AI model to automatically generate an explanation for the cause of confusion.
[0334] 6. Displaying explanatory and emotional feedback:
[0335] The generated explanation is displayed on the robot's display, providing visual assistance to the worker.
[0336] Specific examples
[0337] As an example, here is the process that would occur if a worker becomes confused while performing "Step 5" in a factory.
[0338] 1. The server captures the worker's video feed through the robot's camera.
[0339] 2. The server analyzes the video feed and detects the worker's face using the dlib library.
[0340] 3. Use EmotionEngine to analyze facial expression and emotion data and recognize when the worker is confused.
[0341] 4. Using the timestamps from the video feed of the task, identify the task step ("Step 5") that caused the confusion.
[0342] 5. Using TextGenerator, enter the "Explanation of points where mistakes are likely to occur in step 5" as a prompt sentence into the generative AI model.
[0343] 6. The generated instructions are displayed on the robot's display, allowing the worker to refer to them and continue the work.
[0344] Example prompt sentence:
[0345] "Explanation of points where mistakes are likely to occur in step 5"
[0346] Through the above series of processes, it is possible to detect situations in which a worker is confused in real time and provide appropriate support, thereby improving work efficiency and reducing worker stress.
[0347] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0348] Step 1:
[0349] The server captures the worker's video feed in real time using the robot's camera. The input is the video footage obtained from the camera, and the output is the captured video feed. The server converts it into an analyzable format and stores it.
[0350] Step 2:
[0351] The server uses the dlib library to detect the worker's face from the video feed. The input is the video feed saved in step 1, and the output is the location information (coordinates) of the detected face. The server performs face recognition processing based on this location information.
[0352] Step 3:
[0353] The server uses the EmotionEngine to analyze facial expressions. The input is the facial position information and video feed obtained in step 2, and the output is detailed facial expression data. The server analyzes eyebrow movements, eye opening and closing, mouth movements, etc.
[0354] Step 4:
[0355] The server further uses the EmotionEngine to perform emotion recognition. The input is the facial expression data obtained in step 3, and the output is the worker's emotional data (e.g., joy, sadness, anger, confusion, etc.). The server determines whether the worker is confused.
[0356] Step 5:
[0357] The server analyzes the work content based on the identified confusion state and its timestamp. The input is the timestamp of the confusion state and the timestamp of the video feed, and the output is the specific work step that caused the confusion. The server analyzes the conversation and work content to identify the cause of the confusion.
[0358] Step 6:
[0359] The server uses TextGenerator to input an explanation of the cause of confusion as a prompt sentence into the generative AI model, which then generates an appropriate explanation. The input is a prompt sentence related to the identified work step (e.g., "An explanation of the points in step 5 where mistakes are likely to occur"), and the output is the generated explanation sentence. The server inputs this prompt sentence into the generative AI model to obtain the necessary explanation.
[0360] Step 7:
[0361] The server displays the generated explanation on the robot's display. The input is the explanatory text obtained in step 6, and the output is visual support information displayed on the display. The server immediately provides appropriate explanations when a worker becomes confused, thereby improving work efficiency and ensuring safety.
[0362] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0363] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0364] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0365] [Second embodiment]
[0366] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0367] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0368] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0369] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0370] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0371] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0372] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0373] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0374] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0375] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0376] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0377] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0378] The present invention describes a system for quickly identifying words or concepts that confuse online meeting participants and providing explanations for them. A specific embodiment of this system is described below.
[0379] System initialization
[0380] The server starts the system, which includes a facial recognition unit, a facial expression analysis unit, a confusion state detection unit, a generation unit, and a display unit, and loads the models and necessary libraries for each unit. The server also works with an online conference tool to prepare for capturing the user's video feed in real time.
[0381] Capturing a video feed
[0382] The server captures each user's video feed and converts it into an analyzable format, providing the data necessary for face detection and facial expression analysis of participants.
[0383] Face detection and facial expression analysis
[0384] The server analyzes each frame of the captured video feed, detects the faces of participants using a facial recognition means, and acquires facial expression data of the detected faces using an expression analysis means to determine whether the user is confused or confused.
[0385] Detecting confusion
[0386] If a participant is found to have a confused expression, the server records the user's face ID and the current timestamp, allowing it to determine the exact moment of confusion.
[0387] Identifying confusing words
[0388] The server analyzes the audio data from the online meeting based on the recorded timestamps and identifies the words or concepts that are causing confusion. This analysis uses speech recognition and natural language processing technologies.
[0389] Generate word descriptions
[0390] The server sends the identified words and concepts to a generator (e.g., a generative AI) that generates a concise and accurate explanation for them, presented in a format that is easy for the user to understand.
[0391] Display Description
[0392] Finally, the server displays the generated explanation on the confused user's device in a pop-up format for immediate review, and can provide the same explanation to other conference participants if necessary.
[0393] Specific examples
[0394] For example, if the term "ecosystem" is used during a meeting and User A finds it difficult to understand, the following will work:
[0395] 1. The server detects that User A has a confused expression.
[0396] 2. Analyze the conversation based on timestamps and identify that the word "ecosystem" is the source of confusion.
[0397] 3. Use the generative tool to generate the statement, "An ecosystem is a collection of interconnected organisms and their surrounding environment."
[0398] 4. This explanation will be displayed in real time on User A's screen.
[0399] This allows the question of User A to be resolved immediately, allowing the flow of the meeting to proceed without being interrupted. This system also works effectively when other participants have the same problem.
[0400] The processing flow will be explained below.
[0401] Step 1:
[0402] The server loads models and libraries for face recognition, facial expression analysis, confusion detection, generation, and display. It also establishes a connection with the online conference tool and prepares to acquire the video feed.
[0403] Step 2:
[0404] The server captures each user's video feed and converts it into a format that can be analyzed in real time, thereby continuously receiving video data from participants.
[0405] Step 3:
[0406] The server analyzes each frame of the captured video feed and detects participants' faces using facial recognition techniques, and the detected face information is recorded for each frame.
[0407] Step 4:
[0408] The server uses an expression analysis means to obtain facial expression data of the detected face, which includes detailed facial expression information such as eyebrow movements, eye opening and closing, and mouth movements.
[0409] Step 5:
[0410] The server analyzes the facial expression data acquired by the server using a confusion detection method to identify the participant's confused expression. The server records the timestamp and face ID of the identified confused expression.
[0411] Step 6:
[0412] The server analyzes the audio data of the conversation based on the recorded timestamps to identify words or concepts that may be confusing, and uses speech recognition technology to convert the audio data into text.
[0413] Step 7:
[0414] The server uses a generator to automatically generate explanations for the identified words and concepts in a concise and easy-to-understand format.
[0415] Step 8:
[0416] The server sends the generated explanation to the terminal of the confused participant via a display means. The explanation is displayed in a pop-up format so that the participant can immediately check it.
[0417] Step 9:
[0418] Users can review explanations displayed on their devices to gain a better understanding of words or concepts they were confused about, allowing them to keep up with the progress of the meeting.
[0419] Step 10:
[0420] The server continuously performs this process for each frame of the video feed, detecting and addressing confusion in real time, allowing for immediate response if new questions arise during the meeting.
[0421] Example 1
[0422] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0423] In online meetings, participants may be confused by the technical terms and concepts used, which can cause communication to be disrupted. This confusion reduces the effectiveness of the meeting and hinders participants' understanding. Therefore, a system that can instantly provide easy-to-understand explanations for content that participants are confused about during a meeting is desirable.
[0424] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0425] In this invention, the server includes a face recognition means, an expression analysis means, a confusion detection means, a generation means, a display means, a video feed acquisition means, an audio data analysis means, a natural language processing means, and a pop-up display means, which enable the server to detect when a participant shows a confused expression, identify target words or concepts based on that timing, and provide explanations to resolve the confusion in real time.
[0426] "Facial Recognition Method" refers to technology for detecting participants' faces within a video feed and analyzing their features.
[0427] The "facial expression analysis means" is a technology for analyzing detected facial expression data and determining the emotional state of the participant.
[0428] The "confusion state detection means" is a technology that uses facial expression analysis means to determine whether a participant is confused and detects that state.
[0429] "Generative means" refers to technology for generating explanations for identified words or concepts, including generative AI models.
[0430] "Display means" refers to a technique for presenting the generated explanation to participants.
[0431] The "video feed acquisition means" is a technology that acquires a video feed from an online conference tool and converts it into an analyzable format.
[0432] "Audio data analysis means" is a technology that analyzes the audio data of a conference and converts it into text.
[0433] "Natural language processing means" is a technology that analyzes acquired text data and identifies words and concepts that cause confusion.
[0434] "Pop-up display means" is a technique for visually presenting the generated explanation to participants in real time.
[0435] The present invention is a system that instantly identifies words or concepts that participants in an online meeting find confusing and provides explanations for them. This system is built using a combination of specific hardware and software.
[0436] Hardware and software used
[0437] The server runs the system using the following hardware and software:
[0438] Face recognition method (e.g. OpenCV)
[0439] Facial expression analysis means (e.g. Facial Expression Recognition)
[0440] Confusion detection means
[0441] Generation method (e.g. GPT-3)
[0442] Display means
[0443] Video feed acquisition method (e.g. Zoom, Teams)
[0444] Voice data analysis method (e.g., Google Speech-to-Text API)
[0445] Natural language processing tools (e.g., NLTK, Transformers)
[0446] Pop-up display method
[0447] Program processing overview
[0448] The server starts the various components of the system and loads the necessary models and libraries, including the facial recognition model, facial expression analysis model, confusion detection model, natural language processing library, and generative AI model. The server also connects with online meeting tools to obtain user video feeds.
[0449] The server captures each user's video feed in real time via the online conferencing tool, converts the video feed into an analyzable format, and uses it as data for facial recognition and facial expression analysis.
[0450] The server sequentially analyzes each frame of the captured video feed and detects the participant's face using a facial recognition model, then uses an expression analysis model to obtain facial expression data from the detected face and determine whether the user is confused or not.
[0451] If the user shows a confused expression, the server records the confused user's face ID and the current timestamp, allowing the specific moment of confusion to be identified.
[0452] The server analyzes the audio data of online meetings based on the recorded timestamps, converts the audio data into text using speech recognition technology, and then uses natural language processing technology to extract words or concepts that are causing confusion.
[0453] The server uses a generative AI model to generate an explanation for the identified word or concept: for example, for "ecosystem," it generates the explanation "an ecosystem is a collection of interconnected organisms and their surrounding environment."
[0454] Finally, the server sends the generated explanation to the confused user's terminal, where it displays the explanation in a pop-up format and can provide the same explanation to other conference participants.
[0455] Specific examples
[0456] For example, if the term "ecosystem" is used during a meeting and it is difficult for User A to understand, the following will work:
[0457] 1. The server detects that User A has a confused expression.
[0458] 2. Analyze the conversation based on timestamps and identify that the word "ecosystem" is the source of confusion.
[0459] 3. The server uses a generative AI model to generate the explanation that "an ecosystem is a collection of interconnected organisms and their surrounding environments."
[0460] 4. This explanation will be displayed in real time on User A's screen.
[0461] Prompt Sentence Examples
[0462] "The term 'ecosystem' comes up during a meeting and the user looks confused. Generate a concise explanation of this term and display it to the user in a popup."
[0463] This system makes it possible to instantly resolve any confusion participants may have during online meetings and ensure smooth communication.
[0464] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0465] Step 1: Initialize the system
[0466] The server initiates the system's various functions, including preparing to load models and libraries for facial recognition, facial expression analysis, confusion detection, generative AI models, video feed acquisition, audio data analysis, and display, such as OpenCV, Facial Expression Recognition, NLTK, Transformers, and GPT-3.
[0467] Input: Models and libraries of various methods
[0468] Output: Initialized system
[0469] Step 2: Capture and convert your video feed
[0470] The server captures the user's video feed in real time through the online conferencing tool and converts it into a format that can be analyzed: the video feed is converted into a series of JPEG frames.
[0471] Input: Video feed from online meeting tool
[0472] Output: Video frames converted to a parsable format
[0473] Step 3: Face detection
[0474] The server analyzes each video frame and detects the participants' faces using a facial recognition method, such as OpenCV.
[0475] Input: Transformed video frames
[0476] Output: Detected face data
[0477] Step 4: Facial Expression Analysis
[0478] The server uses the facial expression analysis model to analyze the facial expressions detected by the facial recognition means, thereby determining the emotional state of the participant.
[0479] Input: Detected face data
[0480] Output: Analyzed facial expression data
[0481] Step 5: Detecting confusion
[0482] If a participant shows a confused expression, the server records the user's face ID and the current timestamp, allowing the specific moment of confusion to be identified.
[0483] Input: Analyzed facial expression data
[0484] Output: Confusion detection result (face ID and timestamp)
[0485] Step 6: Analyze the audio data
[0486] The server analyzes the audio data of the online meeting based on the recorded timestamps, and the audio data is converted into text using speech recognition technology.
[0487] Input: Online meeting audio data and timestamps
[0488] Output: Audio data converted to text
[0489] Step 7: Identify the confusing word
[0490] The server analyzes the text of the speech data and identifies confusing words and concepts using natural language processing means.
[0491] Input: Audio data converted to text
[0492] Output: Words or concepts that cause confusion
[0493] Step 8: Generate word descriptions
[0494] The server sends the identified words and concepts to a generative AI model to generate a concise and accurate description of them. For example, a generative AI model can be used to generate a description of an "ecosystem."
[0495] Input: Identified words or concepts
[0496] Output: Generated description
[0497] Step 9: View the description
[0498] The server sends the generated explanation to the confused user's device, which displays it in a pop-up format. If necessary, the same explanation can be provided to other conference participants.
[0499] Input: Generated Description
[0500] Output: Description displayed on the user's terminal
[0501] This series of processes makes it possible to provide an appropriate explanation immediately and support smooth communication even if a user becomes confused during a meeting.
[0502] (Application example 1)
[0503] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0504] In modern factories, workers are required to quickly adapt to new machines and technologies to maintain or increase productivity. However, if workers are confused by technical terms and operating procedures, time is wasted and work efficiency is reduced. To address this issue, a system that allows workers to instantly resolve their questions is needed. Therefore, a system and method that allows workers to obtain explanations of technical terms and operating procedures in real time is needed.
[0505] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0506] In this invention, the server includes a face recognition means, an expression analysis means, a confusion state detection means, a generation means, a display means, a real-time video feed acquisition means, an audio analysis means, and a visual display means for displaying the generated explanation, thereby enabling visual explanations of technical terms and operating procedures to be provided in real time when a worker is confused.
[0507] "Facial recognition tools" are technologies for detecting faces in video feeds and identifying individual people.
[0508] "Facial expression analysis means" is a technology for analyzing recognized facial expressions and inferring the emotional state of the person.
[0509] The "confusion state detection means" is a technology that determines whether the user is confused or not based on the results of facial expression analysis.
[0510] A "generative means" is a technology for generating specific information or explanations, which in this case is a generative AI model.
[0511] "Display means" refers to a technique for visually displaying the generated information and explanations to the user.
[0512] "Real-time video feed acquisition means" refers to technology for acquiring and processing video of the current user in real time.
[0513] "Speech analysis means" refers to technology that analyzes speech data in real time and recognizes specific words and phrases.
[0514] "Visual display means" refers to the display of smart glasses or other devices and the technology used to display information thereon.
[0515] A "generative AI model" is an artificial intelligence model that generates appropriate content based on input information.
[0516] The present invention is a system for providing appropriate explanations in real time to factory workers so that they can continue working without being confused by technical terms and operating procedures. The system includes the following means.
[0517] The server includes a facial recognition means, an expression analysis means, a confusion detection means, a generation means, a display means, a real-time video feed acquisition means, an audio analysis means and a visual display means.
[0518] Hardware and software used
[0519] For facial recognition and facial expression analysis, Google MediaPipe and Azure Face API are used, for example. Real-time video feeds are acquired using the smart glasses' camera. Voice data analysis is performed using Google Cloud Speech-to-Text and Amazon Transcribe. Explanation generation utilizes generative AI models such as OpenAI's GPT-4, which generate appropriate explanations in real time. The generated explanations are then provided to the user using visual display means such as the smart glasses' display.
[0520] Data processing and data calculation
[0521] The server processes various data in real time. When a video feed is sent from the smart glasses, the face of the worker is identified using a facial recognition means, and confusion is detected using an expression analysis means. If a confusion state is detected, the voice data is analyzed based on the timestamp to identify the technical term or operation procedure that is causing the confusion. The technical term or operation procedure is then sent to a generative AI model, which generates a concise explanation for it. This explanation is displayed in real time on the visual display means (the display of the smart glasses).
[0522] Examples of concrete examples and prompts
[0523] As a concrete example, if a factory worker is operating a new machine and is confused by the technical term "scan speed," the system works as follows:
[0524] 1. Detect when a worker has a confused expression.
[0525] 2. Analyze the conversation based on timestamps and identify that the technical term "scanning rate" is the source of confusion.
[0526] 3. Use a generative AI model to generate the explanation, "Scanning speed is the speed at which the scan head scans the surface. It is typically expressed in meters per second (m / s)."
[0527] 4. This explanation will be displayed in real time on the smart glasses display.
[0528] Example prompt sentence:
[0529] "Scanning speed is the speed at which the scan head scans the surface."
[0530] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0531] Step 1:
[0532] Getting a Video Feed
[0533] The device (smart glasses) captures a real-time video feed of the user and sends it to the server. The input is the video feed, which serves as the data source for facial recognition and facial expression analysis. The output is image data for analysis.
[0534] Step 2:
[0535] Facial Recognition and Expression Analysis
[0536] The server uses a facial recognition mechanism to identify faces in the video feed and a facial expression analysis mechanism to analyze the user's facial expressions. The input is image data from the video feed, and the output is the user's face ID and its facial expression data. Processing is performed using Google MediaPipe and Azure Face API.
[0537] Step 3:
[0538] Detecting confusion
[0539] The server uses a confusion detection means to determine whether the user is confused based on the results of the facial expression analysis. The input is facial expression data, and the output is the user's confusion state and its timestamp. If a confusion state is detected, the timestamp is recorded.
[0540] Step 4:
[0541] Analysis of audio data
[0542] The server uses voice analysis tools to analyze the audio data of online meetings based on the recorded timestamps and identify confusing technical terms and operating procedures. The input is the audio data and timestamps, and the output is text data of the technical terms and operating procedures. Google Cloud Speech-to-Text and Amazon Transcribe are used.
[0543] Step 5:
[0544] Generate Description
[0545] The server sends the identified technical terms and operating procedures to a generative AI model, which generates a concise explanation for them. The input is text data of the technical terms and operating procedures, and the output is explanatory text. The explanation is generated using OpenAI's GPT-4 or similar software.
[0546] Step 6:
[0547] Display Description
[0548] The server sends the generated explanation to the display of the terminal (smart glasses) and displays it to the user using a visual display means. The input is the explanation text, and the output is the information displayed in the user's field of view. In concrete terms, the smart glasses display displays "Scanning speed is the speed at which the scan head scans the surface. It is typically expressed in meters per second (m / s)."
[0549] By repeating this series of steps, workers can proceed with their work without confusion.
[0550] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0551] The present invention aims to further improve the efficiency of online meetings and the satisfaction of participants by incorporating an emotion engine that recognizes the emotions of users in addition to a system that quickly detects confusion among participants in online meetings and provides appropriate explanations. Specific embodiments of the present invention will be described below.
[0552] System initialization
[0553] The server loads models and libraries for face recognition, facial expression analysis, confusion detection, generation, display, and emotion recognition, performs necessary initialization, establishes a connection with the online conferencing tool, and prepares to capture each user's video feed in real time.
[0554] Capturing a video feed
[0555] The server captures each user's video feed, converts it into an analyzable format, and stores it. This video feed is the key data used for face detection and facial expression and emotion analysis.
[0556] Face detection and facial expression analysis
[0557] The server analyzes the captured video feed and detects each participant's face using a facial recognition means, and acquires facial expression data of the detected faces using an expression analysis means, including detailed facial expression information such as eyebrow movements, eye opening and closing, and mouth movements.
[0558] Detecting embarrassment and emotions
[0559] The server uses both facial expression analysis and emotion recognition to determine whether the participant is confused or not. In addition, it simultaneously analyzes the user's emotions (happiness, sadness, anger, etc.) and obtains detailed emotional data.
[0560] Identifying confusing words
[0561] The server analyzes the conversation based on the identified confusion state and its timestamp, and identifies words and concepts that are the subject of confusion. This analysis process uses speech recognition and natural language processing technologies.
[0562] Generate word descriptions
[0563] The server uses a generator to automatically generate explanations for the identified words and concepts, and the generated explanations are provided to the participants in a concise and easy-to-understand format.
[0564] Displaying explanations and emotional feedback
[0565] The server sends the generated explanation to the confused user's device via a display means, and may provide similar feedback to other participants in some cases. It may also provide feedback regarding the progress of the meeting based on the emotion recognition results.
[0566] Specific examples
[0567] For example, if User B looks confused when the term "ecosystem" is used during a meeting:
[0568] 1. The server detects User B's confusion state and timestamp.
[0569] 2. Using facial expression analysis and emotion recognition means, we recognize that User B is not only confused but also a little nervous.
[0570] 3. Analyze the conversation and identify that the word "ecosystem" is the source of the question.
[0571] 4. Use the generative tools to generate the statement, "An ecosystem is a collection of interconnected organisms and their surrounding environment."
[0572] 5. This description may pop up on User B's screen and be shared with other participants.
[0573] 6. Based on emotional feedback, advice and tips on how to proceed with the meeting are also provided.
[0574] This series of processes quickly resolves any questions or confusion that meeting participants may have, allowing everyone to participate comfortably in the meeting. Feedback based on emotional data can also be used to improve the atmosphere of the meeting.
[0575] The processing flow will be explained below.
[0576] Step 1:
[0577] The server loads models and libraries for face recognition, facial expression analysis, confusion state detection, generation, display, and emotion recognition, performs necessary initial settings, and prepares to establish a connection with the online conference tool.
[0578] Step 2:
[0579] The server captures each participant's video feed in real time through the online conferencing tool, converts this data into a parsable format, and prepares it for processing.
[0580] Step 3:
[0581] The server analyzes each frame of the captured video feed and uses facial recognition to detect participants' faces, which are then stored for use in the next step.
[0582] Step 4:
[0583] The server acquires the facial expression data detected using the facial expression analysis means, and analyzes detailed facial expression information such as eyebrow movements, eye opening and closing, and mouth movements.
[0584] Step 5:
[0585] The server uses emotion recognition means to analyze facial expression data and other biometric signals to identify the user's emotions (happiness, sadness, anger, confusion, etc.), thereby determining with high accuracy whether the user is confused.
[0586] Step 6:
[0587] If the server detects a confused state, it records the timestamp and the face ID of the user. Based on this information, the next step is to analyze the conversation.
[0588] Step 7:
[0589] The server analyzes the audio data of the meeting based on the recorded timestamps, identifies words or concepts that are confusing, converts the audio data into text using speech recognition technology, and extracts keywords using natural language processing technology.
[0590] Step 8:
[0591] The server uses a generator to automatically generate explanations for the identified words and concepts, which are then summarized in a concise and easy-to-understand format.
[0592] Step 9:
[0593] The server sends the generated explanation to the confused user's terminal and displays it in a pop-up format through a display means, and provides similar feedback to other participants as needed.
[0594] Step 10:
[0595] Users can review explanations displayed on their devices to gain a better understanding of words or concepts they were confused about, allowing them to keep up with the progress of the meeting.
[0596] Step 11:
[0597] The server uses the emotion data to provide feedback and advice on how the meeting should proceed. For example, if the user is nervous, it will encourage them to relax, and if they are tired, it will suggest taking a break.
[0598] Step 12:
[0599] The server runs this process continuously for each frame of the video feed, detecting and addressing confusion and emotions in real time, allowing for immediate response as new questions or emotional states arise.
[0600] Example 2
[0601] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0602] In online meetings, participants often become confused by technical terms and complex concepts they do not understand. This confusion can hinder the progress of the meeting and reduce participant satisfaction. It is also problematic when the meeting proceeds without noticing that some participants are confused. Furthermore, if participants' emotional states cannot be addressed, the meeting atmosphere becomes awkward and productivity decreases. An efficient solution to this problem is needed.
[0603] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0604] In this invention, the server includes face recognition means, facial expression analysis means, confusion detection means, generation means, display means, emotion recognition means, and conversation content analysis means, which enable the server to quickly detect participants' confusion, provide appropriate explanations, and analyze participants' emotional states in real time to provide feedback.
[0605] "Facial recognition means" is a technology for identifying the faces of each user participating in an online conference.
[0606] The "facial expression analysis means" is a technology for analyzing the facial expressions of each detected user and acquiring facial expression data.
[0607] The "confusion state detection means" is a technique for determining whether the user is confused or not based on the facial expression data obtained by the facial expression analysis means.
[0608] A "generator" is a technology that automatically generates explanations for identified words or concepts.
[0609] The "display means" is a technique for displaying the generated explanation on the user's terminal.
[0610] The "emotion recognition means" is a technology that analyzes the user's emotional state (joy, sadness, anger, etc.) based on the facial expression data obtained by the facial expression analysis means.
[0611] The "conversation content analysis means" is a technology that analyzes the conversation content before and after the time when a confused state is detected, and identifies the specific words or phrases that cause the confusion.
[0612] The present invention provides a system for quickly detecting participant confusion in an online conference and providing appropriate explanations. This system further incorporates an emotion engine that recognizes user emotions, thereby improving the efficiency of the conference and participant satisfaction. A specific embodiment of the system is described below.
[0613] The server first loads models and libraries for face recognition, facial expression analysis, confusion detection, generation, display, and emotion recognition. This process uses machine learning libraries such as TensorFlow and Keras. It also establishes a connection with online conferencing tools (e.g., Zoom, Microsoft Teams) and prepares to capture each user's video feed in real time. At this time, it obtains the necessary API keys and access tokens.
[0614] The server then captures each user's video feed and converts it into a format that can be analyzed using a conversion tool such as FFmpeg. This video data is crucial for facial recognition and facial expression analysis.
[0615] The server analyzes the captured video feed frame by frame and detects each participant's face using the OpenCV library. It then uses the dlib library to extract facial feature points and perform a detailed analysis of the movements of each part of the face. Based on this data, the facial expression analyzer performs a detailed analysis of the user's facial expressions.
[0616] Furthermore, the server uses the Emotion Recognition API to analyze the user's emotional state based on the acquired facial expression data, for example, determining whether the participant is confused or expressing emotions such as joy, sadness, or anger.
[0617] If a confusion state is detected, the server analyzes the conversation content based on the timestamp. The conversation content is converted into text using online conferencing tools or speech recognition technology (e.g., Google Cloud Speech-to-Text API), and analyzed with a natural language processing engine (e.g., the BERT model) to identify words and phrases that cause confusion.
[0618] The server then uses a generative method to automatically generate explanations for the identified words and phrases. Specifically, it inputs the prompt "What is an ecosystem?" into a generative AI model such as GPT-3 and uses the generated answer to create an explanation. For example, it generates an explanation such as "An ecosystem is a collection of interconnected organisms and their surrounding environment."
[0619] Finally, the server sends the generated explanation to the user's device and displays it as a pop-up on the display. In some cases, the same explanation can be shared with other participants. The server also provides feedback on the progress of the meeting based on the emotion recognition results. For example, it may give advice such as, "The participants seem nervous, so it would be more effective if you spoke a little more slowly."
[0620] This series of processes quickly resolves any questions or confusion that meeting participants may have, allowing everyone to participate comfortably in the meeting. Feedback based on emotional data can also be used to improve the atmosphere of the meeting.
[0621] Prompt Sentence Examples
[0622] "What is an ecosystem?"
[0623] "Please provide additional explanation for the confusion."
[0624] "Please analyze participants' facial expression data and perform emotion recognition."
[0625] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0626] Step 1:
[0627] The server loads models and libraries for face recognition, facial expression analysis, confusion detection, generation, display, and emotion recognition. Specifically, it loads the necessary models into memory using machine learning libraries such as TensorFlow and Keras. It also establishes a connection using the online conference tool's API and prepares to obtain each user's video feed. The input is a request to load the models and libraries, and the output is the results of those loads.
[0628] Step 2:
[0629] The server captures each user's video feed in real time. Specifically, it obtains the video stream through the API of a conferencing tool such as Zoom or Microsoft Teams, and converts it into an analyzable format (e.g., MP4) using FFmpeg. This video data is used for subsequent processing. The input is each user's video feed, and the output is the video data converted into an analyzable format.
[0630] Step 3:
[0631] The server analyzes the captured video feed frame by frame and detects each participant's face using the OpenCV library. It then uses the dlib library to extract facial feature points and perform detailed analysis of eyebrow, eye, and mouth movements. The input is each frame of video data, and the output is facial data with extracted feature points.
[0632] Step 4:
[0633] The server uses facial expression analysis means to obtain facial expression data based on the acquired facial data. Specifically, it uses the Emotion Recognition API to analyze the user's emotional state based on information such as eyebrow movements, eye opening and closing, and mouth movements. The input is facial data, and the output is detailed data on each facial expression and the results of that emotion analysis.
[0634] Step 5:
[0635] The server uses the confusion detection means to determine the user's confusion state. Based on the data obtained by the facial expression analysis means, it applies a specific algorithm (e.g., random forest or support vector machine) to determine whether the user is confused. The input is facial expression data and the emotion analysis result, and the output is the confusion state determination result.
[0636] Step 6:
[0637] The server analyzes the conversation content based on the timestamp of the user who was detected as confused. It acquires the audio data using the conferencing tool's API and converts the audio into text using the Google Cloud Speech-to-Text API. It then analyzes the text data using a natural language processing engine such as the BERT model to identify words and phrases that cause confusion. The input is the audio data, and the output is the text data and the identified words and phrases.
[0638] Step 7:
[0639] The server automatically generates explanations for the identified words and phrases using a generation method. Specifically, a prompt such as "What is an ecosystem?" is input to a generative AI model such as GPT-3, and an explanation is generated based on the response. The input is the prompt, and the output is the generated explanation.
[0640] Step 8:
[0641] The server sends the generated explanation to the confused user's device and displays it as a pop-up using a display device. It also shares the same explanation with other participants as needed. It also provides feedback on the progress of the meeting based on the emotion recognition results. The inputs are the generated explanation and the emotion analysis results, and the outputs are the display of the explanation and the provision of feedback.
[0642] (Application example 2)
[0643] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0644] Conventional online conference systems and factory work support systems have the problem of making it difficult to quickly detect situations and provide appropriate feedback when workers or meeting participants are confused or emotionally troubled. In particular, in factories, confusion can lead to reduced work efficiency and incorrect work procedures, which can also affect safety. To solve this problem, the present invention provides a system that detects a worker's emotions and confusion in real time and provides appropriate support.
[0645] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes a face recognition means, a facial expression analysis means, an emotion recognition means, a confusion state detection means, a generation means, a display means, a video feed acquisition means, a real-time processing means, a means for generating an explanation using a generative AI model, and a means for generating an explanation using a prompt sentence. This makes it possible to detect in real time that a worker is confused, and to instantly generate and display appropriate explanations regarding specific work procedures. This makes it possible to improve work efficiency and ensure safety.
[0646] "Facial recognition tools" are technologies and devices used to detect an individual's face from a video feed and identify that person.
[0647] "Facial expression analysis means" refers to technology and devices for analyzing subtle changes in facial expressions from a detected face and identifying the person's emotional state.
[0648] "Emotion recognition means" refers to technology and devices that use facial expression analysis data to determine a person's emotions (joy, sadness, anger, confusion, etc.).
[0649] The "confusion state detection means" refers to a technique and device for identifying whether a person is confused or not based on emotion recognition and facial expression analysis.
[0650] The "generation means" refers to technology and devices for automatically generating necessary explanations and information based on the identified confusion state and emotions.
[0651] "Display means" refers to the technology and devices used to visually present the generated explanations and information to the user.
[0652] "Video Feed Acquisition Measures" means the technology and equipment used to capture video footage in real time and convert and store that data in an analyzable format.
[0653] "Real-time processing means" refers to techniques and devices that rapidly process captured video feeds and analytical data to provide immediate feedback.
[0654] "Means for generating explanations using generative AI models" refers to technologies and devices that use artificial intelligence models to automatically generate explanations based on specific situations and contexts.
[0655] The "means for generating an explanation using a prompt sentence" refers to a technique and device that allows an artificial intelligence model to generate an appropriate explanation based on a specific prompt sentence given as input.
[0656] The present invention is a system for detecting the emotions and confusion of factory workers in real time and providing appropriate support. Specific embodiments of the system will be described below.
[0657] System Configuration
[0658] The server includes a facial recognition means, a facial expression analysis means, an emotion recognition means, a confusion state detection means, a generation means, a display means, a video feed acquisition means, a real-time processing means, a means for generating an explanation using a generative AI model, and a means for generating an explanation using a prompt sentence.
[0659] Hardware and software used
[0660] This system uses the following hardware and software:
[0661] Camera: Mounted on the robot, it captures a video feed of the worker in real time.
[0662] Server: Processes the video feed and performs facial recognition, facial expression analysis, and emotion recognition.
[0663] dlib: A library for face detection.
[0664] EmotionEngine: A proprietary module for facial expression analysis and emotion recognition.
[0665] TextGenerator: A module for generating explanations using generative AI models.
[0666] Data processing and calculation
[0667] The server processes and calculates the data as follows:
[0668] 1. Capturing and saving video feed:
[0669] The robot's camera captures a real-time video feed of the worker, converts it into an analyzable format, and stores it.
[0670] 2. Facial Recognition and Expression Analysis:
[0671] The dlib library is used to detect the worker's face from the video feed, and the EmotionEngine is used to obtain detailed facial expression data.
[0672] 3. Detecting embarrassment and emotions:
[0673] The EmotionEngine analyzes facial expression data to determine whether the worker is confused and obtains emotional data (e.g., tension, anxiety, joy, etc.).
[0674] 4. Identify the task steps that cause confusion:
[0675] Based on the identified confusion state and its timestamp, the work content is analyzed to identify the part that causes the confusion.
[0676] 5. Generating explanations using generative means:
[0677] Using TextGenerator, a prompt sentence is input into the generative AI model to automatically generate an explanation for the cause of confusion.
[0678] 6. Displaying explanatory and emotional feedback:
[0679] The generated explanation is displayed on the robot's display, providing visual assistance to the worker.
[0680] Specific examples
[0681] As an example, here is the process that would occur if a worker becomes confused while performing "Step 5" in a factory.
[0682] 1. The server captures the worker's video feed through the robot's camera.
[0683] 2. The server analyzes the video feed and detects the worker's face using the dlib library.
[0684] 3. Use EmotionEngine to analyze facial expression and emotion data and recognize when the worker is confused.
[0685] 4. Using the timestamps from the video feed of the task, identify the task step ("Step 5") that caused the confusion.
[0686] 5. Using TextGenerator, enter the "Explanation of points where mistakes are likely to occur in step 5" as a prompt sentence into the generative AI model.
[0687] 6. The generated instructions are displayed on the robot's display, allowing the worker to refer to them and continue the work.
[0688] Example prompt sentence:
[0689] "Explanation of points where mistakes are likely to occur in step 5"
[0690] Through the above series of processes, it is possible to detect situations in which a worker is confused in real time and provide appropriate support, thereby improving work efficiency and reducing worker stress.
[0691] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0692] Step 1:
[0693] The server captures the worker's video feed in real time using the robot's camera. The input is the video footage obtained from the camera, and the output is the captured video feed. The server converts it into an analyzable format and stores it.
[0694] Step 2:
[0695] The server uses the dlib library to detect the worker's face from the video feed. The input is the video feed saved in step 1, and the output is the location information (coordinates) of the detected face. The server performs face recognition processing based on this location information.
[0696] Step 3:
[0697] The server uses the EmotionEngine to analyze facial expressions. The input is the facial position information and video feed obtained in step 2, and the output is detailed facial expression data. The server analyzes eyebrow movements, eye opening and closing, mouth movements, etc.
[0698] Step 4:
[0699] The server further uses the EmotionEngine to perform emotion recognition. The input is the facial expression data obtained in step 3, and the output is the worker's emotional data (e.g., joy, sadness, anger, confusion, etc.). The server determines whether the worker is confused.
[0700] Step 5:
[0701] The server analyzes the work content based on the identified confusion state and its timestamp. The input is the timestamp of the confusion state and the timestamp of the video feed, and the output is the specific work step that caused the confusion. The server analyzes the conversation and work content to identify the cause of the confusion.
[0702] Step 6:
[0703] The server uses TextGenerator to input an explanation of the cause of confusion as a prompt sentence into the generative AI model, which then generates an appropriate explanation. The input is a prompt sentence related to the identified work step (e.g., "An explanation of the points in step 5 where mistakes are likely to occur"), and the output is the generated explanation sentence. The server inputs this prompt sentence into the generative AI model to obtain the necessary explanation.
[0704] Step 7:
[0705] The server displays the generated explanation on the robot's display. The input is the explanatory text obtained in step 6, and the output is visual support information displayed on the display. The server immediately provides appropriate explanations when a worker becomes confused, thereby improving work efficiency and ensuring safety.
[0706] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0707] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0708] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0709] [Third embodiment]
[0710] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0711] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[0712] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0713] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0714] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0715] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0716] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0717] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0718] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0719] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0720] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0721] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0722] The present invention describes a system for quickly identifying words or concepts that confuse online meeting participants and providing explanations for them. A specific embodiment of this system is described below.
[0723] System initialization
[0724] The server starts the system, which includes a facial recognition unit, a facial expression analysis unit, a confusion state detection unit, a generation unit, and a display unit, and loads the models and necessary libraries for each unit. The server also works with an online conference tool to prepare for capturing the user's video feed in real time.
[0725] Capturing a video feed
[0726] The server captures each user's video feed and converts it into an analyzable format, providing the data necessary for face detection and facial expression analysis of participants.
[0727] Face detection and facial expression analysis
[0728] The server analyzes each frame of the captured video feed, detects the faces of participants using a facial recognition means, and acquires facial expression data of the detected faces using an expression analysis means to determine whether the user is confused or confused.
[0729] Detecting confusion
[0730] If a participant is found to have a confused expression, the server records the user's face ID and the current timestamp, allowing it to determine the exact moment of confusion.
[0731] Identifying confusing words
[0732] The server analyzes the audio data from the online meeting based on the recorded timestamps and identifies the words or concepts that are causing confusion. This analysis uses speech recognition and natural language processing technologies.
[0733] Generate word descriptions
[0734] The server sends the identified words and concepts to a generator (e.g., a generative AI) that generates a concise and accurate explanation for them, presented in a format that is easy for the user to understand.
[0735] Display Description
[0736] Finally, the server displays the generated explanation on the confused user's device in a pop-up format for immediate review, and can provide the same explanation to other conference participants if necessary.
[0737] Specific examples
[0738] For example, if the term "ecosystem" is used during a meeting and User A finds it difficult to understand, the following will work:
[0739] 1. The server detects that User A has a confused expression.
[0740] 2. Analyze the conversation based on timestamps and identify that the word "ecosystem" is the source of confusion.
[0741] 3. Use the generative tool to generate the statement, "An ecosystem is a collection of interconnected organisms and their surrounding environment."
[0742] 4. This explanation will be displayed in real time on User A's screen.
[0743] This allows the question of User A to be resolved immediately, allowing the flow of the meeting to proceed without being interrupted. This system also works effectively when other participants have the same problem.
[0744] The processing flow will be explained below.
[0745] Step 1:
[0746] The server loads models and libraries for face recognition, facial expression analysis, confusion detection, generation, and display. It also establishes a connection with the online conference tool and prepares to acquire the video feed.
[0747] Step 2:
[0748] The server captures each user's video feed and converts it into a format that can be analyzed in real time, thereby continuously receiving video data from participants.
[0749] Step 3:
[0750] The server analyzes each frame of the captured video feed and detects participants' faces using facial recognition techniques, and the detected face information is recorded for each frame.
[0751] Step 4:
[0752] The server uses an expression analysis means to obtain facial expression data of the detected face, which includes detailed facial expression information such as eyebrow movements, eye opening and closing, and mouth movements.
[0753] Step 5:
[0754] The server analyzes the facial expression data acquired by the server using a confusion detection method to identify the participant's confused expression. The server records the timestamp and face ID of the identified confused expression.
[0755] Step 6:
[0756] The server analyzes the audio data of the conversation based on the recorded timestamps to identify words or concepts that may be confusing, and uses speech recognition technology to convert the audio data into text.
[0757] Step 7:
[0758] The server uses a generator to automatically generate explanations for the identified words and concepts in a concise and easy-to-understand format.
[0759] Step 8:
[0760] The server sends the generated explanation to the terminal of the confused participant via a display means. The explanation is displayed in a pop-up format so that the participant can immediately check it.
[0761] Step 9:
[0762] Users can review explanations displayed on their devices to gain a better understanding of words or concepts they were confused about, allowing them to keep up with the progress of the meeting.
[0763] Step 10:
[0764] The server continuously performs this process for each frame of the video feed, detecting and addressing confusion in real time, allowing for immediate response if new questions arise during the meeting.
[0765] Example 1
[0766] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0767] In online meetings, participants may be confused by the technical terms and concepts used, which can cause communication to be disrupted. This confusion reduces the effectiveness of the meeting and hinders participants' understanding. Therefore, a system that can instantly provide easy-to-understand explanations for content that participants are confused about during a meeting is desirable.
[0768] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0769] In this invention, the server includes a face recognition means, an expression analysis means, a confusion detection means, a generation means, a display means, a video feed acquisition means, an audio data analysis means, a natural language processing means, and a pop-up display means, which enable the server to detect when a participant shows a confused expression, identify target words or concepts based on that timing, and provide explanations to resolve the confusion in real time.
[0770] "Facial Recognition Method" refers to technology for detecting participants' faces within a video feed and analyzing their features.
[0771] The "facial expression analysis means" is a technology for analyzing detected facial expression data and determining the emotional state of the participant.
[0772] The "confusion state detection means" is a technology that uses facial expression analysis means to determine whether a participant is confused and detects that state.
[0773] "Generative means" refers to technology for generating explanations for identified words or concepts, including generative AI models.
[0774] "Display means" refers to a technique for presenting the generated explanation to participants.
[0775] The "video feed acquisition means" is a technology that acquires a video feed from an online conference tool and converts it into an analyzable format.
[0776] "Audio data analysis means" is a technology that analyzes the audio data of a conference and converts it into text.
[0777] "Natural language processing means" is a technology that analyzes acquired text data and identifies words and concepts that cause confusion.
[0778] "Pop-up display means" is a technique for visually presenting the generated explanation to participants in real time.
[0779] The present invention is a system that instantly identifies words or concepts that participants in an online meeting find confusing and provides explanations for them. This system is built using a combination of specific hardware and software.
[0780] Hardware and software used
[0781] The server runs the system using the following hardware and software:
[0782] Face recognition method (e.g. OpenCV)
[0783] Facial expression analysis means (e.g. Facial Expression Recognition)
[0784] Confusion detection means
[0785] Generation method (e.g. GPT-3)
[0786] Display means
[0787] Video feed acquisition method (e.g. Zoom, Teams)
[0788] Voice data analysis method (e.g., Google Speech-to-Text API)
[0789] Natural language processing tools (e.g., NLTK, Transformers)
[0790] Pop-up display method
[0791] Program processing overview
[0792] The server starts the various components of the system and loads the necessary models and libraries, including the facial recognition model, facial expression analysis model, confusion detection model, natural language processing library, and generative AI model. The server also connects with online meeting tools to obtain user video feeds.
[0793] The server captures each user's video feed in real time via the online conferencing tool, converts the video feed into an analyzable format, and uses it as data for facial recognition and facial expression analysis.
[0794] The server sequentially analyzes each frame of the captured video feed and detects the participant's face using a facial recognition model, then uses an expression analysis model to obtain facial expression data from the detected face and determine whether the user is confused or not.
[0795] If the user shows a confused expression, the server records the confused user's face ID and the current timestamp, allowing the specific moment of confusion to be identified.
[0796] The server analyzes the audio data of online meetings based on the recorded timestamps, converts the audio data into text using speech recognition technology, and then uses natural language processing technology to extract words or concepts that are causing confusion.
[0797] The server uses a generative AI model to generate an explanation for the identified word or concept: for example, for "ecosystem," it generates the explanation "an ecosystem is a collection of interconnected organisms and their surrounding environment."
[0798] Finally, the server sends the generated explanation to the confused user's terminal, where it displays the explanation in a pop-up format and can provide the same explanation to other conference participants.
[0799] Specific examples
[0800] For example, if the term "ecosystem" is used during a meeting and it is difficult for User A to understand, the following will work:
[0801] 1. The server detects that User A has a confused expression.
[0802] 2. Analyze the conversation based on timestamps and identify that the word "ecosystem" is the source of confusion.
[0803] 3. The server uses a generative AI model to generate the explanation that "an ecosystem is a collection of interconnected organisms and their surrounding environments."
[0804] 4. This explanation will be displayed in real time on User A's screen.
[0805] Prompt Sentence Examples
[0806] "The term 'ecosystem' comes up during a meeting and the user looks confused. Generate a concise explanation of this term and display it to the user in a popup."
[0807] This system makes it possible to instantly resolve any confusion participants may have during online meetings and ensure smooth communication.
[0808] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0809] Step 1: Initialize the system
[0810] The server initiates the system's various functions, including preparing to load models and libraries for facial recognition, facial expression analysis, confusion detection, generative AI models, video feed acquisition, audio data analysis, and display, such as OpenCV, Facial Expression Recognition, NLTK, Transformers, and GPT-3.
[0811] Input: Models and libraries of various methods
[0812] Output: Initialized system
[0813] Step 2: Capture and convert your video feed
[0814] The server captures the user's video feed in real time through the online conferencing tool and converts it into a format that can be analyzed: the video feed is converted into a series of JPEG frames.
[0815] Input: Video feed from online meeting tool
[0816] Output: Video frames converted to a parsable format
[0817] Step 3: Face detection
[0818] The server analyzes each video frame and detects the participants' faces using a facial recognition method, such as OpenCV.
[0819] Input: Transformed video frames
[0820] Output: Detected face data
[0821] Step 4: Facial Expression Analysis
[0822] The server uses the facial expression analysis model to analyze the facial expressions detected by the facial recognition means, thereby determining the emotional state of the participant.
[0823] Input: Detected face data
[0824] Output: Analyzed facial expression data
[0825] Step 5: Detecting confusion
[0826] If a participant shows a confused expression, the server records the user's face ID and the current timestamp, allowing the specific moment of confusion to be identified.
[0827] Input: Analyzed facial expression data
[0828] Output: Confusion detection result (face ID and timestamp)
[0829] Step 6: Analyze the audio data
[0830] The server analyzes the audio data of the online meeting based on the recorded timestamps, and the audio data is converted into text using speech recognition technology.
[0831] Input: Online meeting audio data and timestamps
[0832] Output: Audio data converted to text
[0833] Step 7: Identify the confusing word
[0834] The server analyzes the text of the speech data and identifies confusing words and concepts using natural language processing means.
[0835] Input: Audio data converted to text
[0836] Output: Words or concepts that cause confusion
[0837] Step 8: Generate word descriptions
[0838] The server sends the identified words and concepts to a generative AI model to generate a concise and accurate description of them. For example, a generative AI model can be used to generate a description of an "ecosystem."
[0839] Input: Identified words or concepts
[0840] Output: Generated description
[0841] Step 9: View the description
[0842] The server sends the generated explanation to the confused user's device, which displays it in a pop-up format. If necessary, the same explanation can be provided to other conference participants.
[0843] Input: Generated Description
[0844] Output: Description displayed on the user's terminal
[0845] This series of processes makes it possible to provide an appropriate explanation immediately and support smooth communication even if a user becomes confused during a meeting.
[0846] (Application example 1)
[0847] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0848] In modern factories, workers are required to quickly adapt to new machines and technologies to maintain or increase productivity. However, if workers are confused by technical terms and operating procedures, time is wasted and work efficiency is reduced. To address this issue, a system that allows workers to instantly resolve their questions is needed. Therefore, a system and method that allows workers to obtain explanations of technical terms and operating procedures in real time is needed.
[0849] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0850] In this invention, the server includes a face recognition means, an expression analysis means, a confusion state detection means, a generation means, a display means, a real-time video feed acquisition means, an audio analysis means, and a visual display means for displaying the generated explanation, thereby enabling visual explanations of technical terms and operating procedures to be provided in real time when a worker is confused.
[0851] "Facial recognition tools" are technologies for detecting faces in video feeds and identifying individual people.
[0852] "Facial expression analysis means" is a technology for analyzing recognized facial expressions and inferring the emotional state of the person.
[0853] The "confusion state detection means" is a technology that determines whether the user is confused or not based on the results of facial expression analysis.
[0854] A "generative means" is a technology for generating specific information or explanations, which in this case is a generative AI model.
[0855] "Display means" refers to a technique for visually displaying the generated information and explanations to the user.
[0856] "Real-time video feed acquisition means" refers to technology for acquiring and processing video of the current user in real time.
[0857] "Speech analysis means" refers to technology that analyzes speech data in real time and recognizes specific words and phrases.
[0858] "Visual display means" refers to the display of smart glasses or other devices and the technology used to display information thereon.
[0859] A "generative AI model" is an artificial intelligence model that generates appropriate content based on input information.
[0860] The present invention is a system for providing appropriate explanations in real time to factory workers so that they can continue working without being confused by technical terms and operating procedures. The system includes the following means.
[0861] The server includes a facial recognition means, an expression analysis means, a confusion detection means, a generation means, a display means, a real-time video feed acquisition means, an audio analysis means and a visual display means.
[0862] Hardware and software used
[0863] For facial recognition and facial expression analysis, Google MediaPipe and Azure Face API are used, for example. Real-time video feeds are acquired using the smart glasses' camera. Voice data analysis is performed using Google Cloud Speech-to-Text and Amazon Transcribe. Explanation generation utilizes generative AI models such as OpenAI's GPT-4, which generate appropriate explanations in real time. The generated explanations are then provided to the user using visual display means such as the smart glasses' display.
[0864] Data processing and data calculation
[0865] The server processes various data in real time. When a video feed is sent from the smart glasses, the face of the worker is identified using a facial recognition means, and confusion is detected using an expression analysis means. If a confusion state is detected, the voice data is analyzed based on the timestamp to identify the technical term or operation procedure that is causing the confusion. The technical term or operation procedure is then sent to a generative AI model, which generates a concise explanation for it. This explanation is displayed in real time on the visual display means (the display of the smart glasses).
[0866] Examples of concrete examples and prompts
[0867] As a concrete example, if a factory worker is operating a new machine and is confused by the technical term "scan speed," the system works as follows:
[0868] 1. Detect when a worker has a confused expression.
[0869] 2. Analyze the conversation based on timestamps and identify that the technical term "scanning rate" is the source of confusion.
[0870] 3. Use a generative AI model to generate the explanation, "Scanning speed is the speed at which the scan head scans the surface. It is typically expressed in meters per second (m / s)."
[0871] 4. This explanation will be displayed in real time on the smart glasses display.
[0872] Example prompt sentence:
[0873] "Scanning speed is the speed at which the scan head scans the surface."
[0874] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0875] Step 1:
[0876] Getting a Video Feed
[0877] The device (smart glasses) captures a real-time video feed of the user and sends it to the server. The input is the video feed, which serves as the data source for facial recognition and facial expression analysis. The output is image data for analysis.
[0878] Step 2:
[0879] Facial Recognition and Expression Analysis
[0880] The server uses a facial recognition mechanism to identify faces in the video feed and a facial expression analysis mechanism to analyze the user's facial expressions. The input is image data from the video feed, and the output is the user's face ID and its facial expression data. Processing is performed using Google MediaPipe and Azure Face API.
[0881] Step 3:
[0882] Detecting confusion
[0883] The server uses a confusion detection means to determine whether the user is confused based on the results of the facial expression analysis. The input is facial expression data, and the output is the user's confusion state and its timestamp. If a confusion state is detected, the timestamp is recorded.
[0884] Step 4:
[0885] Analysis of audio data
[0886] The server uses voice analysis tools to analyze the audio data of online meetings based on the recorded timestamps and identify confusing technical terms and operating procedures. The input is the audio data and timestamps, and the output is text data of the technical terms and operating procedures. Google Cloud Speech-to-Text and Amazon Transcribe are used.
[0887] Step 5:
[0888] Generate Description
[0889] The server sends the identified technical terms and operating procedures to a generative AI model, which generates a concise explanation for them. The input is text data of the technical terms and operating procedures, and the output is explanatory text. The explanation is generated using OpenAI's GPT-4 or similar software.
[0890] Step 6:
[0891] Display Description
[0892] The server sends the generated explanation to the display of the terminal (smart glasses) and displays it to the user using a visual display means. The input is the explanation text, and the output is the information displayed in the user's field of view. In concrete terms, the smart glasses display displays "Scanning speed is the speed at which the scan head scans the surface. It is typically expressed in meters per second (m / s)."
[0893] By repeating this series of steps, workers can proceed with their work without confusion.
[0894] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0895] The present invention aims to further improve the efficiency of online meetings and the satisfaction of participants by incorporating an emotion engine that recognizes the emotions of users in addition to a system that quickly detects confusion among participants in online meetings and provides appropriate explanations. Specific embodiments of the present invention will be described below.
[0896] System initialization
[0897] The server loads models and libraries for face recognition, facial expression analysis, confusion detection, generation, display, and emotion recognition, performs necessary initialization, establishes a connection with the online conferencing tool, and prepares to capture each user's video feed in real time.
[0898] Capturing a video feed
[0899] The server captures each user's video feed, converts it into an analyzable format, and stores it. This video feed is the key data used for face detection and facial expression and emotion analysis.
[0900] Face detection and facial expression analysis
[0901] The server analyzes the captured video feed and detects each participant's face using a facial recognition means, and acquires facial expression data of the detected faces using an expression analysis means, including detailed facial expression information such as eyebrow movements, eye opening and closing, and mouth movements.
[0902] Detecting embarrassment and emotions
[0903] The server uses both facial expression analysis and emotion recognition to determine whether the participant is confused or not. In addition, it simultaneously analyzes the user's emotions (happiness, sadness, anger, etc.) and obtains detailed emotional data.
[0904] Identifying confusing words
[0905] The server analyzes the conversation based on the identified confusion state and its timestamp, and identifies words and concepts that are the subject of confusion. This analysis process uses speech recognition and natural language processing technologies.
[0906] Generate word descriptions
[0907] The server uses a generator to automatically generate explanations for the identified words and concepts, and the generated explanations are provided to the participants in a concise and easy-to-understand format.
[0908] Displaying explanations and emotional feedback
[0909] The server sends the generated explanation to the confused user's device via a display means, and may provide similar feedback to other participants in some cases. It may also provide feedback regarding the progress of the meeting based on the emotion recognition results.
[0910] Specific examples
[0911] For example, if User B looks confused when the term "ecosystem" is used during a meeting:
[0912] 1. The server detects User B's confusion state and timestamp.
[0913] 2. Using facial expression analysis and emotion recognition means, we recognize that User B is not only confused but also a little nervous.
[0914] 3. Analyze the conversation and identify that the word "ecosystem" is the source of the question.
[0915] 4. Use the generative tools to generate the statement, "An ecosystem is a collection of interconnected organisms and their surrounding environment."
[0916] 5. This description may pop up on User B's screen and be shared with other participants.
[0917] 6. Based on emotional feedback, advice and tips on how to proceed with the meeting are also provided.
[0918] This series of processes quickly resolves any questions or confusion that meeting participants may have, allowing everyone to participate comfortably in the meeting. Feedback based on emotional data can also be used to improve the atmosphere of the meeting.
[0919] The processing flow will be explained below.
[0920] Step 1:
[0921] The server loads models and libraries for face recognition, facial expression analysis, confusion state detection, generation, display, and emotion recognition, performs necessary initial settings, and prepares to establish a connection with the online conference tool.
[0922] Step 2:
[0923] The server captures each participant's video feed in real time through the online conferencing tool, converts this data into a parsable format, and prepares it for processing.
[0924] Step 3:
[0925] The server analyzes each frame of the captured video feed and uses facial recognition to detect participants' faces, which are then stored for use in the next step.
[0926] Step 4:
[0927] The server acquires the facial expression data detected using the facial expression analysis means, and analyzes detailed facial expression information such as eyebrow movements, eye opening and closing, and mouth movements.
[0928] Step 5:
[0929] The server uses emotion recognition means to analyze facial expression data and other biometric signals to identify the user's emotions (happiness, sadness, anger, confusion, etc.), thereby determining with high accuracy whether the user is confused.
[0930] Step 6:
[0931] If the server detects a confused state, it records the timestamp and the face ID of the user. Based on this information, the next step is to analyze the conversation.
[0932] Step 7:
[0933] The server analyzes the audio data of the meeting based on the recorded timestamps, identifies words or concepts that are confusing, converts the audio data into text using speech recognition technology, and extracts keywords using natural language processing technology.
[0934] Step 8:
[0935] The server uses a generator to automatically generate explanations for the identified words and concepts, which are then summarized in a concise and easy-to-understand format.
[0936] Step 9:
[0937] The server sends the generated explanation to the confused user's terminal and displays it in a pop-up format through a display means, and provides similar feedback to other participants as needed.
[0938] Step 10:
[0939] Users can review explanations displayed on their devices to gain a better understanding of words or concepts they were confused about, allowing them to keep up with the progress of the meeting.
[0940] Step 11:
[0941] The server uses the emotion data to provide feedback and advice on how the meeting should proceed. For example, if the user is nervous, it will encourage them to relax, and if they are tired, it will suggest taking a break.
[0942] Step 12:
[0943] The server runs this process continuously for each frame of the video feed, detecting and addressing confusion and emotions in real time, allowing for immediate response as new questions or emotional states arise.
[0944] Example 2
[0945] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0946] In online meetings, participants often become confused by technical terms and complex concepts they do not understand. This confusion can hinder the progress of the meeting and reduce participant satisfaction. It is also problematic when the meeting proceeds without noticing that some participants are confused. Furthermore, if participants' emotional states cannot be addressed, the meeting atmosphere becomes awkward and productivity decreases. An efficient solution to this problem is needed.
[0947] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0948] In this invention, the server includes face recognition means, facial expression analysis means, confusion detection means, generation means, display means, emotion recognition means, and conversation content analysis means, which enable the server to quickly detect participants' confusion, provide appropriate explanations, and analyze participants' emotional states in real time to provide feedback.
[0949] "Facial recognition means" is a technology for identifying the faces of each user participating in an online conference.
[0950] The "facial expression analysis means" is a technology for analyzing the facial expressions of each detected user and acquiring facial expression data.
[0951] The "confusion state detection means" is a technique for determining whether the user is confused or not based on the facial expression data obtained by the facial expression analysis means.
[0952] A "generator" is a technology that automatically generates explanations for identified words or concepts.
[0953] The "display means" is a technique for displaying the generated explanation on the user's terminal.
[0954] The "emotion recognition means" is a technology that analyzes the user's emotional state (joy, sadness, anger, etc.) based on the facial expression data obtained by the facial expression analysis means.
[0955] The "conversation content analysis means" is a technology that analyzes the conversation content before and after the time when a confused state is detected, and identifies the specific words or phrases that cause the confusion.
[0956] The present invention provides a system for quickly detecting participant confusion in an online conference and providing appropriate explanations. This system further incorporates an emotion engine that recognizes user emotions, thereby improving the efficiency of the conference and participant satisfaction. A specific embodiment of the system is described below.
[0957] The server first loads models and libraries for face recognition, facial expression analysis, confusion detection, generation, display, and emotion recognition. This process uses machine learning libraries such as TensorFlow and Keras. It also establishes a connection with online conferencing tools (e.g., Zoom, Microsoft Teams) and prepares to capture each user's video feed in real time. At this time, it obtains the necessary API keys and access tokens.
[0958] The server then captures each user's video feed and converts it into a format that can be analyzed using a conversion tool such as FFmpeg. This video data is crucial for facial recognition and facial expression analysis.
[0959] The server analyzes the captured video feed frame by frame and detects each participant's face using the OpenCV library. It then uses the dlib library to extract facial feature points and perform a detailed analysis of the movements of each part of the face. Based on this data, the facial expression analyzer performs a detailed analysis of the user's facial expressions.
[0960] Furthermore, the server uses the Emotion Recognition API to analyze the user's emotional state based on the acquired facial expression data, for example, determining whether the participant is confused or expressing emotions such as joy, sadness, or anger.
[0961] If a confusion state is detected, the server analyzes the conversation content based on the timestamp. The conversation content is converted into text using online conferencing tools or speech recognition technology (e.g., Google Cloud Speech-to-Text API), and analyzed with a natural language processing engine (e.g., the BERT model) to identify words and phrases that cause confusion.
[0962] The server then uses a generative method to automatically generate explanations for the identified words and phrases. Specifically, it inputs the prompt "What is an ecosystem?" into a generative AI model such as GPT-3 and uses the generated answer to create an explanation. For example, it generates an explanation such as "An ecosystem is a collection of interconnected organisms and their surrounding environment."
[0963] Finally, the server sends the generated explanation to the user's device and displays it as a pop-up on the display. In some cases, the same explanation can be shared with other participants. The server also provides feedback on the progress of the meeting based on the emotion recognition results. For example, it may give advice such as, "The participants seem nervous, so it would be more effective if you spoke a little more slowly."
[0964] This series of processes quickly resolves any questions or confusion that meeting participants may have, allowing everyone to participate comfortably in the meeting. Feedback based on emotional data can also be used to improve the atmosphere of the meeting.
[0965] Prompt Sentence Examples
[0966] "What is an ecosystem?"
[0967] "Please provide additional explanation for the confusion."
[0968] "Please analyze participants' facial expression data and perform emotion recognition."
[0969] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0970] Step 1:
[0971] The server loads models and libraries for face recognition, facial expression analysis, confusion detection, generation, display, and emotion recognition. Specifically, it loads the necessary models into memory using machine learning libraries such as TensorFlow and Keras. It also establishes a connection using the online conference tool's API and prepares to obtain each user's video feed. The input is a request to load the models and libraries, and the output is the results of those loads.
[0972] Step 2:
[0973] The server captures each user's video feed in real time. Specifically, it obtains the video stream through the API of a conferencing tool such as Zoom or Microsoft Teams, and converts it into an analyzable format (e.g., MP4) using FFmpeg. This video data is used for subsequent processing. The input is each user's video feed, and the output is the video data converted into an analyzable format.
[0974] Step 3:
[0975] The server analyzes the captured video feed frame by frame and detects each participant's face using the OpenCV library. It then uses the dlib library to extract facial feature points and perform detailed analysis of eyebrow, eye, and mouth movements. The input is each frame of video data, and the output is facial data with extracted feature points.
[0976] Step 4:
[0977] The server uses facial expression analysis means to obtain facial expression data based on the acquired facial data. Specifically, it uses the Emotion Recognition API to analyze the user's emotional state based on information such as eyebrow movements, eye opening and closing, and mouth movements. The input is facial data, and the output is detailed data on each facial expression and the results of that emotion analysis.
[0978] Step 5:
[0979] The server uses the confusion detection means to determine the user's confusion state. Based on the data obtained by the facial expression analysis means, it applies a specific algorithm (e.g., random forest or support vector machine) to determine whether the user is confused. The input is facial expression data and the emotion analysis result, and the output is the confusion state determination result.
[0980] Step 6:
[0981] The server analyzes the conversation content based on the timestamp of the user who was detected as confused. It acquires the audio data using the conferencing tool's API and converts the audio into text using the Google Cloud Speech-to-Text API. It then analyzes the text data using a natural language processing engine such as the BERT model to identify words and phrases that cause confusion. The input is the audio data, and the output is the text data and the identified words and phrases.
[0982] Step 7:
[0983] The server automatically generates explanations for the identified words and phrases using a generation method. Specifically, a prompt such as "What is an ecosystem?" is input to a generative AI model such as GPT-3, and an explanation is generated based on the response. The input is the prompt, and the output is the generated explanation.
[0984] Step 8:
[0985] The server sends the generated explanation to the confused user's device and displays it as a pop-up using a display device. It also shares the same explanation with other participants as needed. It also provides feedback on the progress of the meeting based on the emotion recognition results. The inputs are the generated explanation and the emotion analysis results, and the outputs are the display of the explanation and the provision of feedback.
[0986] (Application example 2)
[0987] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0988] Conventional online conference systems and factory work support systems have the problem of making it difficult to quickly detect situations and provide appropriate feedback when workers or meeting participants are confused or emotionally troubled. In particular, in factories, confusion can lead to reduced work efficiency and incorrect work procedures, which can also affect safety. To solve this problem, the present invention provides a system that detects a worker's emotions and confusion in real time and provides appropriate support.
[0989] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes a face recognition means, a facial expression analysis means, an emotion recognition means, a confusion state detection means, a generation means, a display means, a video feed acquisition means, a real-time processing means, a means for generating an explanation using a generative AI model, and a means for generating an explanation using a prompt sentence. This makes it possible to detect in real time that a worker is confused, and to instantly generate and display appropriate explanations regarding specific work procedures. This makes it possible to improve work efficiency and ensure safety.
[0990] "Facial recognition tools" are technologies and devices used to detect an individual's face from a video feed and identify that person.
[0991] "Facial expression analysis means" refers to technology and devices for analyzing subtle changes in facial expressions from a detected face and identifying the person's emotional state.
[0992] "Emotion recognition means" refers to technology and devices that use facial expression analysis data to determine a person's emotions (joy, sadness, anger, confusion, etc.).
[0993] The "confusion state detection means" refers to a technique and device for identifying whether a person is confused or not based on emotion recognition and facial expression analysis.
[0994] The "generation means" refers to technology and devices for automatically generating necessary explanations and information based on the identified confusion state and emotions.
[0995] "Display means" refers to the technology and devices used to visually present the generated explanations and information to the user.
[0996] "Video Feed Acquisition Measures" means the technology and equipment used to capture video footage in real time and convert and store that data in an analyzable format.
[0997] "Real-time processing means" refers to techniques and devices that rapidly process captured video feeds and analytical data to provide immediate feedback.
[0998] "Means for generating explanations using generative AI models" refers to technologies and devices that use artificial intelligence models to automatically generate explanations based on specific situations and contexts.
[0999] The "means for generating an explanation using a prompt sentence" refers to a technique and device that allows an artificial intelligence model to generate an appropriate explanation based on a specific prompt sentence given as input.
[1000] The present invention is a system for detecting the emotions and confusion of factory workers in real time and providing appropriate support. Specific embodiments of the system will be described below.
[1001] System Configuration
[1002] The server includes a facial recognition means, a facial expression analysis means, an emotion recognition means, a confusion state detection means, a generation means, a display means, a video feed acquisition means, a real-time processing means, a means for generating an explanation using a generative AI model, and a means for generating an explanation using a prompt sentence.
[1003] Hardware and software used
[1004] This system uses the following hardware and software:
[1005] Camera: Mounted on the robot, it captures a video feed of the worker in real time.
[1006] Server: Processes the video feed and performs facial recognition, facial expression analysis, and emotion recognition.
[1007] dlib: A library for face detection.
[1008] EmotionEngine: A proprietary module for facial expression analysis and emotion recognition.
[1009] TextGenerator: A module for generating explanations using generative AI models.
[1010] Data processing and calculation
[1011] The server processes and calculates the data as follows:
[1012] 1. Capturing and saving video feed:
[1013] The robot's camera captures a real-time video feed of the worker, converts it into an analyzable format, and stores it.
[1014] 2. Facial Recognition and Expression Analysis:
[1015] The dlib library is used to detect the worker's face from the video feed, and the EmotionEngine is used to obtain detailed facial expression data.
[1016] 3. Detecting embarrassment and emotions:
[1017] The EmotionEngine analyzes facial expression data to determine whether the worker is confused and obtains emotional data (e.g., tension, anxiety, joy, etc.).
[1018] 4. Identify the task steps that cause confusion:
[1019] Based on the identified confusion state and its timestamp, the work content is analyzed to identify the part that causes the confusion.
[1020] 5. Generating explanations using generative means:
[1021] Using TextGenerator, a prompt sentence is input into the generative AI model to automatically generate an explanation for the cause of confusion.
[1022] 6. Displaying explanatory and emotional feedback:
[1023] The generated explanation is displayed on the robot's display, providing visual assistance to the worker.
[1024] Specific examples
[1025] As an example, here is the process that would occur if a worker becomes confused while performing "Step 5" in a factory.
[1026] 1. The server captures the worker's video feed through the robot's camera.
[1027] 2. The server analyzes the video feed and detects the worker's face using the dlib library.
[1028] 3. Use EmotionEngine to analyze facial expression and emotion data and recognize when the worker is confused.
[1029] 4. Using the timestamps from the video feed of the task, identify the task step ("Step 5") that caused the confusion.
[1030] 5. Using TextGenerator, enter the "Explanation of points where mistakes are likely to occur in step 5" as a prompt sentence into the generative AI model.
[1031] 6. The generated instructions are displayed on the robot's display, allowing the worker to refer to them and continue the work.
[1032] Example prompt sentence:
[1033] "Explanation of points where mistakes are likely to occur in step 5"
[1034] Through the above series of processes, it is possible to detect situations in which a worker is confused in real time and provide appropriate support, thereby improving work efficiency and reducing worker stress.
[1035] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1036] Step 1:
[1037] The server captures the worker's video feed in real time using the robot's camera. The input is the video footage obtained from the camera, and the output is the captured video feed. The server converts it into an analyzable format and stores it.
[1038] Step 2:
[1039] The server uses the dlib library to detect the worker's face from the video feed. The input is the video feed saved in step 1, and the output is the location information (coordinates) of the detected face. The server performs face recognition processing based on this location information.
[1040] Step 3:
[1041] The server uses the EmotionEngine to analyze facial expressions. The input is the facial position information and video feed obtained in step 2, and the output is detailed facial expression data. The server analyzes eyebrow movements, eye opening and closing, mouth movements, etc.
[1042] Step 4:
[1043] The server further uses the EmotionEngine to perform emotion recognition. The input is the facial expression data obtained in step 3, and the output is the worker's emotional data (e.g., joy, sadness, anger, confusion, etc.). The server determines whether the worker is confused.
[1044] Step 5:
[1045] The server analyzes the work content based on the identified confusion state and its timestamp. The input is the timestamp of the confusion state and the timestamp of the video feed, and the output is the specific work step that caused the confusion. The server analyzes the conversation and work content to identify the cause of the confusion.
[1046] Step 6:
[1047] The server uses TextGenerator to input an explanation of the cause of confusion as a prompt sentence into the generative AI model, which then generates an appropriate explanation. The input is a prompt sentence related to the identified work step (e.g., "An explanation of the points in step 5 where mistakes are likely to occur"), and the output is the generated explanation sentence. The server inputs this prompt sentence into the generative AI model to obtain the necessary explanation.
[1048] Step 7:
[1049] The server displays the generated explanation on the robot's display. The input is the explanatory text obtained in step 6, and the output is visual support information displayed on the display. The server immediately provides appropriate explanations when a worker becomes confused, thereby improving work efficiency and ensuring safety.
[1050] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1051] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1052] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1053] [Fourth embodiment]
[1054] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1055] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1056] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1057] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1058] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1059] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1060] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1061] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1062] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1063] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1064] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1065] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1066] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1067] The present invention describes a system for quickly identifying words or concepts that confuse online meeting participants and providing explanations for them. A specific embodiment of this system is described below.
[1068] System initialization
[1069] The server starts the system, which includes a facial recognition unit, a facial expression analysis unit, a confusion state detection unit, a generation unit, and a display unit, and loads the models and necessary libraries for each unit. The server also works with an online conference tool to prepare for capturing the user's video feed in real time.
[1070] Capturing a video feed
[1071] The server captures each user's video feed and converts it into an analyzable format, providing the data necessary for face detection and facial expression analysis of participants.
[1072] Face detection and facial expression analysis
[1073] The server analyzes each frame of the captured video feed, detects the faces of participants using a facial recognition means, and acquires facial expression data of the detected faces using an expression analysis means to determine whether the user is confused or confused.
[1074] Detecting confusion
[1075] If a participant is found to have a confused expression, the server records the user's face ID and the current timestamp, allowing it to determine the exact moment of confusion.
[1076] Identifying confusing words
[1077] The server analyzes the audio data from the online meeting based on the recorded timestamps and identifies the words or concepts that are causing confusion. This analysis uses speech recognition and natural language processing technologies.
[1078] Generate word descriptions
[1079] The server sends the identified words and concepts to a generator (e.g., a generative AI) that generates a concise and accurate explanation for them, presented in a format that is easy for the user to understand.
[1080] Display Description
[1081] Finally, the server displays the generated explanation on the confused user's device in a pop-up format for immediate review, and can provide the same explanation to other conference participants if necessary.
[1082] Specific examples
[1083] For example, if the term "ecosystem" is used during a meeting and User A finds it difficult to understand, the following will work:
[1084] 1. The server detects that User A has a confused expression.
[1085] 2. Analyze the conversation based on timestamps and identify that the word "ecosystem" is the source of confusion.
[1086] 3. Use the generative tool to generate the statement, "An ecosystem is a collection of interconnected organisms and their surrounding environment."
[1087] 4. This explanation will be displayed in real time on User A's screen.
[1088] This allows the question of User A to be resolved immediately, allowing the flow of the meeting to proceed without being interrupted. This system also works effectively when other participants have the same problem.
[1089] The processing flow will be explained below.
[1090] Step 1:
[1091] The server loads models and libraries for face recognition, facial expression analysis, confusion detection, generation, and display. It also establishes a connection with the online conference tool and prepares to acquire the video feed.
[1092] Step 2:
[1093] The server captures each user's video feed and converts it into a format that can be analyzed in real time, thereby continuously receiving video data from participants.
[1094] Step 3:
[1095] The server analyzes each frame of the captured video feed and detects participants' faces using facial recognition techniques, and the detected face information is recorded for each frame.
[1096] Step 4:
[1097] The server uses an expression analysis means to obtain facial expression data of the detected face, which includes detailed facial expression information such as eyebrow movements, eye opening and closing, and mouth movements.
[1098] Step 5:
[1099] The server analyzes the facial expression data acquired by the server using a confusion detection method to identify the participant's confused expression. The server records the timestamp and face ID of the identified confused expression.
[1100] Step 6:
[1101] The server analyzes the audio data of the conversation based on the recorded timestamps to identify words or concepts that may be confusing, and uses speech recognition technology to convert the audio data into text.
[1102] Step 7:
[1103] The server uses a generator to automatically generate explanations for the identified words and concepts in a concise and easy-to-understand format.
[1104] Step 8:
[1105] The server sends the generated explanation to the terminal of the confused participant via a display means. The explanation is displayed in a pop-up format so that the participant can immediately check it.
[1106] Step 9:
[1107] Users can review explanations displayed on their devices to gain a better understanding of words or concepts they were confused about, allowing them to keep up with the progress of the meeting.
[1108] Step 10:
[1109] The server continuously performs this process for each frame of the video feed, detecting and addressing confusion in real time, allowing for immediate response if new questions arise during the meeting.
[1110] Example 1
[1111] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1112] In online meetings, participants may be confused by the technical terms and concepts used, which can cause communication to be disrupted. This confusion reduces the effectiveness of the meeting and hinders participants' understanding. Therefore, a system that can instantly provide easy-to-understand explanations for content that participants are confused about during a meeting is desirable.
[1113] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1114] In this invention, the server includes a face recognition means, an expression analysis means, a confusion detection means, a generation means, a display means, a video feed acquisition means, an audio data analysis means, a natural language processing means, and a pop-up display means, which enable the server to detect when a participant shows a confused expression, identify target words or concepts based on that timing, and provide explanations to resolve the confusion in real time.
[1115] "Facial Recognition Method" refers to technology for detecting participants' faces within a video feed and analyzing their features.
[1116] The "facial expression analysis means" is a technology for analyzing detected facial expression data and determining the emotional state of the participant.
[1117] The "confusion state detection means" is a technology that uses facial expression analysis means to determine whether a participant is confused and detects that state.
[1118] "Generative means" refers to technology for generating explanations for identified words or concepts, including generative AI models.
[1119] "Display means" refers to a technique for presenting the generated explanation to participants.
[1120] The "video feed acquisition means" is a technology that acquires a video feed from an online conference tool and converts it into an analyzable format.
[1121] "Audio data analysis means" is a technology that analyzes the audio data of a conference and converts it into text.
[1122] "Natural language processing means" is a technology that analyzes acquired text data and identifies words and concepts that cause confusion.
[1123] "Pop-up display means" is a technique for visually presenting the generated explanation to participants in real time.
[1124] The present invention is a system that instantly identifies words or concepts that participants in an online meeting find confusing and provides explanations for them. This system is built using a combination of specific hardware and software.
[1125] Hardware and software used
[1126] The server runs the system using the following hardware and software:
[1127] Face recognition method (e.g. OpenCV)
[1128] Facial expression analysis means (e.g. Facial Expression Recognition)
[1129] Confusion detection means
[1130] Generation method (e.g. GPT-3)
[1131] Display means
[1132] Video feed acquisition method (e.g. Zoom, Teams)
[1133] Voice data analysis method (e.g., Google Speech-to-Text API)
[1134] Natural language processing tools (e.g., NLTK, Transformers)
[1135] Pop-up display method
[1136] Program processing overview
[1137] The server starts the various components of the system and loads the necessary models and libraries, including the facial recognition model, facial expression analysis model, confusion detection model, natural language processing library, and generative AI model. The server also connects with online meeting tools to obtain user video feeds.
[1138] The server captures each user's video feed in real time via the online conferencing tool, converts the video feed into an analyzable format, and uses it as data for facial recognition and facial expression analysis.
[1139] The server sequentially analyzes each frame of the captured video feed and detects the participant's face using a facial recognition model, then uses an expression analysis model to obtain facial expression data from the detected face and determine whether the user is confused or not.
[1140] If the user shows a confused expression, the server records the confused user's face ID and the current timestamp, allowing the specific moment of confusion to be identified.
[1141] The server analyzes the audio data of online meetings based on the recorded timestamps, converts the audio data into text using speech recognition technology, and then uses natural language processing technology to extract words or concepts that are causing confusion.
[1142] The server uses a generative AI model to generate an explanation for the identified word or concept: for example, for "ecosystem," it generates the explanation "an ecosystem is a collection of interconnected organisms and their surrounding environment."
[1143] Finally, the server sends the generated explanation to the confused user's terminal, where it displays the explanation in a pop-up format and can provide the same explanation to other conference participants.
[1144] Specific examples
[1145] For example, if the term "ecosystem" is used during a meeting and it is difficult for User A to understand, the following will work:
[1146] 1. The server detects that User A has a confused expression.
[1147] 2. Analyze the conversation based on timestamps and identify that the word "ecosystem" is the source of confusion.
[1148] 3. The server uses a generative AI model to generate the explanation that "an ecosystem is a collection of interconnected organisms and their surrounding environments."
[1149] 4. This explanation will be displayed in real time on User A's screen.
[1150] Prompt Sentence Examples
[1151] "The term 'ecosystem' comes up during a meeting and the user looks confused. Generate a concise explanation of this term and display it to the user in a popup."
[1152] This system makes it possible to instantly resolve any confusion participants may have during online meetings and ensure smooth communication.
[1153] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1154] Step 1: Initialize the system
[1155] The server initiates the system's various functions, including preparing to load models and libraries for facial recognition, facial expression analysis, confusion detection, generative AI models, video feed acquisition, audio data analysis, and display, such as OpenCV, Facial Expression Recognition, NLTK, Transformers, and GPT-3.
[1156] Input: Models and libraries of various methods
[1157] Output: Initialized system
[1158] Step 2: Capture and convert your video feed
[1159] The server captures the user's video feed in real time through the online conferencing tool and converts it into a format that can be analyzed: the video feed is converted into a series of JPEG frames.
[1160] Input: Video feed from online meeting tool
[1161] Output: Video frames converted to a parsable format
[1162] Step 3: Face detection
[1163] The server analyzes each video frame and detects the participants' faces using a facial recognition method, such as OpenCV.
[1164] Input: Transformed video frames
[1165] Output: Detected face data
[1166] Step 4: Facial Expression Analysis
[1167] The server uses the facial expression analysis model to analyze the facial expressions detected by the facial recognition means, thereby determining the emotional state of the participant.
[1168] Input: Detected face data
[1169] Output: Analyzed facial expression data
[1170] Step 5: Detecting confusion
[1171] If a participant shows a confused expression, the server records the user's face ID and the current timestamp, allowing the specific moment of confusion to be identified.
[1172] Input: Analyzed facial expression data
[1173] Output: Confusion detection result (face ID and timestamp)
[1174] Step 6: Analyze the audio data
[1175] The server analyzes the audio data of the online meeting based on the recorded timestamps, and the audio data is converted into text using speech recognition technology.
[1176] Input: Online meeting audio data and timestamps
[1177] Output: Audio data converted to text
[1178] Step 7: Identify the confusing word
[1179] The server analyzes the text of the speech data and identifies confusing words and concepts using natural language processing means.
[1180] Input: Audio data converted to text
[1181] Output: Words or concepts that cause confusion
[1182] Step 8: Generate word descriptions
[1183] The server sends the identified words and concepts to a generative AI model to generate a concise and accurate description of them. For example, a generative AI model can be used to generate a description of an "ecosystem."
[1184] Input: Identified words or concepts
[1185] Output: Generated description
[1186] Step 9: View the description
[1187] The server sends the generated explanation to the confused user's device, which displays it in a pop-up format. If necessary, the same explanation can be provided to other conference participants.
[1188] Input: Generated Description
[1189] Output: Description displayed on the user's terminal
[1190] This series of processes makes it possible to provide an appropriate explanation immediately and support smooth communication even if a user becomes confused during a meeting.
[1191] (Application example 1)
[1192] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1193] In modern factories, workers are required to quickly adapt to new machines and technologies to maintain or increase productivity. However, if workers are confused by technical terms and operating procedures, time is wasted and work efficiency is reduced. To address this issue, a system that allows workers to instantly resolve their questions is needed. Therefore, a system and method that allows workers to obtain explanations of technical terms and operating procedures in real time is needed.
[1194] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1195] In this invention, the server includes a face recognition means, an expression analysis means, a confusion state detection means, a generation means, a display means, a real-time video feed acquisition means, an audio analysis means, and a visual display means for displaying the generated explanation, thereby enabling visual explanations of technical terms and operating procedures to be provided in real time when a worker is confused.
[1196] "Facial recognition tools" are technologies for detecting faces in video feeds and identifying individual people.
[1197] "Facial expression analysis means" is a technology for analyzing recognized facial expressions and inferring the emotional state of the person.
[1198] The "confusion state detection means" is a technology that determines whether the user is confused or not based on the results of facial expression analysis.
[1199] A "generative means" is a technology for generating specific information or explanations, which in this case is a generative AI model.
[1200] "Display means" refers to a technique for visually displaying the generated information and explanations to the user.
[1201] "Real-time video feed acquisition means" refers to technology for acquiring and processing video of the current user in real time.
[1202] "Speech analysis means" refers to technology that analyzes speech data in real time and recognizes specific words and phrases.
[1203] "Visual display means" refers to the display of smart glasses or other devices and the technology used to display information thereon.
[1204] A "generative AI model" is an artificial intelligence model that generates appropriate content based on input information.
[1205] The present invention is a system for providing appropriate explanations in real time to factory workers so that they can continue working without being confused by technical terms and operating procedures. The system includes the following means.
[1206] The server includes a facial recognition means, an expression analysis means, a confusion detection means, a generation means, a display means, a real-time video feed acquisition means, an audio analysis means and a visual display means.
[1207] Hardware and software used
[1208] For facial recognition and facial expression analysis, Google MediaPipe and Azure Face API are used, for example. Real-time video feeds are acquired using the smart glasses' camera. Voice data analysis is performed using Google Cloud Speech-to-Text and Amazon Transcribe. Explanation generation utilizes generative AI models such as OpenAI's GPT-4, which generate appropriate explanations in real time. The generated explanations are then provided to the user using visual display means such as the smart glasses' display.
[1209] Data processing and data calculation
[1210] The server processes various data in real time. When a video feed is sent from the smart glasses, the face of the worker is identified using a facial recognition means, and confusion is detected using an expression analysis means. If a confusion state is detected, the voice data is analyzed based on the timestamp to identify the technical term or operation procedure that is causing the confusion. The technical term or operation procedure is then sent to a generative AI model, which generates a concise explanation for it. This explanation is displayed in real time on the visual display means (the display of the smart glasses).
[1211] Examples of concrete examples and prompts
[1212] As a concrete example, if a factory worker is operating a new machine and is confused by the technical term "scan speed," the system works as follows:
[1213] 1. Detect when a worker has a confused expression.
[1214] 2. Analyze the conversation based on timestamps and identify that the technical term "scanning rate" is the source of confusion.
[1215] 3. Use a generative AI model to generate the explanation, "Scanning speed is the speed at which the scan head scans the surface. It is typically expressed in meters per second (m / s)."
[1216] 4. This explanation will be displayed in real time on the smart glasses display.
[1217] Example prompt sentence:
[1218] "Scanning speed is the speed at which the scan head scans the surface."
[1219] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1220] Step 1:
[1221] Getting a Video Feed
[1222] The device (smart glasses) captures a real-time video feed of the user and sends it to the server. The input is the video feed, which serves as the data source for facial recognition and facial expression analysis. The output is image data for analysis.
[1223] Step 2:
[1224] Facial Recognition and Expression Analysis
[1225] The server uses a facial recognition mechanism to identify faces in the video feed and a facial expression analysis mechanism to analyze the user's facial expressions. The input is image data from the video feed, and the output is the user's face ID and its facial expression data. Processing is performed using Google MediaPipe and Azure Face API.
[1226] Step 3:
[1227] Detecting confusion
[1228] The server uses a confusion detection means to determine whether the user is confused based on the results of the facial expression analysis. The input is facial expression data, and the output is the user's confusion state and its timestamp. If a confusion state is detected, the timestamp is recorded.
[1229] Step 4:
[1230] Analysis of audio data
[1231] The server uses voice analysis tools to analyze the audio data of online meetings based on the recorded timestamps and identify confusing technical terms and operating procedures. The input is the audio data and timestamps, and the output is text data of the technical terms and operating procedures. Google Cloud Speech-to-Text and Amazon Transcribe are used.
[1232] Step 5:
[1233] Generate Description
[1234] The server sends the identified technical terms and operating procedures to a generative AI model, which generates a concise explanation for them. The input is text data of the technical terms and operating procedures, and the output is explanatory text. The explanation is generated using OpenAI's GPT-4 or similar software.
[1235] Step 6:
[1236] Display Description
[1237] The server sends the generated explanation to the display of the terminal (smart glasses) and displays it to the user using a visual display means. The input is the explanation text, and the output is the information displayed in the user's field of view. In concrete terms, the smart glasses display displays "Scanning speed is the speed at which the scan head scans the surface. It is typically expressed in meters per second (m / s)."
[1238] By repeating this series of steps, workers can proceed with their work without confusion.
[1239] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1240] The present invention aims to further improve the efficiency of online meetings and the satisfaction of participants by incorporating an emotion engine that recognizes the emotions of users in addition to a system that quickly detects confusion among participants in online meetings and provides appropriate explanations. Specific embodiments of the present invention will be described below.
[1241] System initialization
[1242] The server loads models and libraries for face recognition, facial expression analysis, confusion detection, generation, display, and emotion recognition, performs necessary initialization, establishes a connection with the online conferencing tool, and prepares to capture each user's video feed in real time.
[1243] Capturing a video feed
[1244] The server captures each user's video feed, converts it into an analyzable format, and stores it. This video feed is the key data used for face detection and facial expression and emotion analysis.
[1245] Face detection and facial expression analysis
[1246] The server analyzes the captured video feed and detects each participant's face using a facial recognition means, and acquires facial expression data of the detected faces using an expression analysis means, including detailed facial expression information such as eyebrow movements, eye opening and closing, and mouth movements.
[1247] Detecting embarrassment and emotions
[1248] The server uses both facial expression analysis and emotion recognition to determine whether the participant is confused or not. In addition, it simultaneously analyzes the user's emotions (happiness, sadness, anger, etc.) and obtains detailed emotional data.
[1249] Identifying confusing words
[1250] The server analyzes the conversation based on the identified confusion state and its timestamp, and identifies words and concepts that are the subject of confusion. This analysis process uses speech recognition and natural language processing technologies.
[1251] Generate word descriptions
[1252] The server uses a generator to automatically generate explanations for the identified words and concepts, and the generated explanations are provided to the participants in a concise and easy-to-understand format.
[1253] Displaying explanations and emotional feedback
[1254] The server sends the generated explanation to the confused user's device via a display means, and may provide similar feedback to other participants in some cases. It may also provide feedback regarding the progress of the meeting based on the emotion recognition results.
[1255] Specific examples
[1256] For example, if User B looks confused when the term "ecosystem" is used during a meeting:
[1257] 1. The server detects User B's confusion state and timestamp.
[1258] 2. Using facial expression analysis and emotion recognition means, we recognize that User B is not only confused but also a little nervous.
[1259] 3. Analyze the conversation and identify that the word "ecosystem" is the source of the question.
[1260] 4. Use the generative tools to generate the statement, "An ecosystem is a collection of interconnected organisms and their surrounding environment."
[1261] 5. This description may pop up on User B's screen and be shared with other participants.
[1262] 6. Based on emotional feedback, advice and tips on how to proceed with the meeting are also provided.
[1263] This series of processes quickly resolves any questions or confusion that meeting participants may have, allowing everyone to participate comfortably in the meeting. Feedback based on emotional data can also be used to improve the atmosphere of the meeting.
[1264] The processing flow will be explained below.
[1265] Step 1:
[1266] The server loads models and libraries for face recognition, facial expression analysis, confusion state detection, generation, display, and emotion recognition, performs necessary initial settings, and prepares to establish a connection with the online conference tool.
[1267] Step 2:
[1268] The server captures each participant's video feed in real time through the online conferencing tool, converts this data into a parsable format, and prepares it for processing.
[1269] Step 3:
[1270] The server analyzes each frame of the captured video feed and uses facial recognition to detect participants' faces, which are then stored for use in the next step.
[1271] Step 4:
[1272] The server acquires the facial expression data detected using the facial expression analysis means, and analyzes detailed facial expression information such as eyebrow movements, eye opening and closing, and mouth movements.
[1273] Step 5:
[1274] The server uses emotion recognition means to analyze facial expression data and other biometric signals to identify the user's emotions (happiness, sadness, anger, confusion, etc.), thereby determining with high accuracy whether the user is confused.
[1275] Step 6:
[1276] If the server detects a confused state, it records the timestamp and the face ID of the user. Based on this information, the next step is to analyze the conversation.
[1277] Step 7:
[1278] The server analyzes the audio data of the meeting based on the recorded timestamps, identifies words or concepts that are confusing, converts the audio data into text using speech recognition technology, and extracts keywords using natural language processing technology.
[1279] Step 8:
[1280] The server uses a generator to automatically generate explanations for the identified words and concepts, which are then summarized in a concise and easy-to-understand format.
[1281] Step 9:
[1282] The server sends the generated explanation to the confused user's terminal and displays it in a pop-up format through a display means, and provides similar feedback to other participants as needed.
[1283] Step 10:
[1284] Users can review explanations displayed on their devices to gain a better understanding of words or concepts they were confused about, allowing them to keep up with the progress of the meeting.
[1285] Step 11:
[1286] The server uses the emotion data to provide feedback and advice on how the meeting should proceed. For example, if the user is nervous, it will encourage them to relax, and if they are tired, it will suggest taking a break.
[1287] Step 12:
[1288] The server runs this process continuously for each frame of the video feed, detecting and addressing confusion and emotions in real time, allowing for immediate response as new questions or emotional states arise.
[1289] Example 2
[1290] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1291] In online meetings, participants often become confused by technical terms and complex concepts they do not understand. This confusion can hinder the progress of the meeting and reduce participant satisfaction. It is also problematic when the meeting proceeds without noticing that some participants are confused. Furthermore, if participants' emotional states cannot be addressed, the meeting atmosphere becomes awkward and productivity decreases. An efficient solution to this problem is needed.
[1292] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1293] In this invention, the server includes face recognition means, facial expression analysis means, confusion detection means, generation means, display means, emotion recognition means, and conversation content analysis means, which enable the server to quickly detect participants' confusion, provide appropriate explanations, and analyze participants' emotional states in real time to provide feedback.
[1294] "Facial recognition means" is a technology for identifying the faces of each user participating in an online conference.
[1295] The "facial expression analysis means" is a technology for analyzing the facial expressions of each detected user and acquiring facial expression data.
[1296] The "confusion state detection means" is a technique for determining whether the user is confused or not based on the facial expression data obtained by the facial expression analysis means.
[1297] A "generator" is a technology that automatically generates explanations for identified words or concepts.
[1298] The "display means" is a technique for displaying the generated explanation on the user's terminal.
[1299] The "emotion recognition means" is a technology that analyzes the user's emotional state (joy, sadness, anger, etc.) based on the facial expression data obtained by the facial expression analysis means.
[1300] The "conversation content analysis means" is a technology that analyzes the conversation content before and after the time when a confused state is detected, and identifies the specific words or phrases that cause the confusion.
[1301] The present invention provides a system for quickly detecting participant confusion in an online conference and providing appropriate explanations. This system further incorporates an emotion engine that recognizes user emotions, thereby improving the efficiency of the conference and participant satisfaction. A specific embodiment of the system is described below.
[1302] The server first loads models and libraries for face recognition, facial expression analysis, confusion detection, generation, display, and emotion recognition. This process uses machine learning libraries such as TensorFlow and Keras. It also establishes a connection with online conferencing tools (e.g., Zoom, Microsoft Teams) and prepares to capture each user's video feed in real time. At this time, it obtains the necessary API keys and access tokens.
[1303] The server then captures each user's video feed and converts it into a format that can be analyzed using a conversion tool such as FFmpeg. This video data is crucial for facial recognition and facial expression analysis.
[1304] The server analyzes the captured video feed frame by frame and detects each participant's face using the OpenCV library. It then uses the dlib library to extract facial feature points and perform a detailed analysis of the movements of each part of the face. Based on this data, the facial expression analyzer performs a detailed analysis of the user's facial expressions.
[1305] Furthermore, the server uses the Emotion Recognition API to analyze the user's emotional state based on the acquired facial expression data, for example, determining whether the participant is confused or expressing emotions such as joy, sadness, or anger.
[1306] If a confusion state is detected, the server analyzes the conversation content based on the timestamp. The conversation content is converted into text using online conferencing tools or speech recognition technology (e.g., Google Cloud Speech-to-Text API), and analyzed with a natural language processing engine (e.g., the BERT model) to identify words and phrases that cause confusion.
[1307] The server then uses a generative method to automatically generate explanations for the identified words and phrases. Specifically, it inputs the prompt "What is an ecosystem?" into a generative AI model such as GPT-3 and uses the generated answer to create an explanation. For example, it generates an explanation such as "An ecosystem is a collection of interconnected organisms and their surrounding environment."
[1308] Finally, the server sends the generated explanation to the user's device and displays it as a pop-up on the display. In some cases, the same explanation can be shared with other participants. The server also provides feedback on the progress of the meeting based on the emotion recognition results. For example, it may give advice such as, "The participants seem nervous, so it would be more effective if you spoke a little more slowly."
[1309] This series of processes quickly resolves any questions or confusion that meeting participants may have, allowing everyone to participate comfortably in the meeting. Feedback based on emotional data can also be used to improve the atmosphere of the meeting.
[1310] Prompt Sentence Examples
[1311] "What is an ecosystem?"
[1312] "Please provide additional explanation for the confusion."
[1313] "Please analyze participants' facial expression data and perform emotion recognition."
[1314] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1315] Step 1:
[1316] The server loads models and libraries for face recognition, facial expression analysis, confusion detection, generation, display, and emotion recognition. Specifically, it loads the necessary models into memory using machine learning libraries such as TensorFlow and Keras. It also establishes a connection using the online conference tool's API and prepares to obtain each user's video feed. The input is a request to load the models and libraries, and the output is the results of those loads.
[1317] Step 2:
[1318] The server captures each user's video feed in real time. Specifically, it obtains the video stream through the API of a conferencing tool such as Zoom or Microsoft Teams, and converts it into an analyzable format (e.g., MP4) using FFmpeg. This video data is used for subsequent processing. The input is each user's video feed, and the output is the video data converted into an analyzable format.
[1319] Step 3:
[1320] The server analyzes the captured video feed frame by frame and detects each participant's face using the OpenCV library. It then uses the dlib library to extract facial feature points and perform detailed analysis of eyebrow, eye, and mouth movements. The input is each frame of video data, and the output is facial data with extracted feature points.
[1321] Step 4:
[1322] The server uses facial expression analysis means to obtain facial expression data based on the acquired facial data. Specifically, it uses the Emotion Recognition API to analyze the user's emotional state based on information such as eyebrow movements, eye opening and closing, and mouth movements. The input is facial data, and the output is detailed data on each facial expression and the results of that emotion analysis.
[1323] Step 5:
[1324] The server uses the confusion detection means to determine the user's confusion state. Based on the data obtained by the facial expression analysis means, it applies a specific algorithm (e.g., random forest or support vector machine) to determine whether the user is confused. The input is facial expression data and the emotion analysis result, and the output is the confusion state determination result.
[1325] Step 6:
[1326] The server analyzes the conversation content based on the timestamp of the user who was detected as confused. It acquires the audio data using the conferencing tool's API and converts the audio into text using the Google Cloud Speech-to-Text API. It then analyzes the text data using a natural language processing engine such as the BERT model to identify words and phrases that cause confusion. The input is the audio data, and the output is the text data and the identified words and phrases.
[1327] Step 7:
[1328] The server automatically generates explanations for the identified words and phrases using a generation method. Specifically, a prompt such as "What is an ecosystem?" is input to a generative AI model such as GPT-3, and an explanation is generated based on the response. The input is the prompt, and the output is the generated explanation.
[1329] Step 8:
[1330] The server sends the generated explanation to the confused user's device and displays it as a pop-up using a display device. It also shares the same explanation with other participants as needed. It also provides feedback on the progress of the meeting based on the emotion recognition results. The inputs are the generated explanation and the emotion analysis results, and the outputs are the display of the explanation and the provision of feedback.
[1331] (Application example 2)
[1332] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1333] Conventional online conference systems and factory work support systems have the problem of making it difficult to quickly detect situations and provide appropriate feedback when workers or meeting participants are confused or emotionally troubled. In particular, in factories, confusion can lead to reduced work efficiency and incorrect work procedures, which can also affect safety. To solve this problem, the present invention provides a system that detects a worker's emotions and confusion in real time and provides appropriate support.
[1334] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes a face recognition means, a facial expression analysis means, an emotion recognition means, a confusion state detection means, a generation means, a display means, a video feed acquisition means, a real-time processing means, a means for generating an explanation using a generative AI model, and a means for generating an explanation using a prompt sentence. This makes it possible to detect in real time that a worker is confused, and to instantly generate and display appropriate explanations regarding specific work procedures. This makes it possible to improve work efficiency and ensure safety.
[1335] "Facial recognition tools" are technologies and devices used to detect an individual's face from a video feed and identify that person.
[1336] "Facial expression analysis means" refers to technology and devices for analyzing subtle changes in facial expressions from a detected face and identifying the person's emotional state.
[1337] "Emotion recognition means" refers to technology and devices that use facial expression analysis data to determine a person's emotions (joy, sadness, anger, confusion, etc.).
[1338] The "confusion state detection means" refers to a technique and device for identifying whether a person is confused or not based on emotion recognition and facial expression analysis.
[1339] The "generation means" refers to technology and devices for automatically generating necessary explanations and information based on the identified confusion state and emotions.
[1340] "Display means" refers to the technology and devices used to visually present the generated explanations and information to the user.
[1341] "Video Feed Acquisition Measures" means the technology and equipment used to capture video footage in real time and convert and store that data in an analyzable format.
[1342] "Real-time processing means" refers to techniques and devices that rapidly process captured video feeds and analytical data to provide immediate feedback.
[1343] "Means for generating explanations using generative AI models" refers to technologies and devices that use artificial intelligence models to automatically generate explanations based on specific situations and contexts.
[1344] The "means for generating an explanation using a prompt sentence" refers to a technique and device that allows an artificial intelligence model to generate an appropriate explanation based on a specific prompt sentence given as input.
[1345] The present invention is a system for detecting the emotions and confusion of factory workers in real time and providing appropriate support. Specific embodiments of the system will be described below.
[1346] System Configuration
[1347] The server includes a facial recognition means, a facial expression analysis means, an emotion recognition means, a confusion state detection means, a generation means, a display means, a video feed acquisition means, a real-time processing means, a means for generating an explanation using a generative AI model, and a means for generating an explanation using a prompt sentence.
[1348] Hardware and software used
[1349] This system uses the following hardware and software:
[1350] Camera: Mounted on the robot, it captures a video feed of the worker in real time.
[1351] Server: Processes the video feed and performs facial recognition, facial expression analysis, and emotion recognition.
[1352] dlib: A library for face detection.
[1353] EmotionEngine: A proprietary module for facial expression analysis and emotion recognition.
[1354] TextGenerator: A module for generating explanations using generative AI models.
[1355] Data processing and calculation
[1356] The server processes and calculates the data as follows:
[1357] 1. Capturing and saving video feed:
[1358] The robot's camera captures a real-time video feed of the worker, converts it into an analyzable format, and stores it.
[1359] 2. Facial Recognition and Expression Analysis:
[1360] The dlib library is used to detect the worker's face from the video feed, and the EmotionEngine is used to obtain detailed facial expression data.
[1361] 3. Detecting embarrassment and emotions:
[1362] The EmotionEngine analyzes facial expression data to determine whether the worker is confused and obtains emotional data (e.g., tension, anxiety, joy, etc.).
[1363] 4. Identify the task steps that cause confusion:
[1364] Based on the identified confusion state and its timestamp, the work content is analyzed to identify the part that causes the confusion.
[1365] 5. Generating explanations using generative means:
[1366] Using TextGenerator, a prompt sentence is input into the generative AI model to automatically generate an explanation for the cause of confusion.
[1367] 6. Displaying explanatory and emotional feedback:
[1368] The generated explanation is displayed on the robot's display, providing visual assistance to the worker.
[1369] Specific examples
[1370] As an example, here is the process that would occur if a worker becomes confused while performing "Step 5" in a factory.
[1371] 1. The server captures the worker's video feed through the robot's camera.
[1372] 2. The server analyzes the video feed and detects the worker's face using the dlib library.
[1373] 3. Use EmotionEngine to analyze facial expression and emotion data and recognize when the worker is confused.
[1374] 4. Using the timestamps from the video feed of the task, identify the task step ("Step 5") that caused the confusion.
[1375] 5. Using TextGenerator, enter the "Explanation of points where mistakes are likely to occur in step 5" as a prompt sentence into the generative AI model.
[1376] 6. The generated instructions are displayed on the robot's display, allowing the worker to refer to them and continue the work.
[1377] Example prompt sentence:
[1378] "Explanation of points where mistakes are likely to occur in step 5"
[1379] Through the above series of processes, it is possible to detect situations in which a worker is confused in real time and provide appropriate support, thereby improving work efficiency and reducing worker stress.
[1380] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1381] Step 1:
[1382] The server captures the worker's video feed in real time using the robot's camera. The input is the video footage obtained from the camera, and the output is the captured video feed. The server converts it into an analyzable format and stores it.
[1383] Step 2:
[1384] The server uses the dlib library to detect the worker's face from the video feed. The input is the video feed saved in step 1, and the output is the location information (coordinates) of the detected face. The server performs face recognition processing based on this location information.
[1385] Step 3:
[1386] The server uses the EmotionEngine to analyze facial expressions. The input is the facial position information and video feed obtained in step 2, and the output is detailed facial expression data. The server analyzes eyebrow movements, eye opening and closing, mouth movements, etc.
[1387] Step 4:
[1388] The server further uses the EmotionEngine to perform emotion recognition. The input is the facial expression data obtained in step 3, and the output is the worker's emotional data (e.g., joy, sadness, anger, confusion, etc.). The server determines whether the worker is confused.
[1389] Step 5:
[1390] The server analyzes the work content based on the identified confusion state and its timestamp. The input is the timestamp of the confusion state and the timestamp of the video feed, and the output is the specific work step that caused the confusion. The server analyzes the conversation and work content to identify the cause of the confusion.
[1391] Step 6:
[1392] The server uses TextGenerator to input an explanation of the cause of confusion as a prompt sentence into the generative AI model, which then generates an appropriate explanation. The input is a prompt sentence related to the identified work step (e.g., "An explanation of the points in step 5 where mistakes are likely to occur"), and the output is the generated explanation sentence. The server inputs this prompt sentence into the generative AI model to obtain the necessary explanation.
[1393] Step 7:
[1394] The server displays the generated explanation on the robot's display. The input is the explanatory text obtained in step 6, and the output is visual support information displayed on the display. The server immediately provides appropriate explanations when a worker becomes confused, thereby improving work efficiency and ensuring safety.
[1395] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1396] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1397] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1398] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1399] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1400] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1401] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1402] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1403] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1404] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1405] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1406] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1407] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1408] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1409] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1410] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1411] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1412] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1413] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1414] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1415] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1416] The following is further disclosed regarding the above embodiment.
[1417] (Claim 1)
[1418] A facial recognition means;
[1419] Facial expression analysis means;
[1420] confusion state detection means;
[1421] generating means;
[1422] A display means;
[1423] A system including:
[1424] (Claim 2)
[1425] 10. The system of claim 1, further comprising a video feed acquisition means.
[1426] (Claim 3)
[1427] 10. The system of claim 1, further comprising real-time processing means.
[1428] "Example 1"
[1429] (Claim 1)
[1430] A facial recognition means;
[1431] Facial expression analysis means;
[1432] confusion state detection means;
[1433] generating means;
[1434] A display means;
[1435] a means for obtaining a video feed;
[1436] A voice data analysis means;
[1437] natural language processing means;
[1438] Pop-up display means;
[1439] A system including:
[1440] (Claim 2)
[1441] 10. The system of claim 1, further comprising real-time processing means.
[1442] (Claim 3)
[1443] 2. The system of claim 1, wherein the generating means uses a generative AI model to generate concise explanations.
[1444] "Application Example 1"
[1445] (Claim 1)
[1446] A facial recognition means;
[1447] Facial expression analysis means;
[1448] confusion state detection means;
[1449] generating means;
[1450] A display means;
[1451] a means for obtaining a real-time video feed;
[1452] A voice analysis means;
[1453] visual display means for displaying the generated description;
[1454] A system including:
[1455] (Claim 2)
[1456] 10. The system of claim 1, further comprising an audio data analysis means.
[1457] (Claim 3)
[1458] 10. The system of claim 1, further comprising a generative AI model that generates explanations for technical terms and operating procedures.
[1459] "Example 2: Combining Emotion Engines"
[1460] (Claim 1)
[1461] A facial recognition means;
[1462] Facial expression analysis means;
[1463] confusion state detection means;
[1464] generating means;
[1465] A display means;
[1466] An emotion recognition means;
[1467] A conversation content analysis means;
[1468] A system including:
[1469] (Claim 2)
[1470] 10. The system of claim 1, further comprising a video feed acquisition means.
[1471] (Claim 3)
[1472] 10. The system of claim 1, further comprising real-time processing means.
[1473] "Application example 2 when combining emotion engines"
[1474] (Claim 1)
[1475] A facial recognition means;
[1476] Facial expression analysis means;
[1477] An emotion recognition means;
[1478] confusion state detection means;
[1479] generating means;
[1480] A display means;
[1481] a means for obtaining a video feed;
[1482] real-time processing means;
[1483] a means for generating an explanation using a generative AI model;
[1484] a means for generating an explanation using a prompt sentence;
[1485] A system including:
[1486] (Claim 2)
[1487] 10. The system of claim 1, which can be used in a factory robot.
[1488] (Claim 3)
[1489] 10. The system of claim 1, further comprising means for detecting a state of confusion in a worker and generating an explanation for a particular work procedure. [Explanation of symbols]
[1490] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. A facial recognition means; Facial expression analysis means; confusion state detection means; generating means; A display means; A system including:
2. The system of claim 1 further comprising a video feed acquisition means.
3. 10. The system of claim 1 further comprising real-time processing means.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A