system

The system uses facial recognition and generation AI to detect and explain unfamiliar words in online meetings, ensuring participants understand without disrupting the flow.

JP2026072304APending Publication Date: 2026-05-01SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-18
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Participants in online meetings often encounter words they don't understand, which can delay the meeting progress.

Method used

A system using facial recognition AI to detect participants' expressions of confusion, identify unfamiliar words, and generate explanations using a generation AI, which are then displayed on the screen tailored to the participant's level.

Benefits of technology

Enables participants to instantly understand unfamiliar words during online meetings, maintaining psychological safety and ensuring smooth meeting progress without interruptions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026072304000001_ABST
    Figure 2026072304000001_ABST
Patent Text Reader

Abstract

The system according to this embodiment aims to enable participants to instantly understand the meaning of unfamiliar words during online meetings. [Solution] The system according to the embodiment comprises a detection unit, an identification unit, a generation unit, and a display unit. The detection unit detects the facial expressions of the participants. The identification unit identifies unfamiliar words based on the facial expressions detected by the detection unit. The generation unit generates explanations for the words identified by the identification unit. The display unit displays the explanations generated by the generation unit on the screen.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In the conventional technology, when participants encounter words they don't understand during an online meeting, there is a risk that the progress of the meeting will be delayed.

[0005] The system according to the embodiment aims to enable participants to immediately understand the meaning of words they don't understand during an online meeting.

Means for Solving the Problems

[0006] The system according to the embodiment comprises a detection unit, an identification unit, a generation unit, and a display unit. The detection unit detects the facial expressions of the participants. The identification unit identifies unfamiliar words based on the facial expressions detected by the detection unit. The generation unit generates explanations for the words identified by the identification unit. The display unit displays the explanations generated by the generation unit on the screen. [Effects of the Invention]

[0007] The system according to this embodiment can enable participants to instantly understand the meaning of unfamiliar words during an online meeting. [Brief explanation of the drawing]

[0008] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Modes for carrying out the invention]

[0009] Hereinafter, an example of an embodiment of the system relating to the technology of this disclosure will be described with reference to the attached drawings.

[0010] First, let's explain the terminology used in the following explanation.

[0011] In the following embodiments, the signed processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Furthermore, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose computing on Graphics Processing Units), APU (Accelerated Processing Unit), or TPU (Tensor Processing Unit).

[0012] In the following embodiments, signed RAM (Random Access Memory) is a memory that temporarily stores information and is used as work memory by the processor.

[0013] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.

[0014] In the following embodiments, the numbered communication I / F (Interface) is an interface including a communication processor, an antenna, and the like. The communication I / F controls communication between multiple computers. Examples of communication standards applicable to the communication I / F include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0015] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B". That is, "A and / or B" means that it may be only A, only B, or a combination of A and B. Also, in this specification, when expressing three or more matters connected by "and / or", the same concept as "A and / or B" is applied.

[0016] [First Embodiment] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0017] As shown in FIG. 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0018] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. Also, the database 24 and the communication I / F 26 are connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0019] The smart device 14 comprises a computer 36, a receiving device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The receiving device 38, output device 40, and camera 42 are also connected to the bus 52.

[0020] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, and accepts user input. The touch panel 38A accepts user input via touch by detecting contact with an object (e.g., a pen or finger). The microphone 38B accepts user input via voice by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 (see Figure 2) acquires the data indicating the user input.

[0021] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user by outputting the data in a form perceptible to the user (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0022] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0023] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0024] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0025] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.

[0026] In the smart device 14, specific processing is performed by the processor 46. The storage 50 stores a specific processing program 60. The specific processing program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the specific processing program 60 from the storage 50 and executes the read specific processing program 60 on the RAM 48. The specific processing is realized by the processor 46 operating as a control unit 46A according to the specific processing program 60 executed on the RAM 48. The smart device 14 also has a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.

[0027] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device (e.g., a generation server) may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device having the data generation model 58. The data processing device 12 may also be a server device or a terminal device owned by a user (e.g., a mobile phone, robot, home appliance, etc.). Next, an example of processing by the data processing system 10 according to the first embodiment will be described.

[0028] (Example of form 1) An online meeting support system according to an embodiment of the present invention is a system that uses facial recognition AI to detect when a participant feels "I don't understand the meaning of this word" during an online meeting, and a generation AI displays an explanation of that word on the screen. When a participant encounters a word they don't understand, the online meeting support system uses facial recognition AI to detect the participant's facial expression. The facial recognition AI reads the emotion of "not understanding" from the participant's facial expression. For example, expressions such as frowning or widening the eyes often indicate the emotion of "not understanding". Next, if the facial recognition AI detects the emotion of "not understanding", the generation AI provides an explanation of the word. The generation AI analyzes the context containing the unfamiliar word and understands the meaning of the word. For example, if the word "Eviction Threshold" is unfamiliar in the context of "If the soft limit of Eviction Threshold is met, a SIGTERM signal will be sent", the generation AI analyzes the meaning of "Eviction Threshold" and generates an explanation. The generated explanation is displayed on the screen according to the participant's level. For example, a simple explanation is given to a new graduate participant, and a detailed explanation is given to a participant with specialized knowledge. This allows information to be provided in a way that is easy for all participants to understand. This system ensures that even if unfamiliar words come up during an online meeting, participants can continue the meeting while maintaining psychological safety. There's no need to stop the meeting every time an unfamiliar word is mentioned, and it doesn't disrupt the flow of the meeting. Furthermore, there's no need to spend time after the meeting looking up unfamiliar words. For example, if a new graduate participant is having an online meeting with team members and the word "Eviction Threshold" comes up, the facial recognition AI detects the participant's expression, and the generative AI displays the meaning of "Eviction Threshold" on the screen. This allows the new graduate participant to understand the meaning of the word without disrupting the meeting. In this way, online meeting support systems are tools that support participants when unfamiliar words come up during online meetings, ensuring psychological safety and enabling smooth meeting progress.This allows online meeting support systems to facilitate smooth online meetings while ensuring the psychological safety of participants.

[0029] The online meeting support system according to the embodiment comprises a detection unit, an identification unit, a generation unit, and a display unit. The detection unit detects the facial expressions of participants. The facial expressions of participants include, for example, smiles, confusion, and surprise, but are not limited to such examples. The detection unit detects the facial expressions of participants using, for example, facial recognition technology. The detection unit can also analyze the facial expressions of participants in real time using a facial expression recognition algorithm. For example, the detection unit analyzes image data captured by a camera to detect the facial expressions of participants. The identification unit identifies unfamiliar words based on the facial expressions detected by the detection unit. The identification unit identifies unfamiliar words by, for example, analyzing patterns of facial expression changes. The identification unit can also identify unfamiliar words by performing contextual analysis using a language model. For example, the identification unit analyzes a combination of participant facial expression data and conversational context data to identify unfamiliar words. The generation unit generates explanations for the words identified by the identification unit. The generation unit generates explanations for unfamiliar words using, for example, a generation AI. The generation AI generates explanations for unfamiliar words using, for example, a text generation AI (e.g., LLM). Furthermore, the generation unit can use generation AI to analyze contexts containing unfamiliar words and understand their meanings. For example, the generation unit analyzes contexts containing unfamiliar words, understands their meanings, and generates appropriate explanations. The display unit displays the explanations generated by the generation unit on the screen. The display unit displays the generated explanations according to the participants' levels, for example. For example, the display unit provides explanations in simple terms for new graduates and detailed explanations for participants with specialized knowledge. The display unit can also display the generated explanations in real time. For example, the display unit displays the generated explanations as pop-ups on the screen. As a result, the online meeting support system according to this embodiment can smoothly conduct online meetings by detecting participants' facial expressions, identifying unfamiliar words, and generating and displaying their explanations.

[0030] The detection unit detects the facial expressions of participants. These expressions include, but are not limited to, smiles, confusion, and surprise. The detection unit uses, for example, facial recognition technology to detect participants' expressions. Specifically, it analyzes image data captured by cameras in real time and extracts facial feature points. This allows for high-precision detection of expressions such as smiles, confusion, and surprise. Furthermore, the detection unit can analyze participants' expressions in real time using facial recognition algorithms. For example, by utilizing a deep learning-based facial recognition model to capture subtle facial changes, it is possible to understand emotional states in more detail. This allows for accurate detection of the emotions participants are experiencing during a meeting and utilize this information for subsequent processing. The detection unit can further improve the accuracy of facial expression detection by acquiring image data from different angles using multiple cameras and constructing a three-dimensional facial model. Additionally, the detection unit dynamically adjusts analysis parameters according to lighting conditions and camera resolution, enabling optimal facial expression detection at all times. This allows the detection unit to detect the facial expressions of online meeting participants with high precision and in real time, providing the data necessary for subsequent processing.

[0031] The identification unit identifies unfamiliar words based on facial expressions detected by the detection unit. For example, the identification unit analyzes patterns of facial expression changes to identify unfamiliar words. Specifically, it detects when a participant shows expressions of confusion or surprise and analyzes the conversation content before and after to identify unfamiliar words. The identification unit can also use language models to perform contextual analysis and identify unfamiliar words. For example, it uses natural language processing techniques to analyze the context of the conversation and determine whether certain words or phrases are difficult for the participant to understand. The identification unit analyzes the participant's facial expression data and the conversation context data in combination to identify unfamiliar words. For example, if a confused expression is detected, it analyzes the content of the statement immediately preceding it to check whether it contains technical terms or difficult phrases. Furthermore, the identification unit can utilize past meeting data and participant profile information to learn which words and phrases certain participants tend to find confusing, enabling more accurate identification. As a result, the identification unit can quickly and accurately identify words and phrases that are difficult for participants to understand and provide the information necessary for the next processing.

[0032] The generation unit generates explanations for words identified by the identification unit. The generation unit generates explanations for unfamiliar words, for example, using a generation AI. Specifically, it uses a text generation AI (e.g., LLM) to generate explanations for unfamiliar words. The generation AI has learned from a large amount of text data and can generate appropriate explanations based on context. For example, the generation unit analyzes the context containing an unfamiliar word, understands its meaning, and generates an appropriate explanation. The generation AI can generate explanations that include not only the meaning of the word but also specific situations and examples in which the word is used. This allows participants to understand the meaning of the word more deeply. Furthermore, the generation unit can adjust the generated explanations to the level of the participants. For example, it can provide simple explanations for new graduates and detailed explanations for participants with specialized knowledge. The generation unit can also update the generated explanations in real time, providing explanations at the appropriate time according to the progress of the meeting. This allows the generation unit to quickly and appropriately generate explanations for words and phrases that are difficult for participants to understand, supporting the smooth progress of online meetings.

[0033] The display unit displays the explanations generated by the generation unit on the screen. For example, the display unit can display the generated explanations according to the participant's level. Specifically, it can explain things in simple terms to new graduates and provide detailed explanations to participants with specialized knowledge. The display unit can also display the generated explanations in real time. For example, it can display the generated explanations as pop-ups on the screen so that participants can check them immediately. Furthermore, the display unit can receive participant feedback and adjust the displayed content. For example, it can check whether participants understood the explanation and provide additional explanations as needed. In addition, the display unit supports multiple display formats and can be customized according to participants' preferences. For example, it can provide explanations using audio and video in addition to text. This allows the display unit to provide explanations in a format that is easy for participants to understand, supporting the smooth progress of online meetings. Furthermore, the display unit can save the generated explanations for later reference. This allows participants to review the explanations after the meeting and deepen their understanding. The display unit can quickly and appropriately provide explanations of unfamiliar words and phrases to online meeting participants, supporting the smooth progress of meetings.

[0034] The generation unit can generate explanations for unfamiliar words using a generation AI. For example, the generation unit can automatically generate explanations for unfamiliar words using a generation AI. For example, the generation unit can generate explanations for unfamiliar words using a text generation AI (e.g., LLM). The generation unit can also analyze the context containing an unfamiliar word and understand the meaning of that word using a generation AI. For example, the generation unit analyzes the context containing an unfamiliar word, understands the meaning of that word, and generates an appropriate explanation. In this way, explanations for unfamiliar words can be automatically generated by using a generation AI. Some or all of the above-described processes in the generation unit may be performed using a generation AI, for example, or without using a generation AI. For example, the generation unit can input contextual data containing an unfamiliar word into a generation AI and have the generation AI generate an explanation for the word.

[0035] The display unit can display the generated explanations according to the participant's level. For example, the display unit can explain the generated explanations in simple terms to new graduates and provide detailed explanations to participants with specialized knowledge. The display unit can customize the generated explanations according to the participant's level. The display unit can also display the generated explanations in real time. For example, the display unit can display the generated explanations as pop-ups on the screen. This makes it possible to provide easily understandable information by displaying explanations tailored to the participant's level. Some or all of the above processing in the display unit may be performed using AI, or not. For example, the display unit can input the generated explanation data into AI and have the AI ​​select a display method according to the participant's level.

[0036] The identification unit can identify unknown words based on detected facial expressions. For example, the identification unit can analyze patterns of facial expression changes to identify unknown words. The identification unit can also identify unknown words by performing contextual analysis using a language model, for example. For example, the identification unit can analyze a combination of participant facial expression data and conversational context data to identify unknown words. This improves the accuracy of identification by identifying unknown words based on detected facial expressions. Some or all of the above processing in the identification unit may be performed using AI, for example, or without AI. For example, the identification unit can input detected facial expression data into AI and have the AI ​​perform the identification of unknown words.

[0037] The generation unit can analyze a context containing an unknown word and understand the meaning of that word. For example, the generation unit can use a generation AI to analyze a context containing an unknown word and understand its meaning. The generation AI, for example, can use a text generation AI (e.g., LLM) to analyze the meaning of the unknown word. Furthermore, the generation unit can also use a generation AI to analyze a context containing an unknown word, understand its meaning, and generate an appropriate explanation. For example, the generation unit can analyze a context containing an unknown word, understand its meaning, and generate an appropriate explanation. This allows for a more accurate understanding of the word's meaning and the generation of an appropriate explanation by analyzing the context. Some or all of the above-described processes in the generation unit may be performed using a generation AI, or they may be performed without a generation AI. For example, the generation unit can input contextual data containing an unknown word into a generation AI and have the generation AI perform the task of understanding the word's meaning.

[0038] The detection unit can analyze the participant's past facial expression data and select the optimal detection algorithm. For example, the detection unit can cluster the past facial expression data and apply the optimal detection algorithm to each cluster. For example, the detection unit can select the optimal detection algorithm based on the participant's past facial expression data indicating "I don't know." The detection unit can also extract specific facial expression patterns from the participant's past facial expression data and adjust the detection algorithm based on them. In this way, the optimal detection algorithm can be selected by analyzing past facial expression data. Some or all of the above processing in the detection unit may be performed using AI, for example, or without AI. For example, the detection unit can input past facial expression data into AI and have the AI ​​perform the selection of the optimal detection algorithm.

[0039] The detection unit can perform filtering based on the participant's current environment and situation when detecting facial expressions. For example, if a participant is in a dark environment, the detection unit adjusts the accuracy of facial expression detection by considering the lighting conditions. For example, if a participant is in a noisy environment, the detection unit improves the accuracy of facial expression detection by filtering background noise. Furthermore, if a participant is moving, the detection unit can maintain the accuracy of facial expression detection by correcting motion blur. In this way, the accuracy of facial expression detection is improved by performing filtering based on the environment and situation. Some or all of the above processing in the detection unit may be performed using AI, for example, or without AI. For example, the detection unit can input current environmental data into the AI ​​and have the AI ​​perform the filtering adjustments.

[0040] The detection unit can prioritize the detection of highly relevant facial expressions by considering the participant's geographical location information when detecting facial expressions. For example, if the participant is in a different cultural area, the detection unit will prioritize the detection of facial expressions specific to that culture. For example, if the participant is in a specific region, the detection unit will prioritize the detection of facial expressions appropriate to the climate and environment of that region. Furthermore, if the participant is traveling, the detection unit can also prioritize the detection of facial expressions based on the culture and customs of the travel destination. In this way, by considering geographical location information, highly relevant facial expressions can be prioritized. Some or all of the above processing in the detection unit may be performed using AI, for example, or without AI. For example, the detection unit can input geographical location information into AI and have the AI ​​perform the detection of highly relevant facial expressions.

[0041] The detection unit can analyze the participant's social media activity when detecting facial expressions and detect relevant expressions. For example, if a participant has posted "I don't know" on social media, the detection unit will prioritize detecting that facial expression. For example, if a participant is discussing a specific topic on social media, the detection unit will detect facial expressions related to that topic. The detection unit can also detect facial expressions related to specific emotions if a participant is expressing those emotions on social media. In this way, relevant facial expressions can be detected by analyzing social media activity. Some or all of the above processing in the detection unit may be performed using AI, for example, or without AI. For example, the detection unit can input social media activity data into AI and have the AI ​​perform the detection of relevant facial expressions.

[0042] The identification unit can optimize its algorithm for identifying unknown words by referring to past meeting data at the time of identification. For example, the identification unit can cluster past meeting data and apply the optimal identification algorithm to each cluster. For example, the identification unit can identify frequently occurring unknown words from past meeting data and optimize their algorithms. The identification unit can also analyze past meeting data and optimize its algorithm for identifying unknown words related to specific topics. In this way, the identification algorithm can be optimized by referring to past meeting data. Some or all of the above processing in the identification unit may be performed using AI, for example, or without AI. For example, the identification unit can input past meeting data into AI and have the AI ​​perform the optimization of the identification algorithm.

[0043] The identification unit can apply different identification methods depending on the context and topic of the meeting at the time of identification. For example, in a technical meeting, the identification unit can apply an identification method specialized in technical terminology. In a business meeting, for example, the identification unit can apply an identification method specialized in business terminology. Furthermore, in an educational meeting, the identification unit can also apply an identification method specialized in educational terminology. This improves the accuracy of identification by applying an identification method according to the context and topic of the meeting. Some or all of the above processing in the identification unit may be performed using AI, for example, or not using AI. For example, the identification unit can input meeting context data into AI and have the AI ​​perform the application of the identification method.

[0044] The identification unit can adjust specific priorities based on the progress of the meeting at the time of identification. For example, in the early stages of the meeting, the identification unit may prioritize identifying basic words. In the middle of the meeting, for example, the identification unit may prioritize identifying words related to important topics. Furthermore, in the final stages of the meeting, the identification unit may also prioritize identifying general words. This allows for the identification of words at the appropriate time by adjusting specific priorities based on the progress of the meeting. Some or all of the above processing in the identification unit may be performed using AI, for example, or not using AI. For example, the identification unit can input meeting progress data into the AI ​​and have the AI ​​perform the adjustment of specific priorities.

[0045] The identification unit can improve the accuracy of identification by considering the expertise level of the meeting participants during the identification process. For example, the identification unit can perform detailed identification for participants with high expertise, simple identification for participants with low expertise, and moderate identification for participants with moderate expertise. This improves the accuracy of identification by considering the expertise level of the participants. Some or all of the above processing in the identification unit may be performed using AI, for example, or without AI. For example, the identification unit can input participant expertise level data into AI and have the AI ​​perform the improvement of identification accuracy.

[0046] The generation unit can adjust the level of detail in the explanations based on the importance of the words during generation. For example, the generation unit can generate detailed explanations for important words, and concise explanations for less important words. It can also generate explanations with a moderate level of detail for words of moderate importance. By adjusting the level of detail in the explanations based on the importance of the words, appropriate explanations can be generated. Some or all of the above processing in the generation unit may be performed using a generation AI, for example, or without a generation AI. For example, the generation unit can input word importance data into a generation AI and have the generation AI perform the adjustment of the level of detail in the explanations.

[0047] The generation unit can apply different generation algorithms depending on the category of the word during generation. For example, the generation unit can apply an algorithm that generates technical explanations to technical terms. For example, the generation unit can apply an algorithm that generates business explanations to business terms. Furthermore, the generation unit can also apply an algorithm that generates educational explanations to educational terms. By applying the appropriate generation algorithm according to the category of the word, the accuracy of the explanation is improved. Some or all of the above processing in the generation unit may be performed using a generation AI, for example, or without a generation AI. For example, the generation unit can input word category data into a generation AI and have the generation AI perform the application of the generation algorithm.

[0048] The generation unit can determine the priority of explanations based on word frequency during generation. For example, the generation unit can prioritize generating explanations for frequently used words. For example, it can postpone generating explanations for less frequently used words. The generation unit can also moderately prioritize generating explanations for words with moderate frequency. In this way, by determining the priority of explanations based on word frequency, explanations for important words can be generated preferentially. Some or all of the above processing in the generation unit may be performed using, for example, a generation AI, or without a generation AI. For example, the generation unit can input word frequency data into a generation AI and have the generation AI perform the determination of explanation priority.

[0049] The generation unit can adjust the order of explanations based on the relevance of words during generation. For example, the generation unit can prioritize generating explanations for highly relevant words. For example, it can postpone generating explanations for less relevant words. It can also moderately prioritize generating explanations for words with a moderate level of relevance. In this way, by adjusting the order of explanations based on the relevance of words, explanations can be provided in an appropriate order. Some or all of the above processing in the generation unit may be performed using, for example, a generation AI, or without a generation AI. For example, the generation unit can input word relevance data into a generation AI and have the generation AI perform the adjustment of the explanation order.

[0050] The display unit can select the optimal display method by referring to the participant's past operation history when displaying information. For example, the display unit may prioritize display methods that the participant has preferred to use in the past. For example, the display unit may select the most efficient display method from the participant's past operation history. The display unit can also exclude display methods that the participant has avoided in the past and provide the optimal display method. In this way, the optimal display method can be selected by referring to past operation history. Some or all of the above processing in the display unit may be performed using AI, for example, or without AI. For example, the display unit may input past operation history data into AI and have the AI ​​perform the selection of the optimal display method.

[0051] The display unit can select the optimal display method when displaying information, taking into account the participant's device information. For example, if a participant is using a smartphone, the display unit provides a display method that matches the screen size. For example, if a participant is using a tablet, the display unit provides a display method optimized for a larger screen. Furthermore, if a participant is using a desktop computer, the display unit can also provide a display method that includes detailed information. In this way, the optimal display method can be provided by taking device information into consideration. Some or all of the above processing in the display unit may be performed using AI, for example, or without AI. For example, the display unit can input device information into AI and have the AI ​​select the optimal display method.

[0052] The display unit can adjust the timing of its display based on the progress of the meeting. For example, in the early stages of the meeting, the display unit prioritizes displaying basic information. In the middle of the meeting, for example, the display unit prioritizes displaying information related to important topics. Furthermore, in the final stages of the meeting, the display unit can also prioritize displaying summary information. By adjusting the timing of the display based on the progress of the meeting, information can be displayed at the appropriate time. Some or all of the above processing in the display unit may be performed using AI, for example, or without AI. For example, the display unit can input meeting progress data into the AI ​​and have the AI ​​perform the adjustment of the display timing.

[0053] The display unit can customize the displayed content to take into account the participants' level of expertise. For example, the display unit can display detailed information to participants with high expertise, simplified information to participants with low expertise, and information of an appropriate level of detail to participants with moderate expertise. This allows for the display of appropriate content by considering the participants' level of expertise. Some or all of the above processing in the display unit may be performed using AI, for example, or without AI. For example, the display unit can input participant expertise level data into AI and have the AI ​​perform the customization of the displayed content.

[0054] The system according to the embodiment is not limited to the example described above, and various modifications are possible, for example, as follows.

[0055] Online meeting support systems can also include a data analysis department. This department analyzes data collected during meetings to evaluate meeting performance and participant contributions. For example, it can analyze the number of contributions and speaking time to visualize each participant's contribution. The data analysis department can also analyze the meeting's progress and discussion content to evaluate its effectiveness. For instance, it can assess the depth and breadth of discussions to quantitatively evaluate the meeting's quality. Furthermore, by comparing current meeting data with past meeting data, the data analysis department can identify areas for improvement and make suggestions for future meetings. This allows for improved meeting performance and a better evaluation of participant contributions.

[0056] Online meeting support systems can also include a scheduling unit. This unit automatically adjusts participants' schedules and suggests the optimal meeting time. For example, it analyzes all participants' calendars to identify time slots when everyone is available. The scheduling unit can also adjust scheduling priorities based on the importance and urgency of the meeting. For instance, it prioritizes urgent meetings and postpones less important ones. Furthermore, it can analyze participants' past schedule data to suggest the optimal meeting time. This reduces the effort required for scheduling and enables more efficient meetings.

[0057] Online meeting support systems can also include a document sharing function. This function automatically collects and shares documents used during the meeting with participants. For example, it can collect the meeting agenda and presentation materials in advance and distribute them to participants. The document sharing function can also share newly added documents in real time during the meeting. For example, if a participant uploads new documents during the meeting, the document sharing function will distribute those documents to everyone. Furthermore, the document sharing function can archive past meeting materials so that participants can refer to them later. This ensures efficient document sharing and smoother meeting progress.

[0058] Online meeting support systems can also include a voting function. This function provides a voting mechanism to gather participants' opinions during a meeting. For example, it can be used to conduct a vote to gather everyone's opinion on important decisions. The voting function can also compile voting results in real time and display them during the meeting. For instance, it can display results immediately after voting ends to support rapid decision-making. Furthermore, the voting function can analyze past voting data to understand trends in participants' opinions. This allows for efficient collection of participant feedback and supports decision-making in meetings.

[0059] Online meeting support systems can also include a breakout room feature. This feature allows participants to be divided into smaller groups for discussion during a meeting. For example, in a large meeting, participants can be divided into multiple breakout rooms to delve deeper into a specific topic. The breakout room feature can also monitor the progress of each room and provide support as needed. For instance, if a discussion in a room stalls, a facilitator can intervene to stimulate the discussion. Furthermore, the breakout room feature can record the discussion content from each room and share it with the general meeting afterward. This allows for more efficient discussions and maximizes the meeting's effectiveness.

[0060] The following briefly describes the processing flow for example form 1.

[0061] Step 1: The detection unit detects the participant's facial expressions. These expressions may include, but are not limited to, smiles, confusion, or surprise. The detection unit uses facial recognition technology and facial expression recognition algorithms to analyze image data captured by the camera and detect the participant's facial expressions in real time. Step 2: The identification unit identifies unfamiliar words based on the facial expressions detected by the detection unit. The identification unit analyzes the patterns of facial expression changes to identify unfamiliar words. It also performs contextual analysis using a language model, combining the participant's facial expression data with the conversation context data to identify unfamiliar words. Step 3: The generation unit generates explanations for the words identified by the identification unit. The generation unit uses a generation AI (e.g., a text generation AI) to generate explanations for unfamiliar words. The generation unit also analyzes the context containing the unfamiliar words, understands their meaning, and generates appropriate explanations. Step 4: The display unit displays the explanation generated by the generation unit on the screen. The display unit displays the generated explanation according to the participant's level. For example, it explains in simple terms to new graduates and provides a detailed explanation to participants with specialized knowledge. It also displays the generated explanation as a pop-up on the screen in real time.

[0062] (Example of form 2) An online meeting support system according to an embodiment of the present invention is a system that uses facial recognition AI to detect when a participant feels "I don't understand the meaning of this word" during an online meeting, and a generation AI displays an explanation of that word on the screen. When a participant encounters a word they don't understand, the online meeting support system uses facial recognition AI to detect the participant's facial expression. The facial recognition AI reads the emotion of "not understanding" from the participant's facial expression. For example, expressions such as frowning or widening the eyes often indicate the emotion of "not understanding". Next, if the facial recognition AI detects the emotion of "not understanding", the generation AI provides an explanation of the word. The generation AI analyzes the context containing the unfamiliar word and understands the meaning of the word. For example, if the word "Eviction Threshold" is unfamiliar in the context of "If the soft limit of Eviction Threshold is met, a SIGTERM signal will be sent", the generation AI analyzes the meaning of "Eviction Threshold" and generates an explanation. The generated explanation is displayed on the screen according to the participant's level. For example, a simple explanation is given to a new graduate participant, and a detailed explanation is given to a participant with specialized knowledge. This allows information to be provided in a way that is easy for all participants to understand. This system ensures that even if unfamiliar words come up during an online meeting, participants can continue the meeting while maintaining psychological safety. There's no need to stop the meeting every time an unfamiliar word is mentioned, and it doesn't disrupt the flow of the meeting. Furthermore, there's no need to spend time after the meeting looking up unfamiliar words. For example, if a new graduate participant is having an online meeting with team members and the word "Eviction Threshold" comes up, the facial recognition AI detects the participant's expression, and the generative AI displays the meaning of "Eviction Threshold" on the screen. This allows the new graduate participant to understand the meaning of the word without disrupting the meeting. In this way, online meeting support systems are tools that support participants when unfamiliar words come up during online meetings, ensuring psychological safety and enabling smooth meeting progress.This allows online meeting support systems to facilitate smooth online meetings while ensuring the psychological safety of participants.

[0063] The online meeting support system according to the embodiment comprises a detection unit, an identification unit, a generation unit, and a display unit. The detection unit detects the facial expressions of participants. The facial expressions of participants include, for example, smiles, confusion, and surprise, but are not limited to such examples. The detection unit detects the facial expressions of participants using, for example, facial recognition technology. The detection unit can also analyze the facial expressions of participants in real time using a facial expression recognition algorithm. For example, the detection unit analyzes image data captured by a camera to detect the facial expressions of participants. The identification unit identifies unfamiliar words based on the facial expressions detected by the detection unit. The identification unit identifies unfamiliar words by, for example, analyzing patterns of facial expression changes. The identification unit can also identify unfamiliar words by performing contextual analysis using a language model. For example, the identification unit analyzes a combination of participant facial expression data and conversational context data to identify unfamiliar words. The generation unit generates explanations for the words identified by the identification unit. The generation unit generates explanations for unfamiliar words using, for example, a generation AI. The generation AI generates explanations for unfamiliar words using, for example, a text generation AI (e.g., LLM). Furthermore, the generation unit can use generation AI to analyze contexts containing unfamiliar words and understand their meanings. For example, the generation unit analyzes contexts containing unfamiliar words, understands their meanings, and generates appropriate explanations. The display unit displays the explanations generated by the generation unit on the screen. The display unit displays the generated explanations according to the participants' levels, for example. For example, the display unit provides explanations in simple terms for new graduates and detailed explanations for participants with specialized knowledge. The display unit can also display the generated explanations in real time. For example, the display unit displays the generated explanations as pop-ups on the screen. As a result, the online meeting support system according to this embodiment can smoothly conduct online meetings by detecting participants' facial expressions, identifying unfamiliar words, and generating and displaying their explanations.

[0064] The detection unit detects the facial expressions of participants. These expressions include, but are not limited to, smiles, confusion, and surprise. The detection unit uses, for example, facial recognition technology to detect participants' expressions. Specifically, it analyzes image data captured by cameras in real time and extracts facial feature points. This allows for high-precision detection of expressions such as smiles, confusion, and surprise. Furthermore, the detection unit can analyze participants' expressions in real time using facial recognition algorithms. For example, by utilizing a deep learning-based facial recognition model to capture subtle facial changes, it is possible to understand emotional states in more detail. This allows for accurate detection of the emotions participants are experiencing during a meeting and utilize this information for subsequent processing. The detection unit can further improve the accuracy of facial expression detection by acquiring image data from different angles using multiple cameras and constructing a three-dimensional facial model. Additionally, the detection unit dynamically adjusts analysis parameters according to lighting conditions and camera resolution, enabling optimal facial expression detection at all times. This allows the detection unit to detect the facial expressions of online meeting participants with high precision and in real time, providing the data necessary for subsequent processing.

[0065] The identification unit identifies unfamiliar words based on facial expressions detected by the detection unit. For example, the identification unit analyzes patterns of facial expression changes to identify unfamiliar words. Specifically, it detects when a participant shows expressions of confusion or surprise and analyzes the conversation content before and after to identify unfamiliar words. The identification unit can also use language models to perform contextual analysis and identify unfamiliar words. For example, it uses natural language processing techniques to analyze the context of the conversation and determine whether certain words or phrases are difficult for the participant to understand. The identification unit analyzes the participant's facial expression data and the conversation context data in combination to identify unfamiliar words. For example, if a confused expression is detected, it analyzes the content of the statement immediately preceding it to check whether it contains technical terms or difficult phrases. Furthermore, the identification unit can utilize past meeting data and participant profile information to learn which words and phrases certain participants tend to find confusing, enabling more accurate identification. As a result, the identification unit can quickly and accurately identify words and phrases that are difficult for participants to understand and provide the information necessary for the next processing.

[0066] The generation unit generates explanations for words identified by the identification unit. The generation unit generates explanations for unfamiliar words, for example, using a generation AI. Specifically, it uses a text generation AI (e.g., LLM) to generate explanations for unfamiliar words. The generation AI has learned from a large amount of text data and can generate appropriate explanations based on context. For example, the generation unit analyzes the context containing an unfamiliar word, understands its meaning, and generates an appropriate explanation. The generation AI can generate explanations that include not only the meaning of the word but also specific situations and examples in which the word is used. This allows participants to understand the meaning of the word more deeply. Furthermore, the generation unit can adjust the generated explanations to the level of the participants. For example, it can provide simple explanations for new graduates and detailed explanations for participants with specialized knowledge. The generation unit can also update the generated explanations in real time, providing explanations at the appropriate time according to the progress of the meeting. This allows the generation unit to quickly and appropriately generate explanations for words and phrases that are difficult for participants to understand, supporting the smooth progress of online meetings.

[0067] The display unit displays the explanations generated by the generation unit on the screen. For example, the display unit can display the generated explanations according to the participant's level. Specifically, it can explain things in simple terms to new graduates and provide detailed explanations to participants with specialized knowledge. The display unit can also display the generated explanations in real time. For example, it can display the generated explanations as pop-ups on the screen so that participants can check them immediately. Furthermore, the display unit can receive participant feedback and adjust the displayed content. For example, it can check whether participants understood the explanation and provide additional explanations as needed. In addition, the display unit supports multiple display formats and can be customized according to participants' preferences. For example, it can provide explanations using audio and video in addition to text. This allows the display unit to provide explanations in a format that is easy for participants to understand, supporting the smooth progress of online meetings. Furthermore, the display unit can save the generated explanations for later reference. This allows participants to review the explanations after the meeting and deepen their understanding. The display unit can quickly and appropriately provide explanations of unfamiliar words and phrases to online meeting participants, supporting the smooth progress of meetings.

[0068] The generation unit can generate explanations for unfamiliar words using a generation AI. For example, the generation unit can automatically generate explanations for unfamiliar words using a generation AI. For example, the generation unit can generate explanations for unfamiliar words using a text generation AI (e.g., LLM). The generation unit can also analyze the context containing an unfamiliar word and understand the meaning of that word using a generation AI. For example, the generation unit analyzes the context containing an unfamiliar word, understands the meaning of that word, and generates an appropriate explanation. In this way, explanations for unfamiliar words can be automatically generated by using a generation AI. Some or all of the above-described processes in the generation unit may be performed using a generation AI, for example, or without using a generation AI. For example, the generation unit can input contextual data containing an unfamiliar word into a generation AI and have the generation AI generate an explanation for the word.

[0069] The display unit can display the generated explanations according to the participant's level. For example, the display unit can explain the generated explanations in simple terms to new graduates and provide detailed explanations to participants with specialized knowledge. The display unit can customize the generated explanations according to the participant's level. The display unit can also display the generated explanations in real time. For example, the display unit can display the generated explanations as pop-ups on the screen. This makes it possible to provide easily understandable information by displaying explanations tailored to the participant's level. Some or all of the above processing in the display unit may be performed using AI, or not. For example, the display unit can input the generated explanation data into AI and have the AI ​​select a display method according to the participant's level.

[0070] The detection unit can detect the emotion of "I don't know" from the participant's facial expression. The detection unit analyzes the participant's facial expression using, for example, a facial expression recognition algorithm and detects the emotion of "I don't know." For example, the detection unit can detect expressions such as frowning or widening the eyes and identify the emotion of "I don't know." The detection unit can also detect the emotion of "I don't know" using an emotion estimation function with an emotion engine or generative AI. For example, the detection unit can input image data captured by a camera into an emotion engine and have the emotion engine perform emotion estimation. This improves the accuracy of identifying unfamiliar words by detecting the emotion of "I don't know" from the participant's facial expression. Some or all of the above processing in the detection unit may be performed using, for example, AI, or not using AI. For example, the detection unit can input the participant's facial expression data into a generative AI and have the generative AI perform the detection of the emotion of "I don't know."

[0071] The identification unit can identify unknown words based on detected facial expressions. For example, the identification unit can analyze patterns of facial expression changes to identify unknown words. The identification unit can also identify unknown words by performing contextual analysis using a language model, for example. For example, the identification unit can analyze a combination of participant facial expression data and conversational context data to identify unknown words. This improves the accuracy of identification by identifying unknown words based on detected facial expressions. Some or all of the above processing in the identification unit may be performed using AI, for example, or without AI. For example, the identification unit can input detected facial expression data into AI and have the AI ​​perform the identification of unknown words.

[0072] The generation unit can analyze a context containing an unknown word and understand the meaning of that word. For example, the generation unit can use a generation AI to analyze a context containing an unknown word and understand its meaning. The generation AI, for example, can use a text generation AI (e.g., LLM) to analyze the meaning of the unknown word. Furthermore, the generation unit can also use a generation AI to analyze a context containing an unknown word, understand its meaning, and generate an appropriate explanation. For example, the generation unit can analyze a context containing an unknown word, understand its meaning, and generate an appropriate explanation. This allows for a more accurate understanding of the word's meaning and the generation of an appropriate explanation by analyzing the context. Some or all of the above-described processes in the generation unit may be performed using a generation AI, or they may be performed without a generation AI. For example, the generation unit can input contextual data containing an unknown word into a generation AI and have the generation AI perform the task of understanding the word's meaning.

[0073] The detection unit can estimate the participant's emotions and adjust the accuracy of facial expression detection based on the estimated emotions. The detection unit estimates the participant's emotions using an emotion estimation function, for example, using an emotion engine or a generative AI. For example, the detection unit can input image data captured by a camera into an emotion engine and have the emotion engine perform emotion estimation. For example, if the participant is tense, the detection unit can use a high-precision algorithm to detect subtle changes in facial expression. If the participant is relaxed, the detection unit can also use a low-precision algorithm to detect broader changes in facial expression. Furthermore, if the participant is tired, the detection unit can use a medium-precision algorithm to detect changes in facial expression. This improves detection accuracy by adjusting the accuracy of facial expression detection based on the participant's emotions. Some or all of the above processing in the detection unit may be performed using AI, for example, or without AI. For example, the detection unit can input the participant's emotion data into a generative AI and have the generative AI perform the adjustment of facial expression detection accuracy.

[0074] The detection unit can analyze the participant's past facial expression data and select the optimal detection algorithm. For example, the detection unit can cluster the past facial expression data and apply the optimal detection algorithm to each cluster. For example, the detection unit can select the optimal detection algorithm based on the participant's past facial expression data indicating "I don't know." The detection unit can also extract specific facial expression patterns from the participant's past facial expression data and adjust the detection algorithm based on them. In this way, the optimal detection algorithm can be selected by analyzing past facial expression data. Some or all of the above processing in the detection unit may be performed using AI, for example, or without AI. For example, the detection unit can input past facial expression data into AI and have the AI ​​perform the selection of the optimal detection algorithm.

[0075] The detection unit can perform filtering based on the participant's current environment and situation when detecting facial expressions. For example, if a participant is in a dark environment, the detection unit adjusts the accuracy of facial expression detection by considering the lighting conditions. For example, if a participant is in a noisy environment, the detection unit improves the accuracy of facial expression detection by filtering background noise. Furthermore, if a participant is moving, the detection unit can maintain the accuracy of facial expression detection by correcting motion blur. In this way, the accuracy of facial expression detection is improved by performing filtering based on the environment and situation. Some or all of the above processing in the detection unit may be performed using AI, for example, or without AI. For example, the detection unit can input current environmental data into the AI ​​and have the AI ​​perform the filtering adjustments.

[0076] The detection unit can estimate the participant's emotions and determine the priority of facial expressions to detect based on the estimated emotions. The detection unit estimates the participant's emotions using an emotion estimation function, for example, using an emotion engine or a generative AI. For example, the detection unit can input image data captured by a camera into an emotion engine and have the emotion engine perform emotion estimation. For example, if the participant is nervous, the detection unit will prioritize detecting facial expressions indicating nervousness. The detection unit can also prioritize detecting facial expressions indicating relaxation if the participant is relaxed. Furthermore, if the participant is confused, the detection unit can also prioritize detecting facial expressions indicating confusion. In this way, important facial expressions can be detected preferentially by determining the priority of facial expressions based on the participant's emotions. Some or all of the above processing in the detection unit may be performed using AI, for example, or without using AI. For example, the detection unit can input the participant's emotion data into a generative AI and have the generative AI perform the determination of the priority of facial expressions.

[0077] The detection unit can prioritize the detection of highly relevant facial expressions by considering the participant's geographical location information when detecting facial expressions. For example, if the participant is in a different cultural area, the detection unit will prioritize the detection of facial expressions specific to that culture. For example, if the participant is in a specific region, the detection unit will prioritize the detection of facial expressions appropriate to the climate and environment of that region. Furthermore, if the participant is traveling, the detection unit can also prioritize the detection of facial expressions based on the culture and customs of the travel destination. In this way, by considering geographical location information, highly relevant facial expressions can be prioritized. Some or all of the above processing in the detection unit may be performed using AI, for example, or without AI. For example, the detection unit can input geographical location information into AI and have the AI ​​perform the detection of highly relevant facial expressions.

[0078] The detection unit can analyze the participant's social media activity when detecting facial expressions and detect relevant expressions. For example, if a participant has posted "I don't know" on social media, the detection unit will prioritize detecting that facial expression. For example, if a participant is discussing a specific topic on social media, the detection unit will detect facial expressions related to that topic. The detection unit can also detect facial expressions related to specific emotions if a participant is expressing those emotions on social media. In this way, relevant facial expressions can be detected by analyzing social media activity. Some or all of the above processing in the detection unit may be performed using AI, for example, or without AI. For example, the detection unit can input social media activity data into AI and have the AI ​​perform the detection of relevant facial expressions.

[0079] The identification unit can estimate the participant's emotions and adjust the accuracy of identifying unknown words based on the estimated emotions. The identification unit estimates the participant's emotions using an emotion estimation function, for example, using an emotion engine or a generative AI. For example, the identification unit can input image data captured by a camera into the emotion engine and have the emotion engine perform emotion estimation. For example, if the participant is nervous, the identification unit performs a detailed analysis to improve the accuracy of identifying unknown words. Also, if the participant is relaxed, the identification unit can set the accuracy of identifying unknown words to a lower level and perform a simpler analysis. Furthermore, if the participant is confused, the identification unit can set the accuracy of identifying unknown words to a moderate level and perform a moderate analysis. In this way, the accuracy of identifying unknown words is improved by adjusting the identification accuracy based on the participant's emotions. Some or all of the above processing in the identification unit may be performed using AI, for example, or without using AI. For example, the identification unit can input the participant's emotion data into a generative AI and have the generative AI perform the adjustment of the identification accuracy.

[0080] The identification unit can optimize its algorithm for identifying unknown words by referring to past meeting data at the time of identification. For example, the identification unit can cluster past meeting data and apply the optimal identification algorithm to each cluster. For example, the identification unit can identify frequently occurring unknown words from past meeting data and optimize their algorithms. The identification unit can also analyze past meeting data and optimize its algorithm for identifying unknown words related to specific topics. In this way, the identification algorithm can be optimized by referring to past meeting data. Some or all of the above processing in the identification unit may be performed using AI, for example, or without AI. For example, the identification unit can input past meeting data into AI and have the AI ​​perform the optimization of the identification algorithm.

[0081] The identification unit can apply different identification methods depending on the context and topic of the meeting at the time of identification. For example, in a technical meeting, the identification unit can apply an identification method specialized in technical terminology. In a business meeting, for example, the identification unit can apply an identification method specialized in business terminology. Furthermore, in an educational meeting, the identification unit can also apply an identification method specialized in educational terminology. This improves the accuracy of identification by applying an identification method according to the context and topic of the meeting. Some or all of the above processing in the identification unit may be performed using AI, for example, or not using AI. For example, the identification unit can input meeting context data into AI and have the AI ​​perform the application of the identification method.

[0082] The identification unit can estimate the participant's emotions and determine the priority of words to identify based on the estimated emotions. The identification unit estimates the participant's emotions using an emotion estimation function, for example, using an emotion engine or a generative AI. For example, the identification unit can input image data captured by a camera into an emotion engine and have the emotion engine perform emotion estimation. For example, if the participant is nervous, the identification unit will prioritize identifying words that are important to alleviate the tension. Also, if the participant is relaxed, the identification unit can prioritize identifying words of lower importance to maintain relaxation. Furthermore, if the participant is confused, the identification unit can prioritize identifying words related to resolving the confusion. In this way, important words can be prioritized by determining the priority of words based on the participant's emotions. Some or all of the above processing in the identification unit may be performed using AI, for example, or without AI. For example, the identification unit can input the participant's emotion data into a generative AI and have the generative AI perform the determination of word priority.

[0083] The identification unit can adjust specific priorities based on the progress of the meeting at the time of identification. For example, in the early stages of the meeting, the identification unit may prioritize identifying basic words. In the middle of the meeting, for example, the identification unit may prioritize identifying words related to important topics. Furthermore, in the final stages of the meeting, the identification unit may also prioritize identifying general words. This allows for the identification of words at the appropriate time by adjusting specific priorities based on the progress of the meeting. Some or all of the above processing in the identification unit may be performed using AI, for example, or not using AI. For example, the identification unit can input meeting progress data into the AI ​​and have the AI ​​perform the adjustment of specific priorities.

[0084] The identification unit can improve the accuracy of identification by considering the expertise level of the meeting participants during the identification process. For example, the identification unit can perform detailed identification for participants with high expertise, simple identification for participants with low expertise, and moderate identification for participants with moderate expertise. This improves the accuracy of identification by considering the expertise level of the participants. Some or all of the above processing in the identification unit may be performed using AI, for example, or without AI. For example, the identification unit can input participant expertise level data into AI and have the AI ​​perform the improvement of identification accuracy.

[0085] The generation unit can estimate the participant's emotions and adjust the explanation generation method based on the estimated participant's emotions. The generation unit estimates the participant's emotions using an emotion estimation function, for example, using an emotion engine or a generation AI. For example, the generation unit can input image data captured by a camera into the emotion engine and have the emotion engine perform emotion estimation. For example, if the participant is nervous, the generation unit can generate a concise and easy-to-understand explanation. The generation unit can also generate a detailed and in-depth explanation if the participant is relaxed. Furthermore, if the participant is confused, the generation unit can generate an explanation that includes specific examples. In this way, appropriate explanations can be generated by adjusting the explanation generation method based on the participant's emotions. Some or all of the above processing in the generation unit may be performed using a generation AI, for example, or without a generation AI. For example, the generation unit can input the participant's emotion data into the generation AI and have the generation AI adjust the explanation generation method.

[0086] The generation unit can adjust the level of detail in the explanations based on the importance of the words during generation. For example, the generation unit can generate detailed explanations for important words, and concise explanations for less important words. It can also generate explanations with a moderate level of detail for words of moderate importance. By adjusting the level of detail in the explanations based on the importance of the words, appropriate explanations can be generated. Some or all of the above processing in the generation unit may be performed using a generation AI, for example, or without a generation AI. For example, the generation unit can input word importance data into a generation AI and have the generation AI perform the adjustment of the level of detail in the explanations.

[0087] The generation unit can apply different generation algorithms depending on the category of the word during generation. For example, the generation unit can apply an algorithm that generates technical explanations to technical terms. For example, the generation unit can apply an algorithm that generates business explanations to business terms. Furthermore, the generation unit can also apply an algorithm that generates educational explanations to educational terms. By applying the appropriate generation algorithm according to the category of the word, the accuracy of the explanation is improved. Some or all of the above processing in the generation unit may be performed using a generation AI, for example, or without a generation AI. For example, the generation unit can input word category data into a generation AI and have the generation AI perform the application of the generation algorithm.

[0088] The generation unit can estimate the participant's emotions and adjust the length of the explanation based on the estimated emotions. The generation unit estimates the participant's emotions using an emotion estimation function, for example, using an emotion engine or a generation AI. For example, the generation unit can input image data captured by a camera into an emotion engine and have the emotion engine perform emotion estimation. For example, if the participant is nervous, the generation unit can generate a short, concise explanation. The generation unit can also generate a detailed explanation if the participant is relaxed. Furthermore, if the participant is confused, the generation unit can generate an explanation that includes specific examples. This allows for the generation of an appropriate explanation by adjusting the length of the explanation based on the participant's emotions. Some or all of the above processing in the generation unit may be performed using a generation AI, for example, or without a generation AI. For example, the generation unit can input the participant's emotion data into a generation AI and have the generation AI adjust the length of the explanation.

[0089] The generation unit can determine the priority of explanations based on word frequency during generation. For example, the generation unit can prioritize generating explanations for frequently used words. For example, it can postpone generating explanations for less frequently used words. The generation unit can also moderately prioritize generating explanations for words with moderate frequency. In this way, by determining the priority of explanations based on word frequency, explanations for important words can be generated preferentially. Some or all of the above processing in the generation unit may be performed using, for example, a generation AI, or without a generation AI. For example, the generation unit can input word frequency data into a generation AI and have the generation AI perform the determination of explanation priority.

[0090] The generation unit can adjust the order of explanations based on the relevance of words during generation. For example, the generation unit can prioritize generating explanations for highly relevant words. For example, it can postpone generating explanations for less relevant words. It can also moderately prioritize generating explanations for words with a moderate level of relevance. In this way, by adjusting the order of explanations based on the relevance of words, explanations can be provided in an appropriate order. Some or all of the above processing in the generation unit may be performed using, for example, a generation AI, or without a generation AI. For example, the generation unit can input word relevance data into a generation AI and have the generation AI perform the adjustment of the explanation order.

[0091] The display unit can estimate the participant's emotions and adjust the display method based on the estimated emotions. The display unit estimates the participant's emotions using an emotion estimation function, for example, using an emotion engine or a generative AI. For example, the display unit can input image data captured by a camera into the emotion engine and have the emotion engine perform emotion estimation. For example, if the participant is nervous, the display unit provides a simple and highly visible display method. The display unit can also provide a display method that includes detailed information if the participant is relaxed. Furthermore, if the participant is confused, the display unit can provide a display method that includes specific examples. In this way, an appropriate display can be provided by adjusting the display method based on the participant's emotions. Some or all of the above processing in the display unit may be performed using AI, for example, or without AI. For example, the display unit can input the participant's emotion data into a generative AI and have the generative AI perform the adjustment of the display method.

[0092] The display unit can select the optimal display method by referring to the participant's past operation history when displaying information. For example, the display unit may prioritize display methods that the participant has preferred to use in the past. For example, the display unit may select the most efficient display method from the participant's past operation history. The display unit can also exclude display methods that the participant has avoided in the past and provide the optimal display method. In this way, the optimal display method can be selected by referring to past operation history. Some or all of the above processing in the display unit may be performed using AI, for example, or without AI. For example, the display unit may input past operation history data into AI and have the AI ​​perform the selection of the optimal display method.

[0093] The display unit can select the optimal display method when displaying information, taking into account the participant's device information. For example, if a participant is using a smartphone, the display unit provides a display method that matches the screen size. For example, if a participant is using a tablet, the display unit provides a display method optimized for a larger screen. Furthermore, if a participant is using a desktop computer, the display unit can also provide a display method that includes detailed information. In this way, the optimal display method can be provided by taking device information into consideration. Some or all of the above processing in the display unit may be performed using AI, for example, or without AI. For example, the display unit can input device information into AI and have the AI ​​select the optimal display method.

[0094] The display unit can estimate the emotions of participants and determine the display priority based on the estimated emotions. The display unit estimates the emotions of participants using an emotion estimation function, for example, using an emotion engine or a generative AI. For example, the display unit can input image data captured by a camera into an emotion engine and have the emotion engine perform emotion estimation. For example, if a participant is nervous, the display unit can prioritize displaying important information. Also, if a participant is relaxed, the display unit can prioritize displaying detailed information. Furthermore, if a participant is confused, the display unit can prioritize displaying specific examples. In this way, important information can be prioritized by determining the display priority based on the participant's emotions. Some or all of the above processing in the display unit may be performed using AI, for example, or without using AI. For example, the display unit can input participant emotion data into a generative AI and have the generative AI perform the determination of the display priority.

[0095] The display unit can adjust the timing of its display based on the progress of the meeting. For example, in the early stages of the meeting, the display unit prioritizes displaying basic information. In the middle of the meeting, for example, the display unit prioritizes displaying information related to important topics. Furthermore, in the final stages of the meeting, the display unit can also prioritize displaying summary information. By adjusting the timing of the display based on the progress of the meeting, information can be displayed at the appropriate time. Some or all of the above processing in the display unit may be performed using AI, for example, or without AI. For example, the display unit can input meeting progress data into the AI ​​and have the AI ​​perform the adjustment of the display timing.

[0096] The display unit can customize the displayed content to take into account the participants' level of expertise. For example, the display unit can display detailed information to participants with high expertise, simplified information to participants with low expertise, and information of an appropriate level of detail to participants with moderate expertise. This allows for the display of appropriate content by considering the participants' level of expertise. Some or all of the above processing in the display unit may be performed using AI, for example, or without AI. For example, the display unit can input participant expertise level data into AI and have the AI ​​perform the customization of the displayed content.

[0097] The system according to the embodiment is not limited to the example described above, and various modifications are possible, for example, as follows.

[0098] Online meeting support systems can also be equipped with a speech recognition unit. This unit analyzes audio data during the meeting in real time, detecting specific keywords and phrases. For example, if a participant says, "I don't understand this part," the speech recognition unit will detect this statement and identify the unclear part. The speech recognition unit can also analyze the tone and speed of the participant's speech to estimate their emotions. For instance, a participant who is nervous often speaks more quickly. This allows the speech recognition unit to analyze the content and emotions of the participant's speech and improve the accuracy of identifying the unclear parts. Furthermore, the speech recognition unit can prioritize the detection of specific keywords and phrases according to the progress of the meeting. For example, it might prioritize basic keywords in the early stages of the meeting and keywords related to important topics in the middle stages. This enables appropriate support tailored to the meeting's progress.

[0099] Online meeting support systems can also include a translation function. This function translates spoken content in real time during meetings, facilitating communication between participants who speak different languages. For example, if English-speaking and Japanese-speaking participants are in the same meeting, the translation function will translate English statements into Japanese and Japanese statements into English. The translation function can also include a dictionary database to appropriately translate specialized and industry-specific terminology, enabling accurate translation even in meetings with highly specialized content. Furthermore, the translation function can estimate participants' emotions and adjust the tone and nuances of the translation based on those emotions. For example, if a participant is nervous, the translation tone can be softened to facilitate smoother communication.

[0100] Online meeting support systems can also include a note-taking function. This function automatically records important statements and key points of discussion during a meeting and provides them to participants after the meeting. For example, it can automatically extract important decisions and action items made during the meeting and compile them into notes. The note-taking function can also estimate participants' emotions and adjust the content of the notes based on those emotions. For example, if a participant is confused, it can record that part of the discussion in detail for later review. Furthermore, the note-taking function can update the notes in real time according to the progress of the meeting, allowing participants to check them during the meeting. This allows for efficient recording of meeting content and provides information that is useful for participants to review later.

[0101] Online meeting support systems can also include a reminder function. This function automatically records action items and important tasks discussed during the meeting and sends reminders to participants. For example, when a deadline for a task decided during the meeting approaches, the reminder function sends a notification to participants. Furthermore, the reminder function can estimate participants' emotions and adjust the content and timing of reminders based on those emotions. For instance, if a participant is feeling stressed, the reminder tone can be softened and notifications sent at the appropriate time. Additionally, the reminder function can update reminders in real time according to the meeting's progress, allowing participants to check them during the meeting. This enables efficient management of action items and tasks, maximizing meeting outcomes.

[0102] Online meeting support systems can also include a feedback function. This function collects feedback from participants after the meeting and analyzes areas for improvement and successes. For example, participants evaluate the meeting's progress and content, and the system compiles these evaluations to create a report. The feedback function can also estimate participants' emotions and adjust the feedback based on those emotions. For instance, if participants are satisfied, it can emphasize positive feedback reflecting that emotion. Furthermore, the feedback function can collect feedback in real time as the meeting progresses, allowing participants to express their opinions during the meeting. This can improve the quality of the meeting and increase participant satisfaction.

[0103] Online meeting support systems can also include a data analysis department. This department analyzes data collected during meetings to evaluate meeting performance and participant contributions. For example, it can analyze the number of contributions and speaking time to visualize each participant's contribution. The data analysis department can also analyze the meeting's progress and discussion content to evaluate its effectiveness. For instance, it can assess the depth and breadth of discussions to quantitatively evaluate the meeting's quality. Furthermore, by comparing current meeting data with past meeting data, the data analysis department can identify areas for improvement and make suggestions for future meetings. This allows for improved meeting performance and a better evaluation of participant contributions.

[0104] Online meeting support systems can also include a scheduling unit. This unit automatically adjusts participants' schedules and suggests the optimal meeting time. For example, it analyzes all participants' calendars to identify time slots when everyone is available. The scheduling unit can also adjust scheduling priorities based on the importance and urgency of the meeting. For instance, it prioritizes urgent meetings and postpones less important ones. Furthermore, it can analyze participants' past schedule data to suggest the optimal meeting time. This reduces the effort required for scheduling and enables more efficient meetings.

[0105] Online meeting support systems can also include a document sharing function. This function automatically collects and shares documents used during the meeting with participants. For example, it can collect the meeting agenda and presentation materials in advance and distribute them to participants. The document sharing function can also share newly added documents in real time during the meeting. For example, if a participant uploads new documents during the meeting, the document sharing function will distribute those documents to everyone. Furthermore, the document sharing function can archive past meeting materials so that participants can refer to them later. This ensures efficient document sharing and smoother meeting progress.

[0106] Online meeting support systems can also include a voting function. This function provides a voting mechanism to gather participants' opinions during a meeting. For example, it can be used to conduct a vote to gather everyone's opinion on important decisions. The voting function can also compile voting results in real time and display them during the meeting. For instance, it can display results immediately after voting ends to support rapid decision-making. Furthermore, the voting function can analyze past voting data to understand trends in participants' opinions. This allows for efficient collection of participant feedback and supports decision-making in meetings.

[0107] Online meeting support systems can also include a breakout room feature. This feature allows participants to be divided into smaller groups for discussion during a meeting. For example, in a large meeting, participants can be divided into multiple breakout rooms to delve deeper into a specific topic. The breakout room feature can also monitor the progress of each room and provide support as needed. For instance, if a discussion in a room stalls, a facilitator can intervene to stimulate the discussion. Furthermore, the breakout room feature can record the discussion content from each room and share it with the general meeting afterward. This allows for more efficient discussions and maximizes the meeting's effectiveness.

[0108] The following briefly describes the processing flow for example form 2.

[0109] Step 1: The detection unit detects the participant's facial expressions. These expressions may include, but are not limited to, smiles, confusion, or surprise. The detection unit uses facial recognition technology and facial expression recognition algorithms to analyze image data captured by the camera and detect the participant's facial expressions in real time. Step 2: The identification unit identifies unfamiliar words based on the facial expressions detected by the detection unit. The identification unit analyzes the patterns of facial expression changes to identify unfamiliar words. It also performs contextual analysis using a language model, combining the participant's facial expression data with the conversation context data to identify unfamiliar words. Step 3: The generation unit generates explanations for the words identified by the identification unit. The generation unit uses a generation AI (e.g., a text generation AI) to generate explanations for unfamiliar words. The generation unit also analyzes the context containing the unfamiliar words, understands their meaning, and generates appropriate explanations. Step 4: The display unit displays the explanation generated by the generation unit on the screen. The display unit displays the generated explanation according to the participant's level. For example, it explains in simple terms to new graduates and provides a detailed explanation to participants with specialized knowledge. It also displays the generated explanation as a pop-up on the screen in real time.

[0110] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0111] Data generation model 58 is a form of so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> Examples of generative AI include text generation AI, image generation AI, and multimodal generation AI. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats from audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVMs), k-means clustering, convolutional neural networks (CNNs), recurrent neural networks (RNNs), generative adversarial networks (GANs), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI ​​may be an AI agent. Furthermore, when the processing of each of the above parts is performed by the AI, the processing may be performed by the AI ​​in part or in whole, but is not limited to this example.Furthermore, processing performed by AI, including generative AI, may be replaced with rule-based processing, and rule-based processing may be replaced with processing performed by AI, including generative AI.

[0112] Furthermore, the processing performed by the data processing system 10 described above is carried out by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart device 14, but it may also be carried out by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart device 14. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the smart device 14 or an external device, and the smart device 14 acquires or collects information necessary for processing from the data processing device 12 or an external device.

[0113] Each of the multiple elements described above, including the detection unit, identification unit, generation unit, and display unit, is implemented in at least one of the smart device 14 and the data processing device 12. For example, the detection unit uses the camera 42 of the smart device 14 to detect the participant's facial expressions and the control unit 46A analyzes them in real time. The identification unit is implemented in the identification processing unit 290 of the data processing device 12 and analyzes the detected facial expression data and conversation context data to identify unfamiliar words. The generation unit is implemented in the identification processing unit 290 of the data processing device 12 and generates explanations for unfamiliar words using a generation AI. The display unit displays the generated explanations on the screen using the display 40A of the smart device 14. The correspondence between each unit and the device or control unit is not limited to the example described above and can be modified in various ways.

[0114] [Second Embodiment] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0115] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0116] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.

[0117] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0118] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0119] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).

[0120] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0121] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing by the processor 28. The storage 32 stores the specific processing program 56.

[0122] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0123] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.

[0124] In the smart glasses 214, specific processing is performed by the processor 46. The storage 50 stores a specific processing program 60. The processor 46 reads the specific processing program 60 from the storage 50 and executes the read specific processing program 60 on the RAM 48. The specific processing is realized by the processor 46 acting as a control unit 46A according to the specific processing program 60 executed on the RAM 48. The smart glasses 214 also have a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.

[0125] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).

[0126] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0127] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI ​​may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI ​​in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.

[0128] The data processing system 210 according to the second embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 210 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart glasses 214, but it may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart glasses 214. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the smart glasses 214 or an external device, and the smart glasses 214 acquires or collects information necessary for processing from the data processing device 12 or an external device.

[0129] Each of the multiple elements described above, including the detection unit, identification unit, generation unit, and display unit, is implemented in at least one of the smart glasses 214 and the data processing unit 12. For example, the detection unit uses the camera 42 of the smart glasses 214 to detect the participant's facial expressions and the control unit 46A analyzes them in real time. The identification unit is implemented in the identification processing unit 290 of the data processing unit 12 and analyzes the detected facial expression data in combination with conversation context data to identify unfamiliar words. The generation unit is implemented in the identification processing unit 290 of the data processing unit 12 and generates explanations of unfamiliar words using generation AI. The display unit displays the generated explanations on the screen using the display of the smart glasses 214. The correspondence between each unit and the device or control unit is not limited to the example described above and can be modified in various ways.

[0130] [Third Embodiment] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0131] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0132] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.

[0133] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0134] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0135] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).

[0136] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0137] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0138] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0139] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.

[0140] In the headset terminal 314, specific processing is performed by the processor 46. The storage 50 stores a specific program 60. The processor 46 reads the specific program 60 from the storage 50 and executes the read specific program 60 on the RAM 48. The specific processing is realized by the processor 46 acting as a control unit 46A according to the specific program 60 executed on the RAM 48. The headset terminal 314 also has a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.

[0141] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).

[0142] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0143] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI ​​may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI ​​in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.

[0144] The data processing system 310 according to the third embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 310 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the headset terminal 314, but may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the headset terminal 314. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the headset terminal 314 or an external device, and the headset terminal 314 acquires or collects information necessary for processing from the data processing device 12 or an external device.

[0145] Each of the multiple elements described above, including the detection unit, identification unit, generation unit, and display unit, is implemented in at least one of the headset terminal 314 and the data processing unit 12. For example, the detection unit uses the camera 42 of the headset terminal 314 to detect the participant's facial expressions and the control unit 46A analyzes them in real time. The identification unit is implemented in the identification processing unit 290 of the data processing unit 12, which analyzes the detected facial expression data and conversation context data in combination to identify unfamiliar words. The generation unit is implemented in the identification processing unit 290 of the data processing unit 12, which generates explanations of unfamiliar words using a generation AI. The display unit displays the generated explanations on the screen using the display 343 of the headset terminal 314. The correspondence between each unit and the device or control unit is not limited to the example described above and can be modified in various ways.

[0146] [Fourth Embodiment] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0147] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0148] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.

[0149] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0150] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0151] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS image sensor or CCD image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).

[0152] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0153] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. The robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0154] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0155] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0156] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.

[0157] In robot 414, specific processing is performed by processor 46. A specific program 60 is stored in storage 50. Processor 46 reads the specific program 60 from storage 50 and executes it on RAM 48. The specific processing is achieved by processor 46 acting as a control unit 46A according to the specific program 60 executed on RAM 48. Robot 414 also has data generation model 58 and emotion identification model 59, similar to those of the robot, and can perform processing similar to that of the specific processing unit 290 using these models.

[0158] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).

[0159] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0160] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI ​​may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI ​​in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.

[0161] The data processing system 410 according to the fourth embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 410 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the robot 414, but it may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the robot 414. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the robot 414 or an external device, and the robot 414 acquires or collects information necessary for processing from the data processing device 12 or an external device.

[0162] Each of the multiple elements described above, including the detection unit, identification unit, generation unit, and display unit, is implemented in at least one of the robot 414 and the data processing unit 12. For example, the detection unit uses the camera 42 of the robot 414 to detect the participant's facial expressions and the control unit 46A analyzes them in real time. The identification unit is implemented in the identification processing unit 290 of the data processing unit 12, which analyzes the detected facial expression data and conversation context data in combination to identify unfamiliar words. The generation unit is implemented in the identification processing unit 290 of the data processing unit 12, which generates explanations of unfamiliar words using a generation AI. The display unit displays the generated explanations on the screen using the display of the robot 414. The correspondence between each unit and the device or control unit is not limited to the example described above and can be modified in various ways.

[0163] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0164] Figure 9 shows the emotion map 400, in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0165] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0166] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0167] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, and motorcycles, emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated based, for example, on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0168] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0169] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0170] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing method for the specific process may be used, which includes computer 22 and multiple other computers.

[0171] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0172] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0173] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0174] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0175] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0176] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0177] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0178] Furthermore, although the above-described examples were divided into four embodiments, some or all of these embodiments may be combined. Also, the smart device 14, smart glasses 214, headset terminal 314, and robot 414 are just examples, and they may be combined, or other devices may be used. Also, although the above-described examples were divided into two embodiments, Embodiment 1 and Embodiment 2, these may be combined.

[0179] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and other things that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0180] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.

[0181] (Note 1) A detection unit that detects the facial expressions of the participants, An identification unit identifies an unknown word based on the facial expression detected by the detection unit, A generation unit that generates an explanation of the word identified by the identification unit, The system includes a display unit that displays the explanation generated by the generation unit on the screen. A system characterized by the following features. (Note 2) The generating unit is Use generative AI to generate explanations for unfamiliar words. The system described in Appendix 1, characterized by the features described herein. (Note 3) The aforementioned display unit is Display the generated explanation according to the participant's level. The system described in Appendix 1, characterized by the features described herein. (Note 4) The detection unit is Detecting the emotion of "I don't understand" from the participants' facial expressions. The system described in Appendix 1, characterized by the features described herein. (Note 5) The specified part is, Identify unknown words based on detected facial expressions. The system described in Appendix 1, characterized by the features described herein. (Note 6) The generating unit is Analyze the context containing unfamiliar words and understand the meaning of those words. The system described in Appendix 1, characterized by the features described herein. (Note 7) The detection unit is The system estimates the emotions of the participants and adjusts the accuracy of facial expression detection based on the estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 8) The detection unit is The system analyzes participants' past facial expression data and selects the optimal detection algorithm. The system described in Appendix 1, characterized by the features described herein. (Note 9) The detection unit is When detecting facial expressions, filtering is performed based on the participant's current environment and situation. The system described in Appendix 1, characterized by the features described herein. (Note 10) The detection unit is The system estimates the emotions of the participants and determines the priority of facial expressions to detect based on the estimated emotions of the participants. The system described in Appendix 1, characterized by the features described herein. (Note 11) The detection unit is When detecting facial expressions, the system prioritizes detecting expressions that are highly relevant, taking into account the participant's geographical location. The system described in Appendix 1, characterized by the features described herein. (Note 12) The detection unit is When detecting facial expressions, the system analyzes the participant's social media activity and identifies relevant facial expressions. The system described in Appendix 1, characterized by the features described herein. (Note 13) The specified part is, The system estimates the participants' emotions and adjusts the accuracy of identifying unfamiliar words based on the estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 14) The specified part is, At specific times, the algorithm for identifying unfamiliar words is optimized by referring to past meeting data. The system described in Appendix 1, characterized by the features described herein. (Note 15) The specified part is, At specific times, different specific methods are applied depending on the context and topic of the meeting. The system described in Appendix 1, characterized by the features described herein. (Note 16) The specified part is, The system estimates the emotions of the participants and determines the priority of words to identify based on the estimated emotions of the participants. The system described in Appendix 1, characterized by the features described herein. (Note 17) The specified part is, At specific times, adjust certain priorities based on the progress of the meeting. The system described in Appendix 1, characterized by the features described herein. (Note 18) The specified part is, At specific times, improve certain accuracy by taking into account the expertise level of meeting participants. The system described in Appendix 1, characterized by the features described herein. (Note 19) The generating unit is We estimate the participants' emotions and adjust the explanation generation method based on the estimated participants' emotions. The system described in Appendix 1, characterized by the features described herein. (Note 20) The generating unit is During generation, adjust the level of detail in the explanation based on the importance of the words. The system described in Appendix 1, characterized by the features described herein. (Note 21) The generating unit is During generation, different generation algorithms are applied depending on the word category. The system described in Appendix 1, characterized by the features described herein. (Note 22) The generating unit is The system estimates the participants' emotions and adjusts the length of the explanation based on the estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 23) The generating unit is During generation, the priority of explanations is determined based on the frequency of word usage. The system described in Appendix 1, characterized by the features described herein. (Note 24) The generating unit is During generation, the order of explanations is adjusted based on the relevance of the words. The system described in Appendix 1, characterized by the features described herein. (Note 25) The aforementioned display unit is The system estimates the participants' emotions and adjusts the display method based on the estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 26) The aforementioned display unit is When displaying information, the system will refer to the participant's past operation history to select the most suitable display method. The system described in Appendix 1, characterized by the features described herein. (Note 27) The aforementioned display unit is When displaying the information, the system selects the optimal display method, taking into account the participants' device information. The system described in Appendix 1, characterized by the features described herein. (Note 28) The aforementioned display unit is The system estimates the participants' emotions and determines the display priority based on the estimated emotions of the participants. The system described in Appendix 1, characterized by the features described herein. (Note 29) The aforementioned display unit is When displaying information, the timing of the display will be adjusted based on the progress of the meeting. The system described in Appendix 1, characterized by the features described herein. (Note 30) The aforementioned display unit is When displaying information, the content displayed will be customized to take into account the participants' level of expertise. The system described in Appendix 1, characterized by the features described herein. [Explanation of symbols]

[0182] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots

Claims

1. A detection unit that detects the facial expressions of the participants, An identification unit identifies an unknown word based on the facial expression detected by the detection unit, A generation unit that generates an explanation of the word identified by the identification unit, The system includes a display unit that displays the explanation generated by the generation unit on the screen. A system characterized by the following features.

2. The generating unit is Use generative AI to generate explanations for unfamiliar words. The system according to feature 1.

3. The aforementioned display unit is Display the generated explanation according to the participant's level. The system according to feature 1.

4. The detection unit is Detecting the emotion of "I don't understand" from the participant's facial expression. The system according to feature 1.

5. The specified part is, Identify unknown words based on detected facial expressions. The system according to feature 1.

6. The generating unit is Analyze the context containing unfamiliar words and understand the meaning of those words. The system according to feature 1.

7. The detection unit is The system estimates the emotions of the participants and adjusts the accuracy of facial expression detection based on the estimated emotions. The system according to feature 1.

8. The detection unit is The system analyzes participants' past facial expression data and selects the optimal detection algorithm. The system according to feature 1.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A