Meeting session control based on attention detection.
The system addresses the challenge of maintaining attendee attention in remote learning and meetings by using machine learning to analyze participant activities and generate personalized recommendations, improving engagement and learning outcomes.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- SONY GROUP CORP
- Filing Date
- 2023-01-24
- Publication Date
- 2026-05-15
AI Technical Summary
Existing remote learning and meeting systems fail to effectively maintain attendee attention levels, focusing primarily on content delivery rather than participant engagement.
A system and method for meeting session control based on attention determination, utilizing machine learning models trained on attendee images and activities to generate real-time recommendations for enhancing attention, including personalized models considering historical data and environmental factors.
Enhances attendee engagement by providing real-time recommendations that improve attention levels, allowing educators to adapt content and interactions, and providing detailed reports for individual and group analysis.
Smart Images

Figure 0007859007000011 
Figure 0007859007000012 
Figure 0007859007000013
Abstract
Description
[Technical Field]
[0001] [Incorporation by cross-referencing / citation with related applications] This application claims priority to U.S. Patent Application No. 17 / 671,495, filed with the U.S. Patent and Trademark Office on 14 February 2022. Each of the above applications is incorporated herein by reference in its entirety.
[0002] Various embodiments of this disclosure relate to meeting sessions. Specifically, various embodiments of this disclosure relate to systems and methods for meeting session control based on attention determination. [Background technology]
[0003] Advances in the field of machine learning have led to multiple applications in various aspects of daily life. With the outbreak of certain pandemics, almost all educational institutions, offices, and meeting spaces were closed or significantly impacted. Consequently, many students and / or employees were forced to attend remote learning and / or meetings via the internet. However, remote learning and remote meetings can present different challenges, such as (but not limited to) the difficulty in maintaining each student's or employee's level of attention. Some solutions focus on content delivery rather than participant attention. Therefore, there is a need for impactful, interactive systems that take into account the attention levels of attendees present within a meeting session and further enhance the overall learning process for attendees. [Overview of the project] [Problems that the invention aims to solve]
[0004] Those skilled in the art will be able to see the limitations and disadvantages of conventional methods by comparing the described system with some aspects of the disclosure shown with reference to the drawings in the remainder of this application. [Means for solving the problem]
[0005] The present invention provides a system and method for meeting session control based on attention determination, as illustrated and / or described in substantially relation to at least one figure and more fully provided in the claims.
[0006] These and other features and advantages of the disclosure can be understood by considering the following detailed description of the disclosure with reference to the accompanying drawings, which indicate the same elements throughout by the same reference numerals. [Brief explanation of the drawing]
[0007] [Figure 1] This block diagram shows an exemplary network environment for meeting session control based on attention determination according to an embodiment of the present disclosure. [Figure 2] This block diagram shows an exemplary system for meeting session control based on attention determination according to an embodiment of the present disclosure. [Figure 3] This figure shows an exemplary operation of training an ML model for each of multiple attendees associated with multiple meeting sessions, according to an embodiment of the present disclosure. [Figure 4] This figure shows exemplary behavior for meeting session control based on attention determination and the application of an ML model trained for specific attendees, according to embodiments of the present disclosure. [Figure 5] This figure shows an exemplary scenario for generating a simulated view of multiple attendees related to a meeting session, according to embodiments of the present disclosure. [Figure 6] This figure shows an exemplary user interface for rendering dashboard information according to an embodiment of the present disclosure. [Figure 7] This flowchart shows an exemplary operation of training an ML model for meeting session control based on attention determination, according to an embodiment of the present disclosure. [Figure 8] This flowchart shows exemplary operation for meeting session control based on attention determination and application of ML models for specific attendees, according to embodiments of the present disclosure. [Modes for carrying out the invention]
[0008] The disclosed systems and methods for meeting session control based on attention determination may include embodiments described below. Such methods can be configured to train a machine learning (ML) model for attendees (such as students or employees) based on the level of attendee attention determined in different meeting sessions (such as classes or meetings in offline or online mode) when run on the system. The system can also generate one or more recommendations in real time to enhance the attendee's attention level based on the application of the trained ML model associated with the corresponding attendee. The system can receive multiple images of multiple attendees related to multiple meeting sessions in order to train the ML model. The multiple images can be received from different imaging devices (such as cameras) placed during the meeting sessions. The meeting categories of one or more of the multiple meeting sessions may be different. The meeting category may be, for example, the type of meeting (i.e., class / offline-based or online-based), the number of attendees in the meeting session, the duration of the meeting session, the average age of the attendees, the topic of the meeting session, the content presented during the meeting session, or the experience of the educator in the meeting session. Based on the received images, the system can further detect one or more activities performed by each of the attendees during a corresponding meeting session (e.g., raising a hand, reading aloud, writing, typing, talking, standing up, asking / answering a question). One or more activities can be detected over a period of time (e.g., over several minutes / hours in a particular meeting session). Based on the one or more detected activities associated with each attendee, the system can further calculate attention scores for each of the attendees over the corresponding period. The attention scores can indicate each attendee's level of attention during the corresponding meeting session.The system can be further configured to train a machine learning (ML) model for each of multiple attendees based on the calculated attention score for each of the multiple attendees and the meeting category of the corresponding meeting session. The disclosed system can train multiple ML models for multiple attendees according to the calculated attention scores for corresponding meeting sessions in different meeting categories. Thus, the disclosed system can be configured to generate and train a personalized machine learning (ML) model for each attendee. The disclosed system can train an ML model based on the level of attention determined during multiple meeting sessions in the same or different meeting categories (i.e., the level of attention calculated based on one or more activities performed by the attendee). The generated ML model can also be further trained based on several other factors, such as the attendee's facial expressions in different meeting sessions, experience information related to the educator, environmental information (such as lighting conditions in the meeting session), and the attendee's preferences / interests. Thus, a personalized ML model can be trained for a specific attendee based on the attention scores calculated (or tracked) for attendees in different meeting sessions with different meeting categories and different other relevant factors. In other words, a trained ML model for attendees can show historical data and fluctuations / patterns related to attention scores that can be determined for attendees in past meeting sessions of different categories and meeting sessions with other relevant factors. Therefore, the disclosed system can propose a characteristic learning model for each student (i.e., attendee) by processing all historical data (related to attention scores) generated from a series of classroom activities over a certain period. This data is useful for performing a cumulative analysis of each student, and thus can identify each student's strengths and weaknesses. The disclosed system can further consider students' tendencies in certain subjects and their responsiveness to teaching methods.
[0009] The disclosed system can, based on the generation of an ML model, further apply the trained ML model to attendee images or calculated attention scores (i.e., those related to the current or real-time meeting session) to further output or control a set of recommendations for attendees, educators, or content presented in the meeting session. One or more recommendations, if followed, can increase the attention level of attendees in the current or upcoming meeting session. Based on a real-time analysis of the attention level and the output recommendations, the corresponding meeting session can be controlled to increase the attention level of attendees.
[0010] The disclosed system can be an artificial intelligence (AI)-based intelligent system that aims to engage students (i.e., attendees) in the learning process (without interfering with the essence of a normal class) by enabling educators in meeting sessions to effectively transmit / impart knowledge to multiple attendees. The disclosed system can also be implemented as a smart classroom mechanism that monitors each attendee's attention level (using cameras and depth sensors) and aims to further improve the overall learning process in real time based on the monitored attention levels of each attendee, as well as recommendations for output from attendees, educators, or presented content. Furthermore, the disclosed system can utilize cumulative data (i.e., an ML model trained on determined attention scores from different meeting sessions) to evaluate the quality of lectures and instructors in meeting sessions using real-time analysis of attendee attention levels. In addition, the disclosed system can provide detailed time-varied individual reports for each attendee, in addition to cumulative group analysis. Thus, the disclosed system can focus on each individual attendee to improve their learning process. Furthermore, the disclosed system can generate various reports and statistics based on multiple parameters such as subjects and grades, and perform comparative analysis between each attendee and other attendees present in the meeting session. Additionally, the disclosed system can detect illegal activities (such as cheating during exams) that may occur by attendees or educators during the meeting session. The disclosed system may also have a predictive mechanism that can adapt to each student's learning pattern.
[0011] FIG. 1 is a block diagram showing an exemplary network environment for meeting session control based on attention determination according to an embodiment of the present disclosure. FIG. 1 shows a network environment 100. The network environment 100 can include a system 102, a plurality of image capture devices 104, a plurality of machine learning (ML) models 106, an audio capture device 108, a server 110, and a communication network 112. FIG. 1 further shows a plurality of images 114 of a plurality of attendees 116.
[0012] The system 102 can include suitable logic, circuitry, interfaces, and / or code configured to receive a plurality of images 114 of a plurality of attendees 116 associated with a plurality of meeting sessions (such as the meeting session 118 shown in FIG. 1) in which the corresponding attendees are present. The system 102 can calculate an attention score for each of the plurality of attendees based on the received plurality of images 114, and can be configured to further train the plurality of ML models 106 for the plurality of attendees 116 based on the calculated attention scores for each of the plurality of attendees 116. Examples of the system 102 can include, but are not limited to, computer devices such as educational engines, computer workstations, mainframe machines, servers, smartphones, cellular phones, personal computers with or without a graphics processing unit (GPU), imaging devices having processing capabilities, and / or consumer electronics (CE) devices. In certain embodiments, the system 102 can also be implemented as a plugin that can be added to existing video conferencing software programs and applications.
[0013] Each of the plurality of image capture devices 104 can include suitable logic, circuitry, and interfaces configured to capture a plurality of images 114 of a plurality of attendees 116. Each of the plurality of image capture devices 104 can further be configured to transmit the plurality of captured images 114 to the system 102. Examples of each of the plurality of image capture devices 104 include, but are not limited to, image sensors, closed-circuit television (CCTV) cameras, webcams (or web cameras), wide-angle cameras, action cameras, camcorders, digital cameras, mobile phones with cameras, time-of-flight cameras (ToF cameras), night vision cameras, and / or other image capture devices. In certain embodiments, each of the plurality of image capture devices 104 can include depth sensors configured to capture depth information / plurality of depth values of the plurality of attendees 116 from a single viewpoint or multiple viewpoints in a corresponding (e.g., class) meeting session.
[0014] Each of the plurality of machine learning (ML) models 106 can be an untrained classifier / regression / clustering model that needs to be trained to identify relationships between inputs such as features within a training dataset and output a set of recommendations. Each of the plurality of ML models 106 can be defined by hyperparameters such as, for example, the number of weights, cost function, input size, and number of layers. The parameters of each of the plurality of ML models 106 can be adjusted to converge to a global minimum of the cost function of the corresponding ML model, and the weights can be updated accordingly. Each of the plurality of ML models 106 can be trained to output prediction / classification results for an input set after being trained for a number of epochs on feature information within the training dataset. The prediction results can indicate the class labels of each input of the input set. In certain embodiments, each of the plurality of ML models 106 can be trained for different attendees based on attention scores calculated from past meeting sessions having different meeting categories in which the corresponding attendees participated.
[0015] Each of the multiple ML Models 106 may include electronic data that can be implemented, for example, as a software component of an application executable on System 102. Each of the multiple ML Models 106 may rely on libraries, external scripts, or other logic / instructions for execution by a device such as a circuit. Each of the multiple ML Models 106 may include code and routines configured to enable a computer device such as System 102 to perform one or more actions to output a recommended set. In addition to or instead of this, each of the multiple ML Models 106 may also be implemented using hardware including a processor, a microprocessor (e.g., one or more actions that are performed or controlled), a field-programmable gate array (FPGA), or an application-specific integrated circuit (ASIC). Alternatively, in some embodiments, each of the multiple ML Models 106 may be implemented using a combination of hardware and software.
[0016] In one embodiment, each of the multiple ML models 106 can be implemented as a neural network model, such as a deep learning model. The neural network model can be defined by its hyperparameters and topology / architecture. For example, the neural network model can be a deep neural network-based model that has the number of nodes (or neurons), (single or multiplicative) activation functions, the number of weights, the cost function, the regularization function, the input size, the learning rate, and the number of layers. Such a model can be called a computational network, or a system of nodes (e.g., artificial neurons). In an implementation of a neural network, the nodes of the neural network model can be arranged in layers as defined by the neural network topology. These layers can include an input layer, one or more hidden layers, and an output layer. Each layer can include one or more nodes (or, for example, artificial neurons represented by circles). The outputs of all nodes in the input layer can be coupled to at least one node in the (single or multiplicative) hidden layers. Similarly, the input of each hidden layer can be coupled to the output of at least one node in the other layers of the model. The output of each hidden layer can be coupled to the input of at least one node in the other layers of the neural network model. The (single or multiplicative) nodes in the final layer can receive input from at least one hidden layer and output a result. The number of layers and the number of nodes in each layer can be determined from hyperparameters that can be set before, during, or after training the neural network model based on the training dataset.
[0017] Each node in a neural network model can correspond to a mathematical function (e.g., a sigmoid function or a rectified linear unit) with a set of parameters that can be tuned during model training. These parameters may include, for example, weight parameters and regularization parameters. Each node can compute an output using the mathematical function based on one or more inputs from nodes in other (one or multiple) layers of the neural network model (e.g., previous (one or multiple) layers). All or some nodes in a neural network model can correspond to the same or different mathematical functions.
[0018] In training a neural network model, one or more parameters of each node can be updated based on whether the output of the final layer for a given input (from the training dataset) matches the correct result based on the neural network model's loss function. This process can be repeated for the same or different inputs until the minimum value of the loss function is achieved and the training error is minimized. Several training methods are known in this field, including gradient descent, stochastic gradient descent, batch gradient descent, gradient boosting, and metaheuristic methods.
[0019] In some embodiments, each of the multiple ML models 106 may be based on a hybrid architecture of multiple deep neural networks (DNNs). Examples of each of the multiple ML models 106 include, but are not limited to, neural network models, or models based on one or more of the following: (single or multiple) regression methods, (single or multiple) instance-based methods, (single or multiple) regularization methods, (single or multiple) decision tree methods, (single or multiple) Bayesian methods, (single or multiple) clustering methods, association rule learning, and (single or multiple) dimensionality reduction methods. Examples of neural network models include, but are not limited to, artificial neural networks (ANNs), deep neural networks (DNNs), convolutional neural networks (CNNs), recurrent neural networks (RNNs), CNN-recurrent neural networks (CNN-RNNs), R-CNNs, Fast R-CNNs, Faster R-CNNs, residual neural networks (Res-Nets), feature pyramid networks (FPNs), and / or combinations thereof.
[0020] The audio capture device 108 may include suitable logic, circuitry, and / or interfaces that can be configured to capture verbal interactions (in the form of audio signals) between each of the multiple attendees 116 and the corresponding meeting session instructor. Examples of the audio capture device 108 include, but are not limited to, a recorder, electret microphone, dynamic microphone, carbon microphone, piezoelectric microphone, fiber microphone, MEMS microphone, or other microphones well known in the art. The audio capture device 108 may be positioned close to the attendees and instructor to capture verbal interactions during different meeting sessions.
[0021] Server 110 may include suitable logic, circuitry, interfaces, and code that can be configured to store multiple images 114 received by multiple attendees 116 related to multiple meeting sessions. In some embodiments, Server 110 may be configured to train and store each of multiple ML models 106 for each of the multiple attendees 116. In some embodiments, Server 110 may be configured to store content presented (or to be presented) during a meeting session, as well as to store profile information about different attendees and instructors. In some embodiments, Server 110 may be implemented as a cloud server capable of performing operations through web applications, cloud applications, HTTP requests, repository operations, file transfers, etc. Other examples of Server 110 include, but are not limited to, database servers, file servers, web servers, media servers, application servers, mainframe servers, cloud servers, or other types of servers. In one or more embodiments, Server 110 may be implemented as multiple distributed cloud-based resources by using multiple technologies well known to those skilled in the art. Those skilled in the art will understand that the scope of this disclosure may not be limited to implementations of the server 110 and system 102 as separate entities. In some embodiments, the functionality of the server 110 can be incorporated into the system 102, either entirely or at least partially, without departing from the scope of this disclosure.
[0022] The communication network 112 may include a communication medium that enables the system 102, multiple image acquisition devices 104, audio acquisition device 108, and server 110 to communicate with each other. The communication network 112 may be a wired or wireless communication network. Examples of the communication network 112 include, but are not limited to, the Internet, a cloud network, a Wireless Fidelity (Wi-Fi) network, a Personal Area Network (PAN), a Local Area Network (LAN), or a Metropolitan Area Network (MAN). Various devices within the network environment 100 may be configured to connect to the communication network 112 according to various wired and wireless communication protocols. Examples of such wired and wireless communication protocols include, but are not limited to, at least one of the following: Transmission Control Protocol and Internet Protocol (TCP / IP), User Datagram Protocol (UDP), Hypertext Transfer Protocol (HTTP), File Transfer Protocol (FTP), ZigBee, EDGE, IEEE 802.11, Light Fidelity (Li-Fi), 802.16, IEEE 802.11s, IEEE 802.11g, Multihop Communication, Wireless Access Point (AP), Device-to-Device Communication, Cellular Communication Protocol, and Bluetooth (BT) Communication Protocol.
[0023] During operation, the system 102 can be configured to receive multiple images 114 of multiple attendees 116 related to multiple meeting sessions (such as meeting sessions attended by different attendees in the past). The multiple images 114 can be received from multiple image acquisition devices 104 positioned in different meeting sessions. The meeting categories of one or more of the multiple meeting sessions can be different and can correspond to at least one of the following: the type of meeting session, the number of attendees in the meeting session, the duration of the meeting session, the average age of the attendees in the meeting session, the topic of the meeting session, the experience of the educator in the meeting session, or the content presented in the meeting session. The system 102 can further detect one or more activities that each of the multiple attendees 116 can perform during the corresponding meeting session. The one or more activities can be detected based on the received multiple images 114. In one embodiment, one or more detected activities performed by each of the multiple attendees 116 may relate to at least one of the following: actions performed by the attendee, gestures performed by the attendee, head posture of the attendee, body position of the attendee, lip movements of the attendee, gaze of the attendee, or facial expressions of the attendee. In other words, one or more activities may indicate at least one of the attendee's level of attention or the level of interaction between the attendee and the educator. Details regarding the detection of one or more activities are shown, for example, in Figure 3.
[0024] In one embodiment, the system 102 can be configured to control an audio capture device 108 to capture interactions between multiple attendees 116 and an educator in a corresponding meeting session. In some embodiments, the system 102 can receive multiple audio signals (i.e., captured by the corresponding audio capture device during different meeting sessions) directly from the server 110. Based on the captured interactions, the system 102 can further determine one or more keywords within the captured interactions. Details regarding the determination of one or more keywords are shown, for example, in Figure 4.
[0025] System 102 can be further configured to calculate the attention score for each of a group of attendees over a corresponding period based on one or more detected activities that occur during different meeting sessions related to the corresponding attendees. In another embodiment, System 102 can be further configured to calculate the attention score for each of a group of attendees over a corresponding period based on one or more detected activities related to the corresponding attendees and the captured interactions between the corresponding attendees and the educator. The attention score can indicate the level of attention of each attendee in the corresponding meeting session. Details of the calculation of the attention level are shown, for example, in Figure 3. System 102 can be further configured to train each of a group of ML models 106 for each of the group of attendees 116 based on the calculated attention score for each of the group of attendees 116 and the meeting category of the corresponding meeting session.
[0026] System 102 can be configured to store multiple trained ML models 106 in memory (i.e., memory 204 in Figure 2) based on the training of multiple ML models 106, and to apply the multiple ML models 106 in a real-world scenario. In a real-world scenario, System 102 can be configured to receive a first set of images of a first attendee among multiple attendees 116. The first attendee may be associated with a first meeting session (i.e., a recent meeting session) which can be different from multiple meeting sessions (i.e., past meeting sessions on which the multiple ML models 106 are trained). System 102 can further detect a first set of activities that the first attendee can perform during the first meeting session over a first period (e.g., several minutes during the meeting session). Based on the detected first set of activities, System 102 can further calculate a first attention score associated with the first attendee over the first period. The first attention score may indicate the first attendee's level of attention during the first meeting session. System 102 can further apply a first machine learning (ML) model from among multiple ML models 106 based on the calculated first attention score. The first machine learning (ML) model can be personalized for the first attendee, or it can be trained based on historical data related to the first attendee (i.e., attention score, meeting category, recommendations, other factors such as attendee's facial expressions (or emotions), educator's experience, attendee's preferences / interests, meeting session environmental conditions, and attendee's response time). System 102 can further determine a first set of recommendations based on the application of the first ML model to the first attendee's calculated first attention score, and can further output the determined first set of recommendations for either the first attendee, the educator of the first meeting session, or the content presented during the first meeting session. Details regarding the application of the trained first ML model and the set of recommendations are shown, for example, in Figure 4.
[0027] Figure 2 is a block diagram illustrating an exemplary system for meeting session control based on attention decisions according to an embodiment of the present disclosure. The description of Figure 2 will be made in relation to the elements of Figure 1. Figure 2 shows a block diagram 200 of system 102. System 102 may include a circuit 202 that performs operations for training multiple ML models 106 and can further apply multiple trained ML models 106 for meeting session control based on attention decisions. System 102 may further include a memory 204, an input / output (I / O) device 206, a network interface 208, one or more NN models 210, and an inference accelerator 212. The memory 204 may include multiple ML models 106 and one or more neural network (NN) models 210. The circuit 202 may be communicatively coupled to the memory 204, the I / O device 206, the network interface 208, and the inference accelerator 212.
[0028] Circuit 202 may include preferred logic, circuits, and interfaces that can be configured to execute program instructions related to different operations performed by system 102. For example, part of an operation could include receiving multiple images 114, detecting one or more activities, calculating an attention score, training multiple ML models 106, applying the trained ML models, and making a recommendation. Circuit 202 may include one or more special processing units that can be implemented as independent processors. In some embodiments, one or more special processing units may be implemented as an integrated processor or group of processors that collectively execute the functions of one or more special processing units. Circuit 202 can be implemented based on several processor technologies well known in the art. Examples of implementations of circuit 202 may be x86-based processors, graphics processing units (GPUs), reduced instruction set computing (RISC) processors, application-specific integrated circuit (ASIC) processors, composite instruction set computing (CISC) processors, microcontrollers, central processing units (CPUs), and / or other control circuits.
[0029] Memory 204 may include suitable logic, circuitry, interfaces, and / or code that can be configured to store multiple received images 114, multiple trained ML models 106, and one or more NN models 210. Memory 204 can be further configured to store a first three-dimensional (3D) map of the meeting session, attendee focus scores, attendee interaction scores, facial expressions of each of the multiple attendees 116, educator experience information for the meeting session, attendee profile information, meeting session environment information, a first set of images, and dashboard information. Examples of implementations of memory 204 include, but are not limited to, random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), hard disk drives (HDDs), solid-state drives (SSDs), CPU caches, and / or secure digital (SD) cards.
[0030] The I / O device 206 may include suitable logic, circuitry, and interfaces that can be configured to receive user input and provide output based on the received user input. User input may include, but is not limited to, requests for providing recommendations for a particular meeting session, requests for providing dashboard information, experience information about a particular educator, or profile information about a particular attendee. The I / O device 206 may include various input and output devices that can be configured to communicate with circuitry 202. Examples of the I / O device 206 may include, but is not limited to, a display device 206A, an audio rendering device, a touch screen, a keyboard, a mouse, a joystick, and a microphone.
[0031] The display device 206A may include, but is not limited to, a suitable logic, circuitry, and interface that can be configured to display dashboard information including statistics on the attention score of a particular attendee (or on the meeting session), or recommendation information from attendees, educators, content creators, or organizations related to the meeting session. The display device 206A may be a touch screen that allows a user to provide user input through the display device 206A. The touch screen may be at least one of a resistive touch screen, a capacitive touch screen, or a thermal touch screen. The display device 206A may be implemented through a number of known technologies, such as liquid crystal display (LCD) displays, light-emitting diode (LED) displays, plasma displays, or organic LED (OLED) display technologies, or at least one of other display devices, but is not limited to. According to one embodiment, the display device 206A may mean a display screen for a head-mounted device (HMD), a smart glasses device, a see-through display, a projected display, an electrochromic display, or a transparent display.
[0032] The network interface 208 may include suitable logic, circuits, and interfaces that can be configured to facilitate communication between the circuit 202, the multiple image acquisition devices 104, the audio acquisition device 108, and the server 110 via the communication network 112. The network interface 208 may be implemented to support wired or wireless communication between the system 102 and the communication network 112 using various known techniques. The network interface 208 may include, but is not limited to, an antenna, a radio frequency (RF) transceiver, one or more amplifiers, a tuner, one or more oscillators, a digital signal processor, a coder-decoder (CODEC) chipset, a subscriber ID module (SIM) card, or a local buffer circuit. The network interface 208 may be configured to communicate wirelessly with networks such as the Internet, an intranet, or wireless networks such as a cellular telephone network, a wireless local area network (LAN), or a metropolitan area network (MAN). Wireless communication can be configured to use one or more of several communication standards, protocols, and technologies, including Global System for Mobile Communications (GSM), Enhanced Data GSM Environment (EDGE), Wideband Code Division Multiple Access (W-CDMA), Long-Term Evolution (LTE), Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Bluetooth, Wireless Fidelity (WiFi) (such as IEEE 802.11a, IEEE 802.11b, IEEE 802.11g, or IEEE 802.11n), Voice over Internet Protocol (VoIP), Light Fidelity (Li-Fi), Worldwide Interoperability for Microwave Access (Wi-MAX), Protocol for Email, Instant Message, and Short Message Service (SMS).
[0033] Each of the one or more NN models 210 can be a system of computer networks or artificial neurons arranged in multiple layers as nodes. Each of the multiple layers of the one or more NN models 210 can include an input layer, one or more hidden layers, and an output layer. Each of the multiple layers can include one or more nodes (or artificial neurons represented, for example, by circles). The outputs of all nodes in the input layer can be coupled to at least one node in the (one or multiple) hidden layers. Similarly, the input of each hidden layer can be coupled to the output of at least one node in each of the other layers of the one or more NN models 210. The output of each hidden layer can be coupled to the input of at least one node in each of the other layers of the one or more NN models 210. The (one or multiple) nodes in the final layer can receive input from at least one hidden layer and output a result. The number of layers and the number of nodes in each layer can be determined from the hyperparameters of each of the one or more NN models 210. Such hyperparameters can be set before, during, or after training each of the one or more NN models 210 based on the training dataset (e.g., training dataset 116). Each node of the one or more NN models 210 can correspond to a mathematical function (e.g., a sigmoid function or a rectified linear unit) with a set of parameters that can be adjusted during network training. The parameter set may include, for example, weight parameters and regularization parameters. Each node can compute an output using a mathematical function based on one or more inputs from the nodes of each of the other (single or multiple) layers (e.g., the previous (single or multiple) layer) of the one or more NN models 210. All or part of each node of the one or more NN models 210 can correspond to the same or different mathematical functions.
[0034] Each training of one or more NN models 210 may include updating one or more parameters of each node based on whether the output of the final layer for a given input (from the training dataset) matches the correct result based on the loss function of each of the one or more NN models 210. The above process can be repeated for the same or different inputs until the minimum value of the loss function is achieved and the training error is minimized. Several training methods are known in the industry, including gradient descent, stochastic gradient descent, batch gradient descent, gradient boosting, and metaheuristic methods.
[0035] Each of the one or more NN models 210 may include electronic data that can be implemented, for example, as a software component of an application executable on system 102. Each of the one or more NN models 210 may rely on libraries, external scripts, or other logic / instructions for execution by a device such as circuit 202. Each of the one or more NN models 210 may include code and routines configured to allow a computer device such as circuit 202 to perform one or more actions to detect one or more activities performed by multiple attendees. In addition to or instead of the above, each of the one or more NN models 210 may also be implemented using hardware including a processor, a microprocessor (for example, one or more actions to perform or control one or more actions), a field-programmable gate array (FPGA), or an application-specific integrated circuit (ASIC). Alternatively, in some embodiments, each of the one or more NN models 210 may also be implemented using a combination of hardware and software. Examples of one or more NN models 210 include, but are not limited to, deep neural networks (DNNs), convolutional neural networks (CNNs), artificial neural networks (ANNs), CNN+ANN, R-CNN, Fast R-CNN, Faster R-CNN, (You Only Look Once) YOLO networks, fully connected neural networks, and / or combinations of such networks.
[0036] The inference accelerator 212 may include suitable logic, circuitry, interfaces, and / or code that can be configured to operate as a coprocessor for circuit 202 to accelerate computations related to the operation of one or more NN models 210 and / or multiple ML models 106. For example, the inference accelerator 212 can accelerate the computations of system 102 so that one or more activities are detected in less time than would normally occur without the inference accelerator 212. The inference accelerator 212 can implement various acceleration techniques, such as parallelizing some or all of the operation of one or more NN models 210 and / or multiple ML models 106. The inference accelerator 212 can be implemented as software, hardware, or a combination thereof. Implementation examples of the inference accelerator 212 include, but are not limited to, GPUs, tensor processing units (TPUs), neuromorphic chips, vision processing units (VPUs), field-programmable gate arrays (FPGAs), reduced instruction set computing (RISC) processors, application-specific integrated circuit (ASIC) processors, composite instruction set computing (CISC) processors, microcontrollers, and / or combinations thereof.
[0037] Figure 3 shows an exemplary operation according to embodiments of the present disclosure for training an ML model for each of multiple attendees related to multiple meeting sessions. The description of Figure 3 is made in relation to the elements of Figures 1 and 2. Figure 3 shows a block diagram 300 illustrating exemplary operations 302A to 302J described herein. The exemplary operations shown in block diagram 300 can begin with 302A and can be performed by any computer system, apparatus or device, such as system 102 in Figure 1 or circuit 202 in Figure 2. The exemplary operations relating to one or more blocks in block diagram 300 are shown as discrete blocks, but these can be divided into further blocks, combined into fewer blocks, or deleted depending on the particular implementation.
[0038] In 302A, a data acquisition operation can be performed. In the data acquisition operation, the circuit 202 can be configured to receive multiple images 114 of multiple attendees 116 related to multiple meeting sessions. In one embodiment, the circuit 202 can control multiple image acquisition devices 104 over a period of time (e.g., several days, several weeks, several months, or several years) to capture multiple images 114 of multiple attendees 116. The circuit 202 can further receive multiple images 114 captured from the multiple image acquisition devices 104. In some embodiments, the circuit 202 can receive multiple images 114 of multiple attendees 116 (i.e., attendees who attended the corresponding meeting sessions) from the server 110.
[0039] In one embodiment, each meeting session (for example, a class as shown in Figure 1) may include a plurality of image acquisition devices 104 positioned at different locations within the meeting session to capture images of attendees from different viewpoints. In such a case, the plurality of image acquisition devices 104 may include, but are not limited to, a first image acquisition device, a second image acquisition device, and a third image acquisition device. The first image acquisition device may perform acquisition in the axial image plane (i.e., on the x and y axes), the second image acquisition device may perform acquisition in the orthogonal image plane (i.e., on the y and z axes), and the third image acquisition device may perform acquisition in the sagittal image plane (i.e., on the x and z axes). In one embodiment, the second and third image acquisition devices may be optional, and the first image acquisition device may capture images of attendees present within a particular meeting session. In one embodiment, the first image acquisition device may include a depth sensor (such as an RGBD camera) to capture images of attendees along with depth information (indicating the distance between the first image acquisition device and each attendee).
[0040] In one embodiment, the meeting session may be a class-based (i.e., offline-based) session, as shown in Figure 1, for example. Examples of such a meeting session include, but are not limited to, classes at educational institutions, physical meetings, conferences, seminars, or training rooms at professional institutions. In another embodiment, the meeting session may be an online-based session, such as an online class, online meeting session, online training session, web conference, or webinar, but are not limited to. In such a case, multiple image capture devices 104 (such as CCTVs) may be placed in a room (or any physical enclosure) from which attendees can participate in the corresponding online meeting session. In some cases, multiple image capture devices 104 (such as webcams) may be integrated (or built into or combined with) a computer device (such as a mobile phone, personal computer, or laptop) from which attendees can participate in the corresponding online meeting session.
[0041] In one embodiment, one or more meeting sessions among a group of meeting sessions may have different meeting categories. A meeting category may correspond to at least one of the following: the type of meeting session, the number of attendees, the duration of the meeting session, the average age of the attendees, the topic of the meeting session, the experience of the educator, or the content presented in the meeting session. In one embodiment, examples of different types of meeting sessions may include, but are not limited to, classroom-based sessions with an educator, classroom-based sessions without an educator, online-based sessions with an educator, and online-based sessions without an educator. The number of attendees may refer to the number of attendees present in a particular meeting session, such as a low-intensity meeting (i.e., 1 to 10 attendees or less than 30% meeting intensity), a medium-intensity meeting (i.e., 11 to 20 attendees or 31 to 60% meeting intensity), or a high-intensity meeting (i.e., 20 or more attendees or more than 60% meeting intensity). The duration of the meeting session indicates whether it is short-term (e.g., less than 30 minutes) or long-term (e.g., more than 1 hour). The average age of attendees may indicate the age group or educational year of the attendees (e.g., preschool attendees, primary standard attendees, higher standard attendees, first-year graduate attendees, final-year graduate attendees, 3-6 year old attendees, 12-17 year old attendees, 21-30 year old attendees, or 35 year old and over attendees). The topic of the meeting session may indicate the agenda of the meeting session (e.g., subject matter of the meeting session, a specific chapter of the curriculum, a technical or industry topic, a specific plan to be discussed during the meeting session, or a specific issue to be discussed during the meeting session, although these are not limited to the following). The experience of the educators may indicate the number of years of teaching (or training) experience of the educators (or teachers, trainers, or instructors), such as less than 1 year of experience, 2-5 years of experience, 6-10 years of experience, or 10 years or more of experience.The content presented in a meeting session can indicate whether it relates to theoretical content, includes multiple practical examples, or is interactive content.
[0042] In one embodiment, different meeting categories may influence or affect the attention levels of attendees, so the disclosed system 102 can consider images of attendees in different categories of meeting sessions. Accordingly, system 102 can track the attention scores of a particular attendee in different categories of meeting sessions over a specific period (e.g., several months or several years) to effectively train a personalized ML model using diverse and robust training data about that particular attendee.
[0043] In 302B, a 3D map generation operation can be performed. In the 3D map generation operation, circuit 202 can be configured to generate a first three-dimensional (3D) map of a corresponding meeting session, including at least one of a plurality of attendees 116. The generation of the first 3D map can correspond to profiling the corresponding (i.e., including at least one of the plurality of attendees 116) meeting session in three dimensions and mapping real-world objects or attendees. Circuit 202 can generate a first 3D map of a meeting session based on a plurality of captured images 114 of the meeting session including the corresponding attendees. In one embodiment, system 102 may include a video (or imaging) multiplexer 304 that can combine a plurality of images 114 of a particular meeting session captured by a first image acquisition sensor, a second image acquisition sensor, and a third image acquisition sensor. System 102 can further control the multiplexer 304 to generate a first 3D map of the meeting session. In one embodiment, the system 102 can use depth information (e.g., captured by a depth sensor of an image acquisition device) to generate a first 3D map of the meeting session, including the corresponding attendees.
[0044] In 302C, an activity detection operation can be performed. In the activity detection operation, the circuit 202 can be configured to detect one or more activities that each of several attendees 116 can perform during a corresponding meeting session. The circuit 202 can detect one or more activities of a particular attendee based on a set of received images 114 relating to the corresponding meeting session in which the attendee is present. In one embodiment, the circuit 202 can be configured to detect one or more activities of a corresponding attendee over a period of time (for example, over several minutes or several hours during a meeting session) by applying one or more neural network (NN) models 210 to the received images 114 or a first 3D map of the generated meeting session.
[0045] In one embodiment, one or more activities performed by each of the multiple attendees 116 may relate to at least one of the following: actions performed by the attendee, gestures performed by the attendee, head posture of the attendee, body position of the attendee, lip movements of the attendee, gaze of the attendee, or facial expressions of the attendee. Examples of one or more activities related to actions performed by the attendee include, but are not limited to, writing, typing, talking, asking questions, yawning, carrying objects, and participating in group activities. One or more activities related to gestures performed by the attendee may indicate whether the attendee is performing a predetermined gesture to perform an action, such as raising their hand to ask for permission to go outside for a break during the meeting session, or pointing at something / someone during a particular period of time. One or more activities related to lip movements of the attendee may indicate whether the attendee is reading aloud or talking to another attendee. One or more activities related to body position of the attendee may indicate whether the attendee is sitting or standing during that period of time, etc. One or more activities related to an attendee's gaze may indicate whether the attendee was focused during the meeting session. One or more activities related to an attendee's head posture may indicate whether the attendee was facing left, facing right, present, absent, or had their head down during the meeting session.
[0046] In one embodiment, the circuit 202 can be configured to detect one or more activities of attendees across a set of time slots within a period. The circuit 202 can detect at least one of the one or more activities (i.e., activities performed by multiple attendees 116) in each of the time slot sets within the period. For example, if the period is "10" minutes, each time slot set may be "1" minute or a few seconds. In one embodiment, each time slot set may correspond to a specific number of image frames in which a particular activity of an attendee is detected using a corresponding set of images 114. The circuit 202 can generate a timeline including the time slot sets for a particular meeting session and the detected activities of each of the multiple attendees 116 in each time slot, as shown in Table 1 below. TIFF0007859007000001.tif35155 Table 1: Exemplary timelines of detected activities for different attendees in corresponding meeting sessions. The parameters in Table 1 are as follows: AQ indicates that attendees are answering questions during the meeting session. F indicates that attendees are focused on the content of the meeting session. RH indicates that attendees are raising their hands during the meeting session. R indicates that an attendee is reading aloud during the meeting session. W indicates that attendees are writing or typing during the meeting session. T indicates that an attendee is talking to someone in the meeting session. E indicates that the attendee is not in their assigned seat at the meeting session (i.e., is not in their place). S indicates that attendees are standing during the meeting session. FM corresponds to frame not detected or activity not detected.
[0047] In one embodiment, the circuit 202 can determine the position of each attendee within a meeting session (for example, in a physically-class-based meeting session) based on a first 3D map of the generated meeting session. The circuit 202 can associate or tag specific seats with attendee positions based on the first 3D map of the generated meeting session. Tagging attendees with specific seats can help identify attendees during the remainder of the meeting session and identify activities performed by the same attendee while sitting in the same seat to which they were tagged. Similarly, the circuit 202 can determine the position of each of several attendees 116 in a corresponding meeting session based on several received images 114 and the determined 3D map of each meeting session. Furthermore, the circuit 202 can associate each of the several attendees 116 with a different seat in each meeting session. Based on such seat tagging, the circuit 202 can determine where in the meeting session (front, back, middle, or corner) a particular attendee was sitting and performing different activities while attending each meeting session. Therefore, based on seat tagging performed by the disclosed system 102, different behaviors (or activities or interactions) can be associated with corresponding attendees in different meeting sessions.
[0048] In 302D, a focus score calculation operation can be performed. In the focus score calculation operation, circuit 202 can be configured to calculate the focus score for each of multiple attendees 116 in corresponding meeting sessions in which each attendee is present. Circuit 202 can calculate the focus score for each of the multiple attendees 116 based on one or more detected activities associated with the corresponding attendee. In other words, circuit 202 can calculate the focus score for a specific attendee in a particular meeting session. Specifically, the focus score can be calculated based on the detection of a first set of one or more activities. Such a first set of activities may include, but is not limited to, actions performed by the attendee (such as writing, typing, talking, or yawning), the attendee's gaze, the attendee's head posture, the attendee's body position, and the attendee's lip movements (such as reading aloud or conversing with other attendees).
[0049] In one embodiment, circuit 202 can be configured to determine a first duration for at least one activity among one or more activities performed by each of several attendees during a corresponding meeting session. Specifically, circuit 202 can be configured to determine a first duration for each of a first set of activities. Circuit 202 can be configured to assign a score to each of the first set of activities based on the determined first duration. In one embodiment, circuit 202 can be configured to assign scores or points (such as, but not limited to, writing-related points, reading-related points, gaze-related points, or speaking-related points) for different activities to each of several attendees 116 based on several captured images 114 of a corresponding meeting session. In one embodiment, the determined first duration may indicate the time attendees spent on different activities, which can be further used to tag attendees as "attention seekers" and / or "attentive attendees."
[0050] Circuit 202 can be configured to determine the intensity (i.e., the number of attendees) of each of multiple meeting sessions in order to assign reading-related points. Circuit 202 can further compare the determined intensity of the meeting sessions with a threshold intensity (e.g., 30% intensity). If the determined intensity is higher than the threshold intensity, system 102 can apply a reading formula to assign reading-related points to each of the multiple attendees 116 that may be present in the meeting session. The reading-related formula can be applied by the following equation (1): TIFF0007859007000002.tif9150(1) Here, R corresponds to reading-related points or score. C corresponds to the type of meeting session, k corresponds to the duration of the reading activity in seconds / minutes (i.e., the first duration), S corresponds to the intensity of the meeting session. C1 supports webinar meeting sessions with educators. C2 supports webinar meeting sessions without an educator present. C3 supports classroom meeting sessions with educators. C4 supports classroom meeting sessions without an instructor present. C5 is designed for interaction / activity meeting sessions with educators. C6 is designed for interaction / activity meeting sessions without an educator. Cn = 0 or 1. If C1 = 1, then C2 = C3 = C4 = C5 = C6 = 0. P1, P2, P3, P4, P5, and P6 correspond to weight variables with values between 0 and 1.
[0051] In one embodiment, if the percentage (i.e., intensity) of attendees (A) present in a meeting session is greater than 50% and the duration of attendees' reading activities is between 10 and 40 seconds, each of the multiple attendees 116 who are reading in a particular meeting session may be assigned "0.9" reading-related points, while the remaining attendees (i.e., those not reading) may be assigned "0.1" reading-related points. In another embodiment, if the percentage (i.e., intensity) of attendees present in a meeting session is greater than 50% and the duration of attendees' reading activities is greater than 40 seconds, each of the multiple attendees 116 who are reading may be assigned "1" reading-related point, while the remaining attendees (i.e., those not reading) may be assigned "0" reading-related points. In another embodiment, if the percentage of attendees (A) who are present / attending the meeting session is greater than 50% and the duration of an attendee's reading activity is less than 10 seconds, each of the multiple attendees 116 who are reading can be assigned 0 reading-related points, while the remaining attendees (i.e., those who were not reading) can also be assigned 0 reading-related points.
[0052] In one embodiment, if the percentage of attendees (A) who are present / in attendance at the meeting session is less than 25% and the duration of an attendee's reading activity is longer than 40 seconds, each of the multiple attendees 116 who are reading can be assigned "0." Reading-related points, while the remaining attendees (i.e., those who were not reading) can also be assigned "0." Reading-related points. In another embodiment, if the percentage of attendees (A) who are present / in attendance at the meeting session is less than 25% and the duration of an attendee's reading activity is less than 10 seconds, each of the multiple attendees 116 who are reading can be assigned "0." Reading-related points, while the remaining attendees (i.e., those who were not reading) can also be assigned "0." In one embodiment, if the percentage of attendees (A) who are present / attending the meeting session is less than 25%, and the duration of an attendee's reading activity is between 10 seconds and 40 seconds, each of the multiple attendees 116 who are reading can be assigned "0" reading-related points, while the remaining attendees (i.e., those who were not reading) can also be assigned 0 reading-related points.
[0053] In one embodiment, if the percentage of attendees (A) present at the meeting session is greater than 25% but less than 50%, and the duration of an attendee's reading activity is between 10 seconds and 40 seconds, each of the multiple attendees 116 who are reading can be assigned 0.8 reading-related points, while the remaining attendees (i.e., those not reading) can be assigned 0.2 reading-related points. In another embodiment, if the percentage of attendees (A) present at the meeting session is greater than 25% but less than 50%, and the duration of an attendee's reading activity is longer than 40 seconds, each of the multiple attendees 116 who are reading can be assigned 0.6 reading-related points, while the remaining attendees (i.e., those not reading) can be assigned 0.4 reading-related points. In another embodiment, if the percentage of attendees present / in attendance at the meeting session (A) is greater than 25% but less than 50%, and the duration of each attendee's reading activity is less than 10 seconds, then each of the multiple attendees 116 who are reading can be assigned 0 reading-related points, while the remaining attendees (i.e., those who were not reading) can also be assigned 0 reading-related points. Table 2 below shows an example relationship between the intensity of the meeting session (i.e., the number of attendees (A)), the duration of the reading activity, and the reading-related points. TIFF0007859007000003.tif67155 Table 2: Exemplary relationships between meeting session intensity, reading activity duration, and reading-related points.
[0054] In one embodiment, the circuit 202 may be configured to determine the intensity of each meeting session in a group of meeting sessions in order to assign writing-related points. The circuit 202 can further compare the determined intensity of the meeting session with a threshold intensity (e.g., 30% intensity). If the determined intensity is higher than the threshold intensity, the system 102 can apply a write formula to assign writing-related points to each of the group of attendees 116 that may be present in the meeting session. The write formula can be applied by the following equation (2): TIFF0007859007000004.tif11157(2) Here, W corresponds to points or scores related to writing, C corresponds to the type of meeting session, k corresponds to the duration of the writing activity in seconds / minutes (i.e., the first duration), S corresponds to the intensity of the meeting session. C1 supports webinar meeting sessions with educators. C2 supports webinar meeting sessions without an educator present. C3 supports classroom meeting sessions with educators. C4 supports classroom meeting sessions without an instructor present. C5 is designed for interaction / activity meeting sessions with educators. C6 is designed for interaction / activity meeting sessions without an educator. Cn = 0 or 1. If C1 = 1, then C2 = C3 = C4 = C5 = C6 = 0. P1, P2, P3, P4, P5, and P6 correspond to weight variables with values between 0 and 1.
[0055] In one embodiment, if the percentage of attendees (A) who are present / in attendance at the meeting session is higher than 50%, and the duration of each attendee's writing activity is between 3 seconds and 15 seconds, each of the multiple attendees 116 who are writing may be assigned "0.9" writing-related points, while the remaining attendees (i.e., those who were not writing) may be assigned "0.1" writing-related points. In another embodiment, if the percentage of attendees (A) who are present / in attendance at the meeting session is higher than 50%, and the duration of each attendee's writing activity is longer than 15 seconds, each of the multiple attendees 116 who are writing may be assigned "1.0" writing-related points, while the remaining attendees (i.e., those who were not writing) may be assigned "0" writing-related points. In another embodiment, if the percentage of attendees (A) who are present / attending the meeting session is greater than 50% and the duration of an attendee's writing activity is less than 3 seconds, each of the multiple attendees 116 who are writing can be assigned 0 writing-related points, while the remaining attendees (i.e., those who were not writing) can also be assigned 0 writing-related points.
[0056] In one embodiment, if the percentage of attendees (A) present at the meeting session is less than 25% and the duration of each attendee's writing activity is longer than 15 seconds, each of the multiple attendees 116 who are writing can be assigned 0.5 writing-related points, while the remaining attendees (i.e., those who were not writing) can also be assigned 0.5 writing-related points. In another embodiment, if the percentage of attendees (A) present at the meeting session is less than 25% and the duration of each attendee's writing activity is less than 3 seconds, each of the multiple attendees 116 who are writing can be assigned 0 writing-related points, while the remaining attendees (i.e., those who were not writing) can also be assigned 0 writing-related points. In one embodiment, if the percentage of attendees (A) who are present / attending the meeting session is less than 25%, and the duration of each attendee's writing activity is between 3 and 15 seconds, then each of the multiple attendees 116 who are writing can be assigned 0 writing-related points, while the remaining attendees (i.e., those who were not writing) can also be assigned 0 writing-related points.
[0057] In one embodiment, if the percentage of attendees (A) who are present / in attendance at the meeting session is greater than 25% but less than 50%, and the duration of each attendee's writing activity is between 3 seconds and 15 seconds, then each of the multiple attendees 116 who are writing can be assigned 0.8 writing-related points, while the remaining attendees (i.e., those who were not writing) can be assigned 0.2 writing-related points. In another embodiment, if the percentage of attendees (A) who are present / in attendance at the meeting session is greater than 25% but less than 50%, and the duration of each attendee's writing activity is longer than 15 seconds, then each of the multiple attendees 116 who are writing can be assigned 0.6 writing-related points, while the remaining attendees (i.e., those who were not writing) can be assigned 0.4 writing-related points. In another embodiment, if the percentage of attendees present / in attendance at the meeting session (A) is greater than 25% but less than 50%, and the duration of each attendee's writing activity is less than 3 seconds, then each of the multiple attendees 116 who are writing can be assigned 0 writing-related points, while the remaining attendees (i.e., those who were not writing) can also be assigned 0 writing-related points. Table 3 below shows an exemplary relationship between the intensity of the meeting session (i.e., the number of attendees (A)), the duration of writing activity, and writing-related points. TIFF0007859007000005.tif67155 Table 3: Exemplary relationships between meeting session intensity, duration of writing activity, and writing-related points.
[0058] In one embodiment, the circuit 202 can be configured to determine the intensity of each meeting session in a group of meeting sessions in order to assign gaze-related points. The circuit 202 can further compare the determined intensity of the meeting session with a threshold intensity (e.g., 30% intensity). If the determined intensity is higher than the threshold intensity, the system 102 can apply a gaze formula to assign gaze-related points to each of the multiple attendees 116 that may be present in the meeting session. The gaze formula can be applied by the following equation (3): TIFF0007859007000006.tif11157(3) Here, X corresponds to the attention-related points or score, C corresponds to the type of meeting session, k corresponds to the duration of the gaze activity in seconds / minutes (i.e., the first duration), S corresponds to the intensity of the meeting session. C1 supports webinar meeting sessions with educators. C2 supports webinar meeting sessions without an educator present. C3 supports classroom meeting sessions with educators. C4 supports classroom meeting sessions without an instructor present. C5 is designed for interaction / activity meeting sessions with educators. C6 is designed for interaction / activity meeting sessions without an educator. Cn = 0 or 1. If C1 = 1, then C2 = C3 = C4 = C5 = C6 = 0. P1, P2, P3, P4, P5, and P6 correspond to weight variables with values between 0 and 1.
[0059] In one embodiment, if the percentage of attendees (A) who are present / in attendance at the meeting session is higher than 50%, and the duration of an attendee's gaze activity is between 30 and 90 seconds, each of the multiple attendees 116 who are gazing can be assigned 0.9 gaze-related points, while the remaining attendees (i.e., those not gazing) can be assigned 0.1 gaze-related points. In another embodiment, if the percentage of attendees (A) who are present / in attendance at the meeting session is higher than 50%, and the duration of an attendee's gaze activity is longer than 90 seconds, each of the multiple attendees 116 who are gazing can be assigned 1.0 gaze-related points, while the remaining attendees (i.e., those not gazing) can be assigned 0 gaze-related points. In another embodiment, if the percentage of attendees (A) who are present / attending the meeting session is greater than 50% and the duration of an attendee's gaze activity is less than 30 seconds, each of the multiple attendees 116 who are gazing can be assigned 0 gaze-related points, while the remaining attendees (i.e., those who were not gazing) can also be assigned 0 gaze-related points.
[0060] In one embodiment, if the percentage of attendees (A) who are present / in attendance at the meeting session is less than 25% and the duration of an attendee's gaze activity is longer than 90 seconds, each of the multiple attendees 116 who are gazing can be assigned 0.5 gaze-related points, while the remaining attendees (i.e., those who were not gazing) can also be assigned 0 gaze-related points. In another embodiment, if the percentage of attendees (A) who are present / in attendance at the meeting session is less than 25% and the duration of an attendee's gaze activity is less than 30 seconds, each of the multiple attendees 116 who are gazing can be assigned 0 gaze-related points, while the remaining attendees (i.e., those who were not gazing) can also be assigned 0 gaze-related points. In one embodiment, if the percentage of attendees (A) who are present / attending the meeting session is less than 25%, and the duration of an attendee's gaze activity is 30 seconds or more but less than 90 seconds, each of the multiple attendees 116 who are gazing can be assigned 0 gaze-related points, while the remaining attendees (i.e., those who were not gazing) can also be assigned 0 gaze-related points.
[0061] In one embodiment, if the percentage of attendees (A) who are present / in attendance at the meeting session is greater than 25% but less than 50%, and the duration of an attendee's gaze activity is 30 seconds or more but less than 90 seconds, then each of the multiple attendees 116 who are gazing can be assigned 0 gaze-related points, while the remaining attendees (i.e., those who were not gazing) can also be assigned 0 gaze-related points. In another embodiment, if the percentage of attendees (A) who are present / in attendance at the meeting session is greater than 25% but less than 50%, and the duration of an attendee's gaze activity is longer than 90 seconds, then each of the multiple attendees 116 who are gazing can be assigned 0.5 gaze-related points, while the remaining attendees (i.e., those who were not gazing) can also be assigned 0.5 gaze-related points. In another embodiment, if the percentage of attendees present / in attendance at the meeting session (A) is greater than 25% but less than 50%, and the duration of an attendee's gaze activity is less than 30 seconds, then each of the multiple attendees 116 who are gazing can be assigned 0 gaze-related points, while the remaining attendees (i.e., those not gazing) can also be assigned 0 gaze-related points. Table 4 below shows an exemplary relationship between the intensity of the meeting session (i.e., the number of attendees (A)), the duration of gaze activity, and gaze-related points. TIFF0007859007000007.tif67155 Table 4: Exemplary relationships between meeting session intensity, duration of attention activity, and attention-related points.
[0062] In one embodiment, the circuit 202 can be configured to calculate an attention score based on reading-related points, writing-related points, and gaze-related points. Specifically, the attention score for each of several attendees 116 can be calculated based on the sum of reading-related points, writing-related points, and gaze-related points assigned to the corresponding attendees in the meeting session. The circuit 202 can further store the calculated attention scores of specific attendees in a particular meeting session in memory 204.
[0063] In 302E, an interaction score calculation operation can be performed. In the interaction score calculation operation, circuit 202 can be configured to calculate the interaction score for each of several attendees 116 for activities performed in a corresponding meeting session. Circuit 202 can calculate the interaction score for each of several attendees 116 based on one or more detected activities performed by the corresponding attendees in different meeting sessions. In other words, circuit 202 can calculate the interaction score for a specific attendee in a particular meeting session. Specifically, the interaction score can be calculated based on the detection of a second set of one or more activities. Such a second set of activities may include, but is not limited to, gestures performed by the attendee (such as raising a hand or participating in an activity), gestures performed by the attendee (such as sitting down or standing up), and lip movements of the attendee (such as when answering a question).
[0064] For example, if the percentage of attendees present at a meeting session (A) is higher than 30%, and attendees raise their hands and speak, those attendees who raised their hands and spoke can be assigned 0.5 interaction points, while the remaining attendees can be assigned 0.1 interaction points. If attendees only speak and do not raise their hands or stand, those attendees can be assigned 0.05 interaction points, while the remaining attendees can be assigned 0.2 interaction points. Each attendee's interaction score (i.e., a score called "RA") can be further based on the assigned interaction points. Table 5 below shows an exemplary relationship between the intensity of the meeting session (i.e., the number of attendees (A)), the activities performed by attendees, and interaction points. TIFF0007859007000008.tif56155 Table 5: Exemplary relationships between meeting session intensity, activities performed by attendees, and interaction points.
[0065] Circuit 202 can be further configured to calculate the median of multiple interaction scores / points calculated for all attendees of a particular meeting session (e.g., a particular class). M The final interaction score for each attendee can be calculated by adding a value called (RA) to each attendee's interaction score / points. This calculation can be performed to normalize the statistical attention curve (i.e., the curve formed by the interaction scores calculated for all attendees) to reduce any bias in the meeting session resulting from the absence of attendees who are negatively affected because there are few high-performing attendees / individuals in the meeting session. In one embodiment, the calculation of the final interaction score (i.e., the sum of a particular attendee's interaction score and the calculated median) can be based on the intensity of attendees in a particular meeting session. For example, if the current intensity of the meeting session is less than 30% of the total intensity, circuit 202 may not consider the calculated median for the calculation of the final interaction score. In such a case, the final interaction score can be the same as the calculated interaction score for each attendee. In another example, if the current intensity of the meeting session is 30% or more of the total intensity, circuit 202 may consider the median based on Table 6 below. M It can be determined as a value. TIFF0007859007000009.tif34153 Table 6: Location of interaction score relative to median and RA M An illustrative relationship between values.
[0066] In 302F, an information determination operation can be performed. In the information determination operation, the circuit 202 can be configured to determine information related to multiple attendees 116, information related to multiple educators in multiple meeting sessions, and information related to multiple meeting sessions. In one embodiment, the circuit 202 can determine experience information that may be related to each of multiple educators in a corresponding meeting session among the multiple meeting sessions. The experience information may represent at least one of the following related to each educator in the corresponding meeting session: experience, evaluation, achievements, or feedback. The circuit 202 can be further configured to store the determined experience information in memory 204. In some embodiments, the circuit 202 can determine information about a particular educator by retrieving experience information about that educator from the server 110 or memory. The educator's experience information may demonstrate the educator's ability to effectively present a particular topic or content and engage attendees in the best possible way.
[0067] Circuit 202 can be further configured to determine content information related to the content presented during each of multiple meeting sessions. The determined content information may include, but is not limited to, the type of content (theoretical, interactive, content with practical / logical / real-time examples and exercises, whether it is fair to new learners), the duration of the content, the subject matter related to the content, the complexity of the content, and the interactivity of the content. Circuit 202 can store the determined content information in memory 204. In some embodiments, circuit 202 can retrieve content information about specific content from server 110 or memory to determine content information related to the content presented during a particular meeting session.
[0068] Circuit 202 can be further configured to retrieve profile information related to each of a plurality of attendees 116. Profile information may indicate preferences or interests for topics or content related to the corresponding meeting session. In some embodiments, profile information may further indicate health information related to the corresponding attendee during a particular meeting session. For example, if an attendee is a student and prefers subjects such as mathematics and science but dislikes social studies, profile information may indicate that this student may be interested in mathematics and science but not in social studies. Thus, the student's level of attention while attending a mathematics and science class may be higher than the student's level of attention while attending a social studies class. Similarly, based on health information, system 102 can determine whether an attendee had any particular health problem (such as a cold, headache, or fever) during the meeting session that resulted in low attention during the meeting session. In some embodiments, circuit 202 can determine profile information by retrieving profile information about a particular attendee from server 110 or memory.
[0069] Circuit 202 may be further configured to determine environmental information related to the geolocation of the educator for each meeting session or at least one of the multiple attendees 116 for each meeting session. The environmental information may include weather information related to the geolocation of the educator for each meeting session or at least one of the multiple attendees 116 for each meeting session. Weather conditions at the geolocation (e.g., heavy rain / wind / storm, or excessive heat / cold) may also affect the attention levels of the multiple attendees 116. The environmental information may further include indoor / outdoor lighting conditions, the teacher's hearing ability, lighting parameters related to the equipment on which the content is presented, etc. Similarly, poor lighting conditions or acoustic problems during the meeting session may also affect the attention levels of specific attendees. In some embodiments, circuit 202 may receive environmental information about the meeting session from one or more sensors (such as temperature sensors, wind sensors, rain sensors, lighting sensors, or audio sensors, but not limited to those below) placed at different locations in the meeting session. Circuit 202 may be further configured to store the determined weather information in memory 204. In some embodiments, the circuit 202 can retrieve environmental information related to a specific meeting session from the server 110 or memory to determine the environmental information.
[0070] In 302G, an expression determination operation can be performed. In the expression determination operation, the circuit 202 can be configured to determine the expression (or emotion) of each of the multiple attendees 116 in the corresponding meeting session based on the multiple images 114 received. In one embodiment, the system 102 can be configured to determine the expression of each of the multiple attendees 116 in the corresponding meeting session by applying one or more NN models 210 to each of the multiple images 114. Based on the determined expression, the system 102 determines the expression parameter (E) for each of the multiple attendees 116. N Circuit 202 can be further configured to initialize the facial expression parameter (E NTo initialize the first facial expression (E) with a value, it can be further configured to determine a first facial expression that can be determined for the majority of attendees from a group of attendees 116 in a particular meeting session. For example, if the total number of attendees 116 in a meeting session is "10" and the determined facial expression of "7" attendees is "smiling", then the first facial expression can be "smiling". Circuit 202 then determines the facial expression parameter (E) of the attendees in the meeting session if the determined facial expression of the corresponding attendee is the same as the first facial expression. N ) can be configured to be initialized to the value "1". Otherwise, the facial expression parameter (E N The ) can be initialized to the value "0". Circuit 202 can initialize the facial expression parameters (E) of multiple attendees 116. N The initialization value of ) can be configured to be stored in memory 204. Examples of facial expressions include, but are not limited to, happiness, sadness, emotion, surprise, anger, neutrality, astonishment, shock, stress, boredom, calmness, excitement, confusion, disgust, or fear.
[0071] In 302H, an object detection operation can be performed. In the object detection operation, the circuit 202 can be configured to detect one or more objects associated with multiple attendees in multiple received images 114 for different meeting sessions in which the corresponding attendees are present. In another embodiment, the circuit 202 can be configured to apply one or more neural network (NN) models to the multiple received images 114 to detect one or more objects associated with multiple attendees 116 in the multiple received images 114. Such objects may be animated objects that can be held by or placed near multiple attendees 116 of the corresponding meeting session. Examples of such objects include, but are not limited to, pens, paper, notebooks, water bottles, toys, ornaments, electrical devices, and communication devices. The circuit 202 can be configured to store information about the detected one or more objects associated with each of the multiple attendees 116 in the memory 204. Furthermore, since certain objects may distract a particular attendee and lower their level of attention (for example, toys may distract a young child, ornaments / paintings may distract an attendee, and a personal cell phone may distract an attendee), system 102 may also consider objects held by the attendee or objects near the attendee.
[0072] In 302I, an attention score calculation operation can be performed. In the attention score calculation operation, circuit 202 can be configured to calculate an attention score for each of a plurality of attendees 116 for each meeting session (i.e., based on the activities detected for each attendee over the corresponding period). The attention score can indicate the level of attention of each attendee in the corresponding meeting session. Circuit 202 can be further configured to calculate an attention score for each of the plurality of attendees 116 based on the calculated concentration score and the calculated interaction score of the corresponding attendee for each activity detected in different meeting sessions. As an example, the calculation of the attention score based on the calculated concentration score and the calculated interaction score can be performed by the following equation (4), AS = FS * (1 - Z)+IS * Z (4) where, AS corresponds to the attention score, FS corresponds to the concentration score, IS corresponds to the interaction score, Z is a multiplication coefficient that can depend on the type of meeting session, and 0 ≤ Z ≤ 1.
[0073] In certain embodiments, as indicated by the time slot tagged as 'E' in the timeline (i.e., the timeline shown in Table 1), an attendee may not be present (i.e., out of place) at the assigned / tagged seat in the meeting session. In such a scenario, the calculation of the attention score can be performed by the following equation (5), AS = FS * (1 - Z)+IS * Z - E (5) where, AS corresponds to the attention score, FS corresponds to the concentration score, IS corresponds to the interaction score, Z corresponds to the multiplication coefficient, and 0 ≤ Z ≤ 1. E corresponds to the out-of-place score, which indicates the possibility that an attendee may not be in their assigned seat (i.e., out of place) during the meeting session in which they are participating.
[0074] In one embodiment, circuit 202 may be configured to calculate an out-of-place score for a particular attendee based on the intensity of the meeting session and the duration of time that particular attendee is outside their designated (or identified or tagged) seat. Circuit 202 may be further configured to calculate a final attention score according to equation (5) based on the focus score, interaction score and out-of-place score. For example, if the current intensity of the meeting session is less than 30% of the total intensity, circuit 202 may not consider the out-of-place score in order to calculate the final interaction score. In such a case, the final interaction score may be the same as the interaction score calculated for each attendee according to equation (4). In another example, if the current intensity of the meeting session is 30% or more of the total intensity, the out-of-place score ("E") may be calculated based on the intensity of the meeting session and the duration of time that the attendee is out of place, as shown in Table 7 below. TIFF0007859007000010.tif51155 Table 7: Illustrative relationship between out-of-place score, number of attendees (intensity), and duration of out-of-place status.
[0075] In one embodiment, the circuit 202 can calculate an interaction score for each attendee in each meeting session in which each attendee is present. The attendee's interaction score can be calculated based on the activities detected over a defined period in each meeting session (i.e., the period described in, for example, 302C). Thus, the disclosed system 102 can calculate multiple attention scores for each attendee in different meeting sessions in which each attendee is present, and one or more meeting sessions of an attendee may belong to different meeting categories (i.e., meeting categories described in, for example, 302A). Multiple attention scores can indicate appropriate patterns or variations in a particular candidate's attention across different categories of meeting sessions (i.e., different types, different intensities, different durations, different ages, different topics / content, different educator experience).
[0076] In one embodiment, the circuit 202 can be configured to calculate the attention score for each of the multiple attendees 116 based on the calculated concentration score, the calculated interaction score of the corresponding attendee, and the stored experience information related to the educator for each meeting session. In this case, the calculation of the attention score can be based on the calculated concentration score, the calculated interaction score, and the determined experience information, and can be performed using the following equation (6): AS=(FS * (1-Z)+IS * ZE) / (R T +R L ) (6) Here, AS corresponds to the attention score, FS corresponds to concentrated scoring, IS corresponds to the interaction score, Z is the multiplication coefficient, and 0 ≤ Z ≤ 1. E corresponds to the time when attendees may not be present in their assigned seats. R TThis corresponds to the determined experience information, R L This corresponds to the evaluation provided by attendees regarding the content presented in the corresponding meeting session.
[0077] For example, attendees are likely to pay more attention to highly-rated educators and / or content. Therefore, an attempt can be made to equalize the attention scores of less experienced educators (teachers) or educators teaching difficult subjects by dividing low attention scores by low educator / content ratings, or by high attention scores by high educator / content ratings.
[0078] In another embodiment, the circuit 202 is used to calculate the concentration score, the calculated interaction score, and the facial expression parameter (E) described in, for example, 302G. N The attention score can be configured to be calculated based on the initialization value of ). Such an attention score can be calculated using the following equation (7): AS / FS * (1-Z)+IS * Z+E N (7) Here, AS corresponds to the attention score, FS corresponds to concentrated scoring, IS corresponds to the interaction score, Z is the multiplication coefficient, and 0 ≤ Z ≤ 1. E N This corresponds to the facial expression parameter, E N = 0 or 1
[0079] In another embodiment, the circuit 202 may be configured to calculate an attention score based on profile information associated with each of the multiple attendees 116, determined environmental information (as described in, for example, 302F), a calculated concentration score, and a calculated interaction score for the corresponding attendee. In yet another embodiment, the circuit 202 may be configured to calculate an attention score associated with each of the multiple attendees 116 based on a calculated concentration score, a calculated interaction score, and one or more detected objects (as described in, for example, 302H) associated with the multiple attendees 116. Thus, in addition to considering attendee concentration and interaction, the disclosed system 102 may consider different influencing factors that may affect the attention score calculated for attendees for different activities performed during each meeting session (such as, but not limited to, educator experience information, attendee profile information, content information, environmental information, and information about objects present in the meeting session). In one embodiment, the disclosed system 102 can fine-tune the attention score (i.e., the attention score calculated according to equation (4) based on the concentration score and the interaction score) using different influencing factors for the attendees' attention (i.e., using equations (5), (6), and (7)).
[0080] In 302J, an ML model training operation can be performed. In the ML model training operation, circuit 202 can be configured to train multiple ML models 106 for multiple attendees 116 based on the calculated attention scores of each of the multiple attendees 116 and the meeting category of the corresponding meeting session. Specifically, circuit 202 can be configured to train multiple ML models 106 for multiple attendees 116. Each of the multiple ML models 106 can be personalized for each attendee. Each personalized ML model can be trained (or memorize attention scores) based on the corresponding attendee's various attention scores determined from different meeting sessions (of various categories). Furthermore, the personalized ML model can also be trained, with respect to each attention score, based on corresponding influencing factors (i.e., educator experience, content, attendee profile, environmental factors, or nearby objects) identified from different meeting sessions previously attended by the corresponding attendee.
[0081] In one embodiment, a personal ML model can be trained with respect to different seating positions (i.e., front, back, middle, corner, window, or door) in which a particular attendee may be seated during a corresponding (at least physically classroom-based) meeting session, or the personal ML model can recognize such different seating positions. The attendee's position can also influence the corresponding attendee's attention level / score. For example, a first attendee sitting near the educator is likely to be more attentive than a second attendee sitting further away from the educator. Therefore, the first attendee's attention score may be higher than that of the second attendee. The circuit 202 can be further configured to train multiple ML models 106 for each of the multiple attendees 116 based on the calculated attention scores and the determined positions of each of the multiple attendees 504 in the corresponding meeting session. As illustrated, for example, in Figure 3(302C), the association of attendees with different positions (or seats) in a meeting session can be called seat tagging.
[0082] Therefore, a trained ML model for a specific attendee can recognize different patterns or fluctuations in the attendee's attention, such as under what circumstances or in what categories of meeting sessions the attendee's attention score increases or decreases. Furthermore, a personalized ML model can be trained on such cumulative data that shows (at least in the form of attention) the learning journey of a specific attendee (such as a student) with respect to factors such as subject / topic, class / standard (age group), meeting session type, meeting session location, seat tagging, information about the educator, environmental conditions, or attendee profile. Each trained model can also show the behavioral characteristics of each attendee (i.e., behavior, gestures, and interactions) by training on the attention scores of past meeting sessions attended by that attendee. Furthermore, multiple ML models 106 created by the disclosed system 102 can function as a ranking and rating library of educational or learning systems that can provide different recommendations for attendees / educators / content by showing real-time performance of attendees / educators in order to develop a robust and holistic educational or learning system. Multiple ML models 106 can be further stored in memory 204 so that they can be applied in real-time scenarios as described in Figure 4.
[0083] Figure 4 shows an exemplary operation for meeting session control based on attention determination and application of an ML model for a specific attendee trained in Figure 3, according to an embodiment of the present disclosure. The description of Figure 4 is made in relation to the elements of Figures 1, 2 and 3. Figure 4 shows a block diagram 400 illustrating exemplary operations 402A to 402F described herein. The exemplary operations shown in block diagram 400 can begin with 402A and can be performed by any computer system, apparatus or device, such as system 102 in Figure 1 or circuit 202 in Figure 2. The exemplary operations relating to one or more blocks in block diagram 400 are shown as discrete blocks, but these can be divided into further blocks, combined into fewer blocks, or deleted depending on the particular implementation.
[0084] In 402A, a data acquisition operation can be performed. In the data acquisition operation, circuit 202 can be configured to receive a first image set 404 of a first attendee 406 among a plurality of attendees 116. The first attendee 406 may be associated with (or attending or have attended) a first meeting session. Specifically, the first attendee 406 may be attending a first meeting session that can be different from a plurality of meeting sessions (i.e., a first ML model 402 can be trained for the first attendee based on this meeting session, for example, as illustrated in Figure 3). As described above, multiple ML models 106 for a plurality of attendees 116 may be trained based on data associated with a plurality of meeting sessions (i.e., attention scores, or activities associated with behavior / gestures / interactions), and may not be trained based on data associated with a first meeting session. In one embodiment, the first image set 404 may be of a plurality of attendees (including the first attendee) who have attended (or are currently attending) a first meeting session. The circuit 202 of system 102 can receive the first image set 404 from server 110, or directly from one or more image acquisition devices located at different locations within the first meeting session. If the first meeting session is an online meeting session, one or more image acquisition devices can be integrated into electronic devices (such as laptops, computers, or mobile phones) used by attendees to participate in the first meeting session.
[0085] In another embodiment, the circuit 202 may be further configured to control an audio capture device 108 (i.e., an audio capture device placed near each participant in the first meeting session) to capture the interaction 408 between a first participant 406 and the educator of the first meeting session. In one embodiment, the captured interaction 408 may relate to content presented in the first meeting session, or questions raised by the educator (or the first participant 406), or responses made by the first participant 406 (or the educator). For example, the first participant may ask one or more questions related to the content, and the educator may answer the first participant 406 about one or more questions asked, or vice versa.
[0086] In another embodiment, circuit 202 may be configured to determine experience information related to each educator in the first meeting session. The experience information may indicate at least one of the following: experience, evaluation, achievements, or feedback related to each educator in the first meeting session. Circuit 202 may be further configured to determine content information related to the content presented during the first meeting session. Circuit 202 may be further configured to retrieve profile information related to a first attendee or other attendees in the first meeting session. The profile information may indicate the first attendee's preferences or interests regarding topics or content presented or discussed in the first meeting session. In another embodiment, circuit 202 may be further configured to determine environmental information related to the geolocation of at least one of the educators or first attendees in the first meeting session. Details of the experience information, content information, profile information and environmental information are shown, for example, in 302F of Figure 3.
[0087] In 402B, an activity detection operation can be performed. In the activity detection operation, circuit 202 can be configured to detect a first set of activities that a first attendee 406 can perform during a first meeting session. The first set of activities can be detected based on applying one or more NN models 210 to a received first set of images 404. The first set of activities can be detected over a first period of the first meeting session (e.g., a few seconds or a few minutes). For example, the first set of activities of the first attendee 406 can be detected over a "10" minute period during the first meeting session. Further details regarding the detection of the first set of activities are shown, for example, in 302C of Figure 3. For example, as illustrated in 302C of Figure 3, the first set of activities of the first attendee 406 may relate to at least one of the following: actions performed by the first attendee 406, gestures performed by the first attendee 406, head posture of the first attendee 406, body position of the first attendee 406, lip movements of the first attendee 406, gaze of the first attendee 406, or facial expressions of the first attendee 406. In one embodiment, (for example, as illustrated in 302B of Figure 3) the circuit 202 may generate a three-dimensional (3D) map of the corresponding first meeting session including the first attendee based on the received first set of images 404. The circuit 202 may further apply one or more NN models 210 to the generated three-dimensional (3D) map of the first meeting session to detect the first set of activities of the first attendee 406.
[0088] In 402C, an interaction duration determination operation can be performed. In the interaction duration determination operation, circuit 202 can be configured to determine a second duration (e.g., in seconds or minutes) of an intercepted interaction 408 between a first attendee 406 and the educator of the first meeting session. Circuit 202 can be further configured to compare the determined second duration with a threshold duration. If the determined second duration is less than the threshold duration, control can proceed to 402E. In other words, in such a case, circuit 202 can discard the intercepted interaction 408. For example, if the threshold duration is 5-10 seconds (i.e., the minimum response time to a question) and the intercepted interaction 408 is shorter (e.g., 2 seconds), circuit 202 can ignore the intercepted interaction 408 because it is considered to be of a short duration that may not be considered a correct response for measuring the attendee's attention or may not demonstrate adequate response quality to the question asked. Otherwise, control can proceed to 402D.
[0089] In one embodiment, circuit 202 can determine the relevance of the captured interaction 408 to the content presented in the first meeting session by comparing the determined second duration with a threshold duration. For example, if one or more questions asked by the first attendee are not related to the content, the educator is likely to simply ignore the questions or give a short response such as "This is irrelevant" or "I've already taught you about this." In such a case, the captured second duration will be short and may be shorter than the threshold duration. On the other hand, if one or more questions asked by the first attendee are related to the content, the educator is likely to give a detailed answer, and therefore the captured second duration is likely to be longer than the threshold duration. Furthermore, the relevant questions asked by the attendee, or the relevant answers given by the attendee, can also indicate to some extent the attendee's attention.
[0090] In one embodiment, the circuit 202 may be further configured to utilize speech recognition technology to identify the voices of the first attendee 406 and the educator in the captured interaction 408. Detailed implementations of the speech recognition technology described above are considered to be well known to those skilled in the art, and therefore, for brevity, a detailed description of the speech recognition technology described above is omitted from this disclosure. In another embodiment, the circuit 202 may be further configured to determine a series of exchanges and the amplitude of those exchanges in the captured interaction 408 between the first attendee 406 and the educator. Based on the series of exchanges and the amplitude of those exchanges, the system 102 may be configured to determine whether the captured interaction 408 of the first attendee 406 is with the educator or with another attendee present in the first meeting session.
[0091] In 402D, a keyword determination operation can be performed. The keyword determination operation can configure circuit 202 to determine one or more keywords in the captured interaction 408. In one embodiment, circuit 202 can be configured to apply one or more NN models 210 to the captured interaction 408 to determine one or more keywords. In another embodiment, circuit 202 can be configured to generate a transcript of the captured interaction 408 to further determine one or more keywords in the captured interaction 408. In such a scenario, one or more NN models 210 may include at least one natural language processing (NLP) model that can be trained to determine one or more keywords in the captured interaction 408. The determined one or more keywords may relate to content presented in the first meeting session. Circuit 202 can be configured to analyze the determined keywords in the captured interaction 408 to determine whether these keywords relate to the topic of the presented content or to a question raised. Based on such analysis and determination, circuit 202 can determine the attention level of the first attendee 406 during the first meeting session.
[0092] In 402E, an attention score calculation operation can be performed. In the attention score calculation operation, circuit 202 can be configured to calculate a first attention score associated with the first attendee 406. The first attention score can indicate the level of attention of the first attendee 406 in the first meeting session and can be calculated over a first period based on a first set of detected activities. System 102 can be configured to calculate a first behavior score and a first interaction score of the first attendee 406 based on the first set of detected activities in order to calculate the attention score. Circuit 202 can be further configured to calculate a first attention score based on the calculated first behavior score and the calculated first interaction score. In another embodiment, circuit 202 can be configured to determine a first facial expression of the first attendee 406 and further calculate a first attention score based on the determined first facial expression. Details regarding the calculation of attention scores (similar to the first attention score) are shown, for example, in 302I of Figure 3.
[0093] In one embodiment, circuit 202 can calculate a first attention score based on a detected first set of activities, a determined second duration, and one or more determined keywords in the captured interaction 408. In another embodiment, the first attention score can be calculated based on experience information, content information, profile information, and environmental information. In one embodiment, circuit 202 can be configured to calculate a first focus score and a first interaction score associated with a first attendee 406 based on a detected first set of activities, and / or a determined second duration, and / or one or more determined keywords, and / or experience information, and / or content information, and / or profile information, and / or environmental information, and / or a determined first facial expression. Circuit 202 can be configured to calculate a first attention score based on the calculated first focus score and the calculated first interaction score. Details regarding the calculation of the first attention score are shown, for example, in Figure 3.
[0094] In 402F, a trained ML model application operation can be performed. In the trained ML model application operation, circuit 202 can be configured to apply a first ML model 402 to a calculated first attention score. Specifically, circuit 202 can be configured to apply a first ML model 402 from among multiple ML models 106 (which can be trained in 302J in Figure 3) to a calculated first attention score. For example, as illustrated in Figure 3, multiple ML models 106 can be trained based on multiple attention scores of multiple attendees 116 associated with multiple meeting sessions in different meeting categories. System 102 can store multiple ML models 106 in memory 204. Multiple attendees 116 may have attended multiple meeting sessions in the past, and multiple ML models 106 are trained based on these meeting sessions.
[0095] In one embodiment, the circuit 202 can apply the first ML model 402 to a calculated first attention score and the meeting category of a first meeting session that the first attendee 406 has attended or is currently attending. The first ML model 402 can be pre-trained based only on data related to the first attendee 406, and thus the first ML model 402 can be personalized for the first attendee. The first ML model 402 can be trained based on historical cumulative data (i.e., a set of past attention scores for the first attendee 406 calculated for different meeting sessions attended in the past, as illustrated, for example, in 302J of Figure 3). In one embodiment, the circuit 202 can be configured to compare the calculated first attention score to a first threshold attention score. The first threshold attention score can be the minimum attention score related to the first attendee 406 and can be determined based on a set of past attention scores that may be related to the first attendee 406. If the calculated first attention score is lower than the first threshold attention score, control can proceed to 402G. Otherwise, control can proceed to termination.
[0096] Circuit 202 can determine, based on a set of past attention scores (from which the first ML model 402 is trained), when the attention of the first attendee 406 significantly decreased (i.e., when the first attendee 406, as indicated by their calculated first attention score, is present in the first meeting session). In other words, based on the first ML model 402 (i.e., an ML model trained on the first attendee 406's past attention score set and different meeting categories), circuit 202 can determine or predict when the calculated first attention score of the first attendee 406 (i.e., the first attendee attending a first meeting session of a particular category) will decrease. For example, the training data for the first ML model 402 shows that the first attendee 406 has a low past attention score in the case of certain meeting sessions (such as online sessions without an instructor, high-intensity meeting sessions, meeting sessions lasting more than one hour, attendees seated in the back, or theoretical content). In such cases, if the first meeting session is of a particular category that is thought to have been trained with a low attention score within the first ML model 402, the circuit 202 may predict in advance that the first attendee 406's first attention score may decrease while attending the first meeting session and may require appropriate action or recommendations to increase the first attendee 406's first attention score during the remainder of the first meeting session.
[0097] In 402G, a recommendation decision operation can be performed. In the recommendation decision operation, circuit 202 can be configured to determine a first set of recommendations. The first set of recommendations may be associated with at least one of the following: a first attendee 406, an instructor in the first meeting session, or first content that can be presented in the first meeting session. Circuit 202 can determine the first set of recommendations based on the application of the first ML model 402 to a calculated first attention score and the meeting category of the first meeting session. In one embodiment, circuit 202 can determine the first set of recommendations based on the determination or prediction (using the trained first ML model 402) that, with respect to the meeting category of the first meeting session, the first attention score of the first attendee 406 in the first meeting session may decrease (i.e., fall below a first threshold attention score). For example, if the first meeting session contains theoretical content, is long, and is conducted by a low-rated instructor, the first attention score may be low or decrease after a certain period of time. Therefore, circuit 202 can determine a first set of recommendations for either the first attendee 406, the educator of the first meeting session, or the content presented in the first meeting session, in order to further avoid a decrease in the first attention score. The first set of recommendations, if followed, can increase the attention score of the first attendee 406 for the remainder of the first meeting session.
[0098] The determined first set of recommendations may include at least one of the following: a first recommendation related to the first attendee 406, a second recommendation related to the educator of the first meeting session, or a third recommendation related to the content presented in the first meeting session. For example, the first recommendation related to the first attendee 406 may include, but are not limited to, one or more individual performance reports with a detailed analysis of attention level breakdowns across content and educators, one or more focus / improvement areas, recommended seating in the meeting session, recommendations for energy drinks (such as coffee) or breaks, and appropriate content or educators based on attendee interests. Second recommendations related to educators in meeting sessions may include, but are not limited to, content modifications (e.g., adding practical or real-time examples), suggestions for similar content with high ratings or many shares / likes (i.e., web links), modifications to content delivery, alerts to educators for significant changes in the attention levels of the first attendees, suggestions for taking breaks or short pauses in the first meeting session, suggestions for updating the duration of the meeting session due to content, suggestions for changing the session duration due to environmental conditions, suggestions for improving environmental conditions (such as improving lighting or acoustic conditions), suggestions for referencing specific qualifications, and suggestions for enhancing teaching ability or evaluation. Third recommendations related to content presented in the first meeting session may include, but are not limited to, content modifications that add other types of content (such as obvious content), content modifications that add interactive content (such as photos, videos, or diagrams), content modifications that remove redundant content, content modifications that remove complex content, and content modifications that add one or more practical sessions to the current theoretical content.
[0099] In some embodiments, the first set of recommendations may function as one or more pieces of feedback (such as 360-degree feedback) for a particular attendee or a particular educator. In other embodiments, one or more pieces of feedback may be two-way feedback between an educator and a particular attendee. In some embodiments, an authority figure (such as an educational institution's management / supervision board, education board, or principal of a school / university) may determine recommendations for making certain decisions (such as, but not limited to, changing to an educator with higher ratings or experience, changing content, adding specific hands-on sessions to the curriculum, shortening the duration of a particular session, adding breaks between meeting sessions, adding a specific educational counselor to the educational staff, updating the overall educational plan, or updating physical conditions or logistics management in meeting sessions).
[0100] In 402H, an output operation can be performed. In the output operation, circuit 202 can be configured to output a determined first set of recommendations. For example, system 102 can control display device 206A to render a first set of recommendations for a first attendee 406 or educator. In some embodiments, circuit 202 can generate an output notification for a first attendee 406 or educator of a first meeting session, based on a calculated first attention score and the application of a first ML model 402 to the meeting category of the first meeting session. The generated notification may include at least one of the first set of recommendations. The notification may be an upfront alert that can be provided to a first attendee 406 or educator of a first meeting session based on a determination or prediction of a low attention score for the first attendee 406 in the first meeting session. The low attention score of the first attendee 406 can be predicted based on the meeting category of the first meeting session and the application of the first ML model 402 (i.e., a personalized model trained for the first attendee 406) to the attention score (i.e., the first attention score) calculated for one or more activities (i.e., behaviors, gestures, interactions, and other factors) performed by the first attendee 406 during the first meeting session. Advance notices can be timely alerted to the educator or the first attendee 406 to take appropriate actions (such as a first set of recommendations) to avoid a further decline in the first attention score or to improve the first attendee 406's first attention score over the remainder of the first meeting session. The first set of recommendations provided to the attendee or educator can further enhance the attendee's learning process. Circuit 202 can further transmit the generated notices to electronic devices that may be relevant to the first attendee 406 or the educator of the first meeting session. The electronic device may include suitable logic, circuits, and interfaces that can be configured to receive notifications from system 102.The electronic device can be further configured to render the received notification. Examples of electronic devices include, but are not limited to, cellular phones, mobile phones, computer devices, smartphones, game devices, mainframe machines, servers, computer workstations, and / or consumer electronic (CE) devices. Accordingly, the disclosed system 102 can effectively control the meeting session to further enhance the attendees' learning process using the output of determined recommendations, based on a real-time assessment of the attendees' attention scores, using a personalized ML model trained for the attendees.
[0101] Figure 5 shows an exemplary scenario for generating a simulated view of multiple attendees related to a meeting session, according to an embodiment of the present disclosure. The description of Figure 5 will be made in relation to the elements of Figures 1, 2, 3, and 4. Figure 5 shows the first figure. Figure 5 shows a system 102 that can include multiple ML models 106. Furthermore, it shows multiple images 502 of multiple attendees 504A to 504N who are able to attend a particular meeting session.
[0102] System 102 can control multiple image acquisition devices 104 (as shown in Figure 1) to acquire multiple images 502 of multiple attendees 504A to 504N. Multiple attendees 504A to 504N may include a first attendee 504A, a second attendee 504B, a third attendee 504C, and an Nth attendee 504N. If the meeting session is held at a physical location, the multiple image acquisition devices 104 may include CCTV cameras that can be installed at specific locations in the physical space. If the meeting session is held virtually (i.e., in online mode), the multiple image acquisition devices 104 may include at least one webcam associated with an electronic device that allows the multiple attendees 504A to 504N to attend the meeting session.
[0103] System 102 can receive multiple images 502 of multiple attendees 504A to 504N from multiple image acquisition devices 104. System 102 can be configured to identify each of the multiple attendees 504A to 504N based on the reception of the multiple images 502. In one embodiment, System 102 can be configured to apply face recognition technology using one or more NN models 210 to each of the multiple images 502 in order to identify each of the multiple attendees 504A to 504N. Detailed implementations of the face recognition technology described above are considered to be well known to those skilled in the art, and therefore, for the sake of brevity, a detailed description of the face recognition technology described above is omitted from this disclosure.
[0104] If the meeting session is held in a physical location (such as a physical classroom), the system 102 can be configured to determine the position of each of the multiple attendees 504A to 504N in the corresponding meeting session based on the multiple images 502 received (or based on a 3D map of the meeting session, as described in 302B of Figure 3, for example). In one embodiment, the circuit 202 of the system 102 can be configured to determine the position of each of the multiple attendees 504A to 504N in the corresponding meeting session based on the captured depth information / multiple depth values. Based on the determined positions, the 3D map of the meeting session, and the attendees' facial recognition, the circuit 202 can tag or associate a particular attendee with a physical seat in the meeting session (such as at the back, front, center, or near a window / door) where the attendee is likely to be sitting while attending the meeting session.
[0105] If a meeting session is virtually open (i.e., such as a web conference or online session held over the internet), system 102 can be configured to generate a simulated view 506 of the corresponding meeting session. In the generated simulated view 506, each of the multiple attendees 504A to 504N can be visualized as sitting next to each other (similar to a physical classroom session). The generated simulated view 506 can be a computer-generated environment that includes realistic-looking scenes and objects, making each of the multiple attendees 504A to 504N feel as if they are immersed in the physical meeting session and environment. In other words, the generated simulated view 506 can mimic a single, comprehensive view of the meeting session for each of the multiple attendees 504A to 504N. In some embodiments, circuit 202 can determine a specific online action (such as an online hand-raising) performed by a first attendee 406 during an online meeting session and further translate this online action into a representation of a similar physical action (such as a physical hand-raising). Circuit 202 can further overlay this representation onto the corresponding attendees visualized within the simulated view 506.
[0106] Figure 6 shows an exemplary user interface for rendering dashboard information according to an embodiment of the present disclosure. The description of Figure 6 will be made in relation to the elements of Figures 1, 2, 3, 4, and 5. Figure 6 shows an electronic UI 600. The electronic UI 600 can be displayed on the display device 206A of the system 102, or on an electronic device associated with the educator or at least one of the multiple attendees 116 based on the reception of a first user input. The first user input can be received via an application interface displayed on the display screen of the display device 206A, or on the display screen of an electronic device associated with the educator or attendee. The application interface can be part of application software such as a software development kit (SDK), a cloud server-based application, a web-based application, an OS-based application / application suite, an enterprise application, or a mobile application.
[0107] In one embodiment, the system 102 can be configured to generate dashboard information. The generated dashboard information may relate to a first attendee 406 among a group of attendees 116 and may be generated based on a calculated attention score of a set of meeting sessions that the first attendee 406 can attend. The generated dashboard may include one or more statistics for at least a set of meeting sessions. The circuit 202 can be further configured to render the generated dashboard information on an electronic UI 600. In some embodiments, for example, dashboard information can be generated for a set of meeting sessions attended by a specific attendee or group of attendees 116 in a particular week, month, or year in the past.
[0108] The electronic UI 600 displays a series of UI elements, such as a first UI element 602, a second UI element 604, a third UI element 606, a fourth UI element 608, a fifth UI element 610, a sixth UI element 612, a seventh UI element 614, and an eighth UI element. The first UI element 602 can be labeled as "Attendance Status" and can show the presence and absence rates of multiple attendees 116 in the first meeting session of a series of meeting sessions. In one example, "Attendance Status" can show the presence / absence rate of the first attendee 406 who have previously attended a series of meeting sessions. The second UI element 604 can be labeled as "Most Attentive Attendee" and can show at least one attendee who has the highest attention score among multiple attendees 116 who have attended a series of meeting sessions. The third UI element 606 can be labeled as "Most Questioning Attendee" and can show at least one attendee who was able to ask the most questions in a series of meeting sessions. Attendees who ask many questions can be determined based on their activities and interactions in past meeting sessions, as illustrated in Figure 3, for example. A fourth UI element 608 can be labeled as an "Attendee in Need of Help" and can indicate at least one attendee who can have the lowest attention score among multiple attendees 116 who have attended a series of meeting sessions. As illustrated in 402G of Figure 4, for example, the disclosed system 102 can suggest different sets of recommendations to such attendees. In one embodiment, the disclosed system 102 can also suggest awards or recognition from the meeting session educator for the "Most Attendee" or the "Attendee Who Asked the Most Questions." Doing so can bring a sense of gamification among multiple attendees in a meeting session and further increase each attendee's attention score in the meeting session.
[0109] The fifth UI element 610 can be labeled as “Best Educators of Last Month” and may show at least one educator who was present in past meeting sessions (e.g., held in the past month) where more attendees achieved a good attention score (or above a certain attention score threshold). The sixth UI element 612 can be labeled as “List of Best Content (Subjects)” and may show at least one piece of content that the first attendee 406 (or the maximum number of attendees) paid more attention to (i.e., could have a high level of attention score). The seventh UI element 614 can be labeled as “Attention Level Graph” and may show a graph (or statistics) between the attention level and time of the first attendee 406 in at least one meeting session out of a series of meeting sessions. For example, the graph may show a decrease in the attention score of a particular attendee based on an increase (or progress) in the time of a particular meeting session. As another example, the statistics may also show patterns or fluctuations in the attention scores of different attendees based on different factors related to the meeting session, such as the category of the meeting session, the duration of the meeting session, the educator, the content, seating position, and environmental conditions. Graphs within dashboard information regarding specific attendees can further alert and / or encourage attendees to effectively follow recommendations in order to increase their attention scores and improve their learning ability and related outcomes. Furthermore, different statistics within the dashboard information (such as the most attentive attendees) can encourage competition among attendees in the meeting session, thereby further increasing the attention scores and learning abilities of different attendees.
[0110] As an example, referring to Figure 6, it can be seen that 70% of the 116 attendees are present in the first meeting session of a series of meeting sessions, while 30% are absent from the first meeting session. Attendees "A" and "B" are thought to have the highest attention scores among the 116 attendees. Attendees "C" and "D" are thought to have asked the most questions in the series of meeting sessions. Attendees "E" and "F" are thought to have the lowest attention scores among the 116 attendees. Educator "A" may be the best educator of the previous month, and their list of best content (subjects) may include science, mathematics, English, computer science, and history.
[0111] In another embodiment, the generated dashboard information may include more statistics that may be relevant to the educator, multiple attendees 116, or the content presented in the corresponding meeting session. The generated dashboard information may also include other statistics based on various parameters such as content, grades, and years. In one embodiment, the generated dashboard information may be presented in a meeting session (e.g., a class) to encourage healthy competition among multiple attendees 116. The generated dashboard information can further be used to evaluate the performance of a first attendee 406 and determine the first attendee's strengths and weaknesses. As shown in Figure 6, the dashboard information may show different sets of recommendations 616, including recommendations for attendees (i.e., attendee 616A), recommendations for the educator (i.e., educator 616B), and recommendations for the presented content (i.e., content 616C), as illustrated, for example, in 402G of Figure 4. The generated dashboard information can also be used by different educational institutions (or educators) to make appropriate decisions (such as following recommendations) to effectively increase attendee attention scores and improve the overall learning process. In some embodiments, circuit 202 can also determine the attention scores of a set of attendees who are located in a specific position within a meeting session, rather than just a specific attendee. For example, circuit 202 can determine the attention scores of a set of attendees who are seated in a specific row (such as the front row or back row) within a meeting session. Based on requests received from attendees, educators, or any educational / professional institution, circuit 202 can control the output of information regarding attention scores (i.e., the attention scores determined for a set of attendees) as dashboard information. Similarly, the dashboard information can present different information regarding attention scores in multiple forms, including (but not limited to) column-level, row-level, section-level, age-level, regional, educator-level, year-level, and designated-level units.Accordingly, the dashboard information generated by the disclosed system 102 functions as a real-time attention level meter for a particular meeting session (e.g., an ongoing lesson), rendering real-time attention scores for different attendees and allowing educators to be further alerted (using different user interface options such as highlighting, or in other formats) about increases or decreases in attendee attention. Thus, the dashboard information can provide real-time visualization of one or more statistics for at least a set of meeting sessions from the perspective of a gamified chart. Similarly, the dashboard information generated by the disclosed system 102 can indicate which specific factors (such as specific rows / columns / hotspot areas, educators, content, meeting session type, session duration, subject / topic) are showing good or need improvement using appropriate real-time output of recommendations. Hotspot areas within a meeting session may correspond to areas where the attention score of each attendee within that range is below a threshold attention score, or areas where it is above a threshold attention score.
[0112] In another embodiment, the disclosed system 102 can identify one or more malpractices that may be committed by a first attendee 406 in a first meeting session. These malpractices may include, but are not limited to, cheating, conversations with other attendees, illegal assistance of an attendee by an instructor (e.g., during an exam), use of unauthorized materials in an examination room, use of illegal and violent gestures by an attendee, and the committing of physical assault. In such a scenario, the circuit 202 can be configured to determine a pattern of the detected first set of activities performed by an attendee (such as the first attendee 406). Based on the determined pattern of the first set of activities, the circuit 202 can be configured to detect one or more malpractices during the first meeting session based on the determined pattern. In one embodiment, the circuit 202 can be configured to apply at least one of a plurality of ML models or one or more NN models 210 to the detected first set of activities in order to detect malpractices. The system 102 can be configured to generate a first notice for the first attendee 406 or the educator of the first meeting session based on the detected illegal activity. In some embodiments, the system 102 may store information about the detected illegal activity in memory 204 for future reference (such as for internal / external investigation or for taking appropriate action to avoid such illegal activity).
[0113] Figure 7 is a flowchart illustrating exemplary operation for training an ML model for meeting session control based on attention determination, according to embodiments of the present disclosure. The description of Figure 7 is made in relation to the elements of Figures 1, 2, 3, 4, 5, and 6. Figure 7 shows flowchart 700. Operations 702-710 can be performed on any computer device, such as system 102 or circuit 202. The operation can start from 702 and proceed to 704.
[0114] In 704, multiple images 114 of multiple attendees 116 associated with multiple meeting sessions, including one or more meeting sessions that belong to different meeting categories, can be received. In one or more embodiments, the circuit 202 can be configured to receive multiple images 114 of multiple attendees 116 associated with multiple meeting sessions, including one or more meeting sessions that belong to different meeting categories. Details regarding the reception of multiple images 114 are shown, for example, in Figures 1, 3 (302A), and 4 (402A).
[0115] In 706, based on the received multiple images 114, it is possible to detect one or more activities performed by each of the multiple attendees during a corresponding meeting session over a period of time. In one or more embodiments, the circuit 202 can be configured to detect one or more activities performed by each of the multiple attendees 116 during a corresponding meeting session over a period of time, based on the received multiple images 114. Details regarding the detection of one or more activities are shown, for example, in Figures 1, 3 (302C), and 4 (402B).
[0116] In 708, based on one or more detected activities associated with the corresponding attendee, an attention score for each of the multiple attendees 116, indicating the attention level of each attendee in the corresponding meeting session, can be calculated over the corresponding period. In one or more embodiments, the circuit 202 can be configured to calculate an attention score for each of the multiple attendees 116, indicating the attention level of each attendee in the corresponding meeting session, based on one or more detected activities associated with the corresponding attendee. Details regarding the calculation of the attention score are shown, for example, in Figure 3(302I).
[0117] In 710, an ML model can be trained for each of the multiple attendees 116 based on the calculated attention score of each of the multiple attendees 116 and the meeting category of the corresponding meeting session. In one or more embodiments, the circuit 202 can be configured to train an ML model for each of the multiple attendees 116 based on the calculated attention score of each of the multiple attendees 116 and the meeting category of the corresponding meeting session. Details regarding the training of multiple ML models 106 for multiple attendees 116 are shown, for example, in Figures 1 and 3 (302J). Control can then proceed to termination.
[0118] Figure 8 is a flowchart illustrating exemplary operations for meeting session control based on attention determination and the application of ML models for specific attendees, according to embodiments of the present disclosure. The description of Figure 8 is made in relation to the elements of Figures 1, 2, 3, 4, 5, 6, and 7. Figure 8 shows flowchart 800. Operations 802-816 can be performed on any computer device, such as system 102 or circuit 202. Operations can start from 802 and proceed to 804.
[0119] In 804, multiple ML models 106 can be stored, which can be trained based on multiple attention scores of multiple attendees related to multiple meeting sessions in different meeting categories. In one or more embodiments, the circuit 202 can be configured to store multiple machine learning (ML) models that can be trained based on multiple attention scores of multiple attendees 116 related to multiple meeting sessions in different meeting categories.
[0120] In 806, a first image set 404 of a first attendee 406 associated with a first meeting session different from multiple meeting sessions among multiple attendees 116 can be received. In one or more embodiments, the circuit 202 can be configured to receive a first image set 404 of a first attendee 406 associated with a first meeting session different from multiple meeting sessions among multiple attendees 116, as described in Figure 4 (402A), for example.
[0121] In 808, based on the received first image set 404, a first set of activities performed by the first attendee 406 during the first meeting session can be detected over a first period of time. In one or more embodiments, the circuit 202 can be configured to detect a first set of activities performed by the first attendee 406 during the first meeting session over a first period of time based on the received first image set 404, as described, for example, in Figure 4(402B).
[0122] In 810, based on the detected first set of activities, a first attention score can be calculated over a first period of time, relating to the first attendee 406 and indicating the attention level of the first attendee 406 in the first meeting session. In one or more embodiments, the circuit 202 can be configured to calculate, for example, a first attention score over a first period of time, relating to the first attendee 406 and indicating the attention level of the first attendee 406 in the first meeting session, based on the detected first set of activities, as described in Figure 4(402B).
[0123] In 812, the first machine learning (ML) model 402 among the multiple ML models 106 can be applied to the calculated first attention score. In one or more embodiments, the circuit 202 can be configured to apply the first machine learning (ML) model 402 among the multiple ML models 106 to the calculated first attention score, as described, for example, in Figure 4 (402F).
[0124] In 814, a first recommendation set can be determined based on the application of the first ML model 402 to the calculated first attention score. In one or more embodiments, the circuit 202 can be configured to determine a first recommendation set based on the application of the first ML model 402 to the calculated first attention score, as described, for example, in Figure 4(402G).
[0125] At 816, the determined first recommended set can be output. In one or more embodiments, the circuit 202 can be configured to output the determined first recommended set, as described, for example, in Figure 4(402H). Control can then proceed to termination.
[0126] Various embodiments of this disclosure can provide a non-temporary computer-readable medium and / or storage medium storing computer-executable instructions that can be executed by a machine and / or computer, such as system 102. The computer-executable instructions can cause a machine and / or computer to perform an operation that may include receiving multiple images (such as multiple images 114) of multiple attendees (such as multiple attendees 116) related to multiple meeting sessions. One or more of the multiple meeting sessions may belong to different meeting categories. The operation may further include detecting, based on the received multiple images 114, one or more activities that each of the multiple attendees 116 can perform during the corresponding meeting session over a period of time. The operation may further include calculating an attention score for each of the multiple attendees 116 over the corresponding period of time based on the detected one or more activities related to the corresponding attendees. The attention score may indicate the level of attention of each attendee in the corresponding meeting session. The operation may further include training a machine learning (ML) model (such as the first ML model 402) for each of the multiple attendees 116, based on the calculated attention score of each of the multiple attendees 116 and the meeting category of the corresponding meeting session.
[0127] Various embodiments of this disclosure can provide a non-temporary computer-readable medium and / or storage medium storing computer-executable instructions that can be executed by a machine and / or computer, such as System 102. The computer-executable instructions can cause a machine and / or computer to perform an operation that may include storing multiple machine learning (ML) models (such as multiple ML models 106) that can be trained on multiple attention scores of multiple attendees associated with multiple meeting sessions of different meeting categories. The operation may further include receiving a first set of images (such as a first set of images 404) of a first attendee (such as a first attendee 406) among the multiple attendees. The first attendee 406 may be associated with a first meeting session which may be different from the multiple meeting sessions. The operation may further include, based on the received first set of images 404, detecting a first set of activities that the first attendee 406 can perform during the first meeting session over a first period of time. The operation may further include calculating a first attention score that may be associated with a first attendee 406 over a first period of time, based on a first set of detected activities. The first attention score may indicate the level of attention of the first attendee 406 in the first meeting session. The operation may further include applying a first machine learning (ML) model 402 from among multiple ML models 106 to the calculated first attention score. The operation may further include determining a first set of recommendations based on the application of the first ML model 402 to the calculated first attention score. The operation may further include outputting the determined first set of recommendations.
[0128] Exemplary embodiments of this disclosure may include systems (such as system 102 in Figure 1) which may include circuits (such as circuit 202). The circuits may be configured to receive multiple images (such as multiple images 114) of multiple attendees (such as multiple attendees 116) related to multiple meeting sessions. One or more meeting sessions among the multiple meeting sessions may have different meeting categories. A meeting category may correspond to at least one of the following: the type of meeting session, the number of attendees in the meeting session, the duration of the meeting session, the average age of the attendees in the meeting session, the topic of the meeting session, the experience of the educator in the meeting session, or the content presented in the meeting session. One or more detected activities performed by each of the multiple attendees may correspond to at least one of the following: actions performed by the attendee, gestures performed by the attendee, head posture of the attendee, body position of the attendee, lip movements of the attendee, gaze of the attendee, or facial expressions of the attendee.
[0129] The circuit 202 can be further configured to generate a first three-dimensional (3D) map of a corresponding meeting session, including at least one of the multiple attendees, based on the multiple images received. Based on the generated first 3D map, the circuit 202 can be further configured to detect one or more activities performed by each of the multiple attendees during the corresponding meeting session over a period of time.
[0130] According to one embodiment, the circuit 202 can be configured to apply one or more neural network (NN) models (such as one or more NN models 210) to a plurality of received images 114. The circuit 202 can be further configured to detect one or more activities performed by each of a plurality of attendees 116 based on the application of one or more NN models 210.
[0131] According to one embodiment, the circuit 202 may be further configured to calculate a concentration score for each of the multiple attendees 116 based on one or more detected activities associated with the corresponding attendee. The circuit 202 may be further configured to calculate an interaction score for each of the multiple attendees 116 based on one or more detected activities associated with the corresponding attendee. The circuit 202 may be further configured to calculate an attention score for each of the multiple attendees 116 based on the calculated concentration score and calculated interaction score of the corresponding attendee.
[0132] According to one embodiment, the circuit 202 can be configured to determine the facial expression of each of the multiple attendees 116 based on the multiple images 114 received, and to further calculate an attention score for each of the multiple attendees 116 based on the determined facial expression of the corresponding attendee, the calculated concentration score, and the calculated interaction score.
[0133] According to one embodiment, circuit 202 can be configured to determine experience information related to each educator in a corresponding meeting session among a plurality of meeting sessions. The experience information may indicate at least one of the following: experience, evaluation, achievements, or feedback related to each educator in the corresponding meeting session. Circuit 202 can be further configured to determine content information related to the content presented during each of the plurality of meeting sessions. Circuit 202 can be further configured to retrieve profile information related to each of the plurality of attendees 116. The profile information may indicate preferences or interests for topics or content related to the corresponding meeting session. Circuit 202 can be further configured to determine environmental information related to the geolocation of at least one of the educators or the plurality of attendees 116 in each meeting session. Circuit 202 can be further configured to calculate an attention score for each of the plurality of attendees 116 based on the determined facial expressions, determined experience information, determined content information, retrieved profile information, determined environmental information, calculated concentration scores, and the calculated interaction scores of the corresponding attendees.
[0134] According to one embodiment, the circuit 202 may be further configured to apply one or more neural network (NN) models to a plurality of received images. Based on the application, the circuit 202 may be further configured to detect one or more objects associated with a plurality of attendees 116 in the plurality of received images 114. Based on the detected one or more objects, the circuit 202 may be further configured to calculate an attention score associated with each of the plurality of attendees 116.
[0135] According to one embodiment, the circuit 202 may be configured to determine a first duration for at least one of one or more activities performed by each of a plurality of attendees 116 during a corresponding meeting session, and to further calculate an attention score for each of the plurality of attendees 116 based on the determined first duration for at least one of the one or more activities.
[0136] According to one embodiment, the circuit 202 may be further configured to generate dashboard information relating to a first attendee 406 among a plurality of attendees 116, based on the calculated attention score of a series of meeting sessions in which the first attendee is present. The generated dashboard may include one or more statistics for at least a series of meeting sessions. The circuit 202 may be further configured to output the generated dashboard information on a display device.
[0137] According to embodiments of the present invention, the circuit 202 can be configured to determine the position of each of the multiple attendees 116 in a corresponding meeting session based on the multiple images 114 received, and to further train an ML model for each of the multiple attendees based on the calculated attention score and the determined position of each of the multiple attendees 116 in the corresponding meeting session.
[0138] Exemplary aspects of the present disclosure include a system (such as system 102 in Figure 1) which can be configured to store multiple machine learning (ML) models (such as multiple ML models 106) which can be trained on multiple attention scores of multiple attendees (such as multiple attendees 116) related to multiple meeting sessions of different meeting categories. System 102 may further include a circuit (such as circuit 202) which can be configured to receive a first set of images (such as a first set of images 404) of a first attendee (such as a first attendee 406) among the multiple attendees 116. The first attendee 406 may be related to a first meeting session which may be different from the multiple meeting sessions. The circuit 202 may further be configured to detect a first set of activities performed by the first attendee 406 during the first meeting session over a first period of time based on the received first set of images 404. Circuit 202 can be further configured to calculate a first attention score over a first period of time, based on the detected first set of activities, which indicates the level of attention of the first attendee 406 in the first meeting session, relating to the first attendee 406. Circuit 202 can be further configured to apply a first machine learning (ML) model (such as the first ML model 402) from among multiple ML models 106 to the calculated first attention score. Circuit 202 can be further configured to determine a first set of recommendations based on the application of the first ML model 402 to the calculated first attention score, and to further output the determined first set of recommendations.
[0139] According to one embodiment, the determined first set of recommendations includes at least one of the following: a first recommendation related to a first attendee 406 among a plurality of attendees 116; a second recommendation related to the educator of the first meeting session; or a third recommendation related to the content presented in the first meeting session.
[0140] According to one embodiment, circuit 202 can be configured to compare a first attention score with a first threshold attention score. Circuit 202 can be further configured to determine a first set of recommendations based on the comparison. Circuit 202 can be further configured to generate a notification containing at least one of the determined first sets of recommendations.
[0141] According to one embodiment, circuit 202 can be configured to control an audio capture device (such as audio capture device 108) to capture interactions (such as interaction 408) between a first attendee 406 and the educator of the first meeting session during the first meeting session. Circuit 202 can be further configured to determine a second duration of the captured interaction 408. Based on the captured interaction 408, circuit 202 can be further configured to determine one or more keywords in the captured interaction 408, and to calculate a first attention score associated with the first attendee 406 based on the determined second duration and the determined one or more keywords in the captured interaction 408.
[0142] According to one embodiment, the circuit 202 can be configured to generate a notification for a first attendee 406 or an educator of the first meeting session, based on the application of a first ML model 402 to a calculated first attention score and the meeting category of the first meeting session. The notification includes at least one of a first set of recommendations.
[0143] According to one embodiment, the circuit 202 may be configured to determine a pattern of a first set of detected activities performed by the first attendee 406. Based on the determined pattern, the circuit 202 may be further configured to detect illegal activity during the first meeting session and, based on the detected illegal activity, generate a notification for the first attendee 406 or the educator of the first meeting session.
[0144] This disclosure can be implemented in hardware or in a combination of hardware and software. This disclosure can be implemented centrally within at least one computer system or in a distributed manner, where different elements can be distributed across multiple interconnected computer systems. A computer system or other device adapted to perform the methods described herein may be suitable. The hardware-software combination may be a general-purpose computer system including a computer program that, when loaded and executed, can control the computer system to perform the methods described herein. This disclosure can also be implemented in hardware, including a portion of an integrated circuit that also performs other functions.
[0145] This disclosure includes all features that enable the implementation of the methods described herein and can be incorporated into a computer program product that can perform these methods when loaded onto a computer system. In this context, a computer program means any expression in any language, code, or notation of an instruction set intended to be executed directly, or after either a) conversion to another language, code, or notation, or b) reproduction in a different content form, on a system having information processing capabilities.
[0146] While this disclosure has been described with reference to several embodiments, those skilled in the art will understand that various modifications can be made and equivalents can be substituted without departing from the scope of this disclosure. Furthermore, many modifications can be made without departing from the scope of this disclosure to suit specific circumstances or content to the teachings of this disclosure. Accordingly, this disclosure is not limited to the specific embodiments disclosed, but is intended to include all embodiments that fall within the scope of the appended claims. [Explanation of Symbols]
[0147] 104 Multiple image acquisition devices 106 Multiple ML Models 114 Multiple images 210 One or more neural network models 300 Block Diagram 302A Data Acquisition 302B 3D Map Generation 302C Activity Detection 302D Concentrated Score Calculation 302E Interaction Score Calculation Information for 302F confirmed. 302G facial expression determination 302H Object Detection 302I Attention Score Calculation 302J ML model training 304 Multiplexer
Claims
1. We receive multiple images of multiple attendees related to multiple meeting sessions, including one or more meeting sessions with different meeting categories. Based on the multiple images received, one or more activities performed by each of the multiple attendees during the corresponding meeting session are detected over a certain period of time. Based on the one or more detected activities associated with the corresponding attendee, an attention score is calculated for each of the multiple attendees over the corresponding period, indicating the level of attention each attendee had during the corresponding meeting session. A machine learning (ML) model is trained for each of the multiple attendees based on the calculated attention score for each of the multiple attendees and the meeting category of the corresponding meeting session. It has a circuit configured as follows: The aforementioned circuit is Based on the one or more detected activities related to the corresponding attendee, the concentration score for each of the multiple attendees is calculated. Based on the one or more detected activities related to the corresponding attendee, the interaction score for each of the multiple attendees is calculated. A system characterized by calculating the attention score of each of the plurality of attendees based on the calculated concentration score and the calculated interaction score of the corresponding attendee.
2. The meeting category corresponds to at least one of the following: the type of meeting session, the number of attendees in the meeting session, the duration of the meeting session, the average age of the attendees in the meeting session, the topic of the meeting session, the experience of the educator in the meeting session, or the content presented in the meeting session. The system according to claim 1.
3. The detected one or more activities performed by each of the aforementioned attendees are related to at least one of the following: an action performed by the attendee, a gesture performed by the attendee, the attendee's head posture, the attendee's body position, the attendee's lip movements, the attendee's gaze, or the attendee's facial expression. The system according to claim 1.
4. The aforementioned circuit is Based on the multiple images received, a first three-dimensional (3D) map of the corresponding meeting session is generated, including at least one of the multiple attendees. Based on the generated first 3D map, the one or more activities performed by the multiple attendees are detected. The system according to claim 1, further configured as follows.
5. The aforementioned circuit is Based on the multiple images received, the facial expressions of each of the multiple attendees are determined. Based on the determined facial expressions, calculated concentration scores, and calculated interaction scores of the corresponding attendees, the attention score of each of the multiple attendees is calculated. The system according to claim 1, further configured as follows.
6. The aforementioned circuit is Determine experience information relating to each educator in a corresponding meeting session among the multiple meeting sessions, showing at least one of the following: experience, evaluation, achievements, or feedback related to each educator in the corresponding meeting session. Determine the content information related to the content presented during each of the aforementioned meeting sessions. Based on the determined experience information, the determined content information, the calculated concentration score and the calculated interaction score of the corresponding attendee, the attention score of each of the multiple attendees is calculated. The system according to claim 1, further configured as follows.
7. The aforementioned circuit is Search for profile information relating to each of the aforementioned multiple attendees, showing their preferences or interests in topics or content related to the corresponding meeting session. Based on the retrieved profile information of the corresponding attendee, the calculated concentration score, and the calculated interaction score, the attention score of each of the multiple attendees is calculated. The system according to claim 1, further configured as follows.
8. The aforementioned circuit is Determine the geolocation-related environmental information for at least one of the educators or attendees of each meeting session. Based on the determined environmental information, the calculated concentration score, and the calculated interaction score of the corresponding attendee, the attention score of each of the multiple attendees is calculated. The system according to claim 1, further configured as follows.
9. The aforementioned circuit is Applying one or more neural network (NN) models to the multiple images received, Based on the application of the one or more NN models, the one or more activities performed by each of the multiple attendees are detected. Based on the one or more detected activities related to the corresponding attendee, the attention score of each of the multiple attendees is calculated over the period. The system according to claim 1, further configured as follows.
10. The aforementioned circuit is Applying one or more neural network (NN) models to the multiple images received, Based on the above application, one or more objects related to the multiple attendees in the multiple received images are detected, Based on the one or more objects detected, an attention score associated with each of the multiple attendees is calculated. The system according to claim 1, further configured as follows.
11. The aforementioned circuit is Determine the first duration of at least one of the one or more activities performed by each of the multiple attendees during the corresponding meeting session. Based on the determined first duration of at least one of the one or more activities, the attention score of each of the multiple attendees is calculated. The system according to claim 1, further configured as follows.
12. The aforementioned circuit is A first set of images of a first attendee, from among the multiple attendees, who is associated with a first meeting session different from the multiple meeting sessions, is received. Based on the received first set of images, a first set of activities performed by the first attendee during the first meeting session is detected over a first period of time. Based on the detected first set of activities, a first attention score associated with the first attendee is calculated over the first period. The first attention score calculated is then applied to a first machine learning (ML) model, which is one of several trained ML models and is trained based on a set of past attention scores associated with the first attendee. Based on the application of the first ML model to the calculated first attention score, a first recommendation set is determined. Output the first recommended set determined above. The system according to claim 1, further configured as follows.
13. The circuit is further configured to generate notifications for the first attendee or the educator of the first meeting session based on the application of the first ML model to the calculated first attention score and the meeting category of the first meeting session. The aforementioned notification includes at least one of the first set of recommendations, The system according to claim 12.
14. The aforementioned circuit is Based on the calculated attention score of a series of meeting sessions attended by the first attendee, dashboard information is generated that includes at least one or more statistical values for the series of meeting sessions related to the first attendee among the multiple attendees. The generated dashboard information is output to a display device. The system according to claim 1, further configured as follows.
15. The aforementioned circuit is Based on the received multiple images, the position of each of the multiple attendees in the corresponding meeting session is determined. The ML model is trained for each of the multiple attendees based on the calculated attention score and the determined position of each of the multiple attendees in the corresponding meeting session. The system according to claim 1, further configured as follows.
16. A memory configured to store multiple machine learning (ML) models trained on multiple attention scores of multiple attendees related to multiple meeting sessions in different meeting categories, A circuit that is communicatively coupled to the memory, The circuit is equipped with, A first set of images of a first attendee, from among the multiple attendees, who is associated with a first meeting session different from the multiple meeting sessions, is received. Based on the received first set of images, a first set of activities performed by the first attendee during the first meeting session is detected over a first period of time. Based on the detected first set of activities, a first attention score is calculated over a first period of time, relating to the first attendee and indicating the first attendee's level of attention during the first meeting session. The first ML model among multiple machine learning (ML) models is applied to the first attention score calculated above. Based on the application of the first ML model to the calculated first attention score, a first recommendation set is determined. Output the first recommended set determined above. During the first meeting session, the audio capture device is controlled to capture the interaction between the first attendee and the educator of the first meeting session. Determine the second duration of the incorporated interaction. Based on the incorporated interactions, one or more keywords in the incorporated interactions are determined. Based on the determined second duration and the determined one or more keywords in the incorporated interaction, the first attention score associated with the first attendee is calculated. It is configured in such a way. A system characterized by the following features.
17. The first set of recommendations determined above includes at least one of the following: a first recommendation relating to a first attendee among the plurality of attendees, a second recommendation relating to the educator of the first meeting session, or a third recommendation relating to the content presented in the first meeting session. The system according to claim 16.
18. The aforementioned circuit is The first attention score is compared with the first threshold attention score. Based on the above comparison, the first recommended set is determined. A notification is generated that includes at least one of the first set of recommendations determined above. The system according to claim 16, further configured as follows.
19. The circuit is further configured to generate notifications for the first attendee or the educator of the first meeting session based on the application of the first ML model to the calculated first attention score and the meeting category of the first meeting session. The aforementioned notification includes at least one of the first set of recommendations, The system according to claim 16.
20. The aforementioned circuit is Determine the pattern of the detected first set of activities performed by the first attendee, Based on the determined pattern, illegal activity during the first meeting session is detected. Based on the detected illegal activity, generate a notice for the first attendee or the educator of the first meeting session. The system according to claim 16, further configured as follows.
21. In the system, Receiving multiple images of multiple attendees related to multiple meeting sessions, including one or more meeting sessions in different meeting categories, Based on the multiple images received, one or more activities performed by each of the multiple attendees during the corresponding meeting session are detected over a certain period of time. Based on the one or more detected activities related to the corresponding attendee, calculate the attention score for each of the multiple attendees over the corresponding period, which indicates the level of attention each attendee had during the corresponding meeting session. Training a machine learning (ML) model for each of the aforementioned attendees based on the calculated attention score of each of the aforementioned attendees and the meeting category of the corresponding meeting session, Includes, Calculating the attention score of each of the aforementioned attendees over the corresponding period is: Based on the one or more detected activities related to the corresponding attendee, the concentration score for each of the multiple attendees is calculated. Based on the one or more detected activities related to the corresponding attendee, the interaction score for each of the multiple attendees is calculated. A method characterized by comprising calculating an attention score for each of the plurality of attendees based on the calculated concentration score and the calculated interaction score of the corresponding attendee.