System and method for improving visual learning through gaze tracking
The gaze tracking technology determines participants' participation in remote and virtual scenarios, adjusts meeting attributes to improve engagement, solves the problem that presenters cannot adjust content delivery, and achieves more effective remote and virtual teaching or collaboration.
Patent Information
- Application Number
- CN202380084844.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-12-12
- Filing Date
- 2023-12-06
- Publication Date
- 2025-07-18
AI Technical Summary
In remote and virtual teaching or collaborative scenarios, the presenter has difficulty determining the level of participants’ participation, resulting in the inability to adjust the style, format, or pace of content delivery to improve participation.
Through gaze tracking techniques, define the gaze focus of predefined targets, establish the gaze pattern of participants, analyze the gaze data to determine the degree of participation, and adjust the attributes of the communication session, such as visual or auditory aspects, based on the degree of participation.
It realizes the real-time or preset intervals of meeting content in remote and virtual scenes to improve participants' participation, solves the problem that the presenter cannot determine the degree of participation, and improves the effectiveness of teaching or collaboration.
Smart Images

Figure CN120344940A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure generally relates to improving visual learning through gaze tracking, and more particularly to systems and methods for improving visual learning by measuring user engagement by means of gaze tracking. Background Art
[0002] For many reasons, remote and / or virtual collaboration and training have become increasingly common methods of presenting and otherwise delivering information. The proliferation of wearable augmented reality and the availability of high-bandwidth networking solutions have facilitated further technological advancements in this regard. Effective engagement of participants in a visual learning or teaching session (i.e., those who view or otherwise receive information) is important for successful remote and / or virtual collaboration. For in-person visual learning or teaching sessions, it is possible to determine the engagement of participants based on general observations of inter-eye contact, head movement, facial expression, and / or verbal responses of the participants.
[0003] Typically, during such in-person meetings involving a visual component, the presenter points to the component while sharing information verbally with the recipient. The presenter obtains cues from where the recipient is looking to know if the recipient is looking at a target different from the intended target to look at, to know if the recipient is following the instructions and uses it as non-verbal feedback. In the case where the recipient is not focusing on where the presenter intends or expects the recipient to look, the presenter can employ reorientation cues to figure out if the recipient has understood or correct the communication style to make the meeting more engaging and effective. However, in a remote and / or virtual scenario, the presenter cannot easily determine the level of engagement of the participants to adjust the style, format, pace, or other attributes of the content delivery. Summary of the Invention
[0004] As described herein, the applicant has recognized and realized that it is possible to use concept-based gaze tracking to evaluate the engagement and associated effectiveness of participants in a remote collaboration or tutoring, training, or teaching session involving a shared screen or virtual reality space. Additionally, the applicant has recognized and realized that using such gaze tracking can improve engagement and other outcomes in the context of remote collaboration, tutoring, training, and / or teaching sessions.
[0005] According to an embodiment of the present disclosure, a method for evaluating user participation during a communication session of a participant related to at least one predefined goal is provided. The method includes: defining a gaze focus associated with the at least one predefined goal; establishing a gaze pattern of the participant during the communication session by tracking the gaze focus relative to the at least one predefined goal; determining the degree of participation of the participant based at least in part on analyzing data associated with the gaze pattern of the participant; and adjusting at least one attribute of the communication session based on the determined degree of participation.
[0006] In one aspect, the method further includes communicating the degree of participation to at least the participant.
[0007] In one aspect, the degree of participation can be communicated substantially in real time.
[0008] In one aspect, the degree of participation can be communicated substantially at a preset interval.
[0009] In one aspect, the communication session can be conducted in a virtual reality environment.
[0010] In one aspect, the gaze pattern of the participant during the communication session can be established by: tracking the gaze focus relative to the manipulation of the at least one predefined goal.
[0011] In one aspect, the method further includes: communicating a set of cues associated with the at least one predefined goal to the participant, wherein the gaze focus of the participant is tracked relative to the participant's response to the set of cues.
[0012] In one aspect, the set of cues can be associated with the manipulation of the at least one predefined goal.
[0013] In one aspect, the method further includes: mapping the relationship between the set of cues and the at least one predefined goal, wherein the mapping is performed using natural language processing and named entity recognition.
[0014] In one aspect, the response of the participant can include verbal feedback.
[0015] In one aspect, the at least one attribute of the communication session can be automatically adjusted based on the determined degree of participation.
[0016] In one aspect, the at least one attribute of the communication session can include a visual aspect of the communication session, an auditory aspect of the communication session, or a combination thereof.
[0017] In one aspect, adjusting the at least one attribute of the communication session based on the determined level of engagement may include: repeating at least a portion of the communication session after the level of engagement drops below a predetermined threshold.
[0018] In one aspect, adjusting the at least one attribute of the communication session based on the determined level of engagement may include: changing the content or delivery of the communication session after the level of engagement drops below a predetermined threshold.
[0019] According to another embodiment of the present disclosure, there is provided a system for evaluating user engagement during a communication session of participants related to at least one predefined goal. The system includes: a server including a memory for storing data related to defining a gaze focus associated with the at least one predefined goal; a camera for capturing the gaze of the participants during the communication session; and one or more processors in communication with the server and the camera, wherein the one or more processors are configured to: establish a gaze pattern of the participants during the communication session by tracking the gaze focus relative to the at least one predefined goal; determine the level of engagement of the participants at least in part based on analyzing data associated with the gaze pattern of the participants; and adjust at least one attribute of the communication session based on the determined level of engagement.
[0020] In one aspect, the one or more processors are further configured to: communicate the level of engagement to at least the participants substantially in real time.
[0021] In one aspect, the one or more processors are further configured to: communicate a set of cues associated with the at least one predefined goal to the participants, wherein the gaze focus of the participants is tracked relative to the participants' response to the set of cues.
[0022] In one aspect, the set of cues may be associated with manipulation of the at least one predefined goal.
[0023] According to another embodiment of the present disclosure, there is provided a virtual reality system for evaluating user participation during a communication session of a participant related to at least one predefined goal. The virtual reality system includes: a head-mounted device that covers at least a part of the field of view; a server that includes a memory for storing data related to defining a gaze focus associated with the at least one predefined goal; and one or more processors that communicate with the head-mounted device and the server, wherein the one or more processors are configured to: establish a gaze pattern of the participant during the communication session by tracking the gaze focus relative to the at least one predefined goal; determine the degree of participation of the participant at least partially based on analyzing data associated with the gaze pattern of the participant; and adjust at least one attribute of the communication session of the participant based on the determined degree of participation.
[0024] In one aspect, the at least one attribute of the communication session can be automatically adjusted based on the determined degree of participation.
[0025] Referring to the embodiments described below, these and other aspects of the various embodiments will be apparent and elucidated. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] In the drawings, like reference numerals generally refer to like parts throughout the different views. Also, the drawings are not necessarily to scale, but generally focus on illustrating the principles of the various embodiments.
[0027] Figure 1 is a schematic diagram illustrating the use of a system for evaluating user participation according to aspects of the present disclosure.
[0028] Figure 2 is a block diagram of a system for evaluating user participation according to aspects of the present disclosure.
[0029] Figure 3 is a block diagram of a participant participation analyzer according to aspects of the present disclosure.
[0030] Figure 4A is an illustration showing a focus pattern according to aspects of the present disclosure.
[0031] Figure 4B is an illustration showing related to Figure 4A the gaze pattern associated with the focus pattern according to aspects of the present disclosure.
[0032] Figure 5A is a schematic diagram illustrating the focus difference in a communication session according to aspects of the present disclosure.
[0033] Figure 5B FIG. is a schematic diagram showing how to calculate the focus difference according to aspects of the present disclosure.
[0034] Figure 6 FIG. is a flowchart showing a method for evaluating user engagement according to aspects of the present disclosure.
[0035] Figure 7 FIG. is a schematic diagram showing a virtual reality system for evaluating user engagement according to aspects of the present disclosure. DETAILED DESCRIPTION
[0036] The present disclosure relates to methods and systems for improving visual learning through gaze tracking. More specifically, the present disclosure relates to methods and systems for evaluating the engagement of users or participants during a telecommunication conference having a shared visual component including at least one predefined target. Examples of such telecommunication conferences include, but are not limited to, teaching demonstrations, technical support, training or tutoring sessions, or team collaboration. In an embodiment, the telecommunication conference can be conducted virtually and / or remotely in a two-dimensional collaborative space (e.g., involving a shared screen) and / or a three-dimensional collaborative space (e.g., a virtual reality or mixed reality environment). In a particular embodiment, the telecommunication conference includes at least a visual component (i.e., visible to the participants), but can also include other components (e.g., an auditory component, a tactile component, an interactive component, etc.).
[0037] As described herein, the methods and systems of the present disclosure achieve improved results in remote and / or virtual teaching scenarios, e.g., improved engagement. Additionally, the methods and systems of the present disclosure can advantageously enable new opportunities for addressing user engagement and potential misunderstandings.
[0038] Turning to Figure 1 , FIG. shows a schematic diagram of a communication conference 100 and one or more systems 102A, 102B configured to evaluate the engagement of one or more participants 104A, 104B according to aspects of the present disclosure. As described herein, the systems 102A, 102B can be configured to evaluate the engagement of one or more participants 104A, 104B during the communication conference 100 using gaze tracking devices 106A, 106B.
[0039] According to the present disclosure, a communication session 100 refers to an information presentation that includes visual components 108 that can be transmitted or otherwise presented to one or more participants 104A, 104B. In some embodiments, the communication session 100 can be a two-dimensional virtual environment, such as, for example, a view of a shared screen or monitor, a two-dimensional virtual reality environment, and / or a two-dimensional augmented reality environment. In additional embodiments, the communication session 100 can be a three-dimensional virtual environment, such as, for example, a shared virtual reality environment and / or an augmented reality environment. In some embodiments, the communication session 100 can be live (i.e., transmitted or presented to one or more participants 104A, 104B in real time), can be pre-recorded, or can be a combination of live and pre-recorded.
[0040] In an embodiment, the communication session 100 can be presented by one or more presenters 110. For example, in an embodiment, the presenter 110 can virtually share a teaching presentation (i.e., including visual components 108) from the presenter device 112. The communication session 100 can be transmitted or otherwise presented to one or more participants 104A, 104B via one or more participant devices 116A, 116B over a communication network 114 (discussed in more detail below). In an embodiment, the communication session 100 can be transmitted or otherwise presented to one or more participants 104A, 104B simultaneously and / or at different times.
[0041] Each of the one or more participants 104A, 104B receives, views, and / or otherwise participates in the communication session 100 near a system 102A, 102B configured to evaluate user engagement. As described herein, user engagement can be evaluated with respect to one or more visual aspects 118, 120 of the visual components 108 of the communication session 100. More specifically, user engagement can be evaluated using gaze tracking related to one or more visual aspects 118, 120 of the visual components 108 of the communication session 100, which are transmitted or otherwise visually presented to one or more participants 104A, 104B (e.g., predefined targets 118A, 118B, 120A, 120B).
[0042] Systems 102A, 102B for evaluating user engagement during a communication session 100 can at least include: (i) remote servers 122A, 122B that include a memory for storing data related to defining a gaze focus associated with at least one predefined goal 104B (e.g., predefined goals 118A, 118B, 120A, 120B); (ii) gaze tracking devices 106A, 106B configured to record multiple images capturing the eye movements of one or more participants 104A, 104B and / or muscle responses / electrical activities associated with the eye movements of one or more participants 104A, 104B; and (iii) one or more processors (e.g., Figure 2 processor 202 as shown) that communicate with servers 122A, 122B and gaze tracking devices 106A, 106B.
[0043] In an embodiment, one or more processors (e.g., Figure 2 processor 202 as shown) can be configured to perform one or more steps of the methods described herein, the one or more steps including but not limited to the following: (i) defining a gaze focus associated with at least one predefined goal (e.g., predefined goals 118A, 118B, 120A, 120B); (ii) establishing a gaze pattern of one or more participants 104A, 104B during the communication session 100 by tracking the gaze focus relative to at least one predefined goal (e.g., predefined goals 118A, 118B, 120A, 120B); (iii) determining the degree of engagement of participants 104A, 104B based at least in part on analyzing data associated with the gaze pattern of one or more participants 104A, 104B; and (iv) adjusting at least one attribute of the communication session 100 based on the determined degree of engagement for one or more participants 104A, 104B.
[0044] In an embodiment, gaze detection devices 106A, 106B can include one or more components configured to measure and record facial features of one or more individuals (e.g., participants 104A, 104B) within corresponding fields of view 124A, 124B. According to certain aspects, gaze detection devices 106A, 106B can include cameras 126A, 126B enabling eye tracking. In additional aspects, gaze detection devices 106A, 106B can include muscle-based eye tracking systems, such as electromyography (EMG) devices. As such, gaze detection devices 106A, 106B can be configured to generate eye gaze data, which can include, for example but not limited to, multiple images capturing the eye movements of one or more participants 104A, 104B and / or muscle responses / electrical activities associated with the eye movements.
[0045] In an embodiment, servers 122A, 122B can include a memory storing instructions for defining one or more gaze foci associated with one or more predefined targets (e.g., predefined targets 118A, 118B, 120A, 120B). In an embodiment, one or more processors (e.g., Figure 2 the processor 202 shown) can be operably connected to and / or communicate with gaze detection devices 106A, 106B such that one or more processors 202 can receive eye gaze data. Similarly, one or more processors 202 can be operably connected to and / or communicate with servers 122A, 122B such that one or more processors 202 can use the instructions to define one or more gaze foci associated with one or more predefined targets 118A, 118B, 120A, 120B and determine the gaze patterns of one or more participants 104A, 104B.
[0046] Further referring Figure 2 , system 102 (e.g., systems 102A, 102B) can include an electronic device 200 that includes one or more processors 202. In an embodiment, device 200 can include one or more processors 202 (also referred to as a central processing unit or CPU), a machine-readable memory 204, an interface bus 206, all of which can be interconnected and / or communicate via a system bus 208 that includes conductive circuit paths through which instructions (e.g., machine-readable signals) can travel to effect communication, tasks, storage, etc. Device 200 can be connected to a power source 210, which can include an internal power source and / or an external power source.
[0047] In an embodiment, one or more processors 202 can include a high-speed data processor sufficient to execute program components, and the high-speed data processor can include various specialized processing units known in the art. The general-purpose processor can be a microprocessor or can also be any conventional processor, controller, microcontroller, or state machine. In some embodiments, one or more of the features described herein can be implemented on components such as application-specific integrated circuits (“ASICs”), digital signal processors (“DSPs”), field-programmable gate arrays (“FPGAs”), graphics processing units (“GPUs”), and / or similar electronic devices.
[0048] In an embodiment, interface bus 206 may include: an input / output interface 212 configured to connect device 200 to one or more peripheral devices (e.g., gaze detection device 106); a network interface 214 configured to connect device 200 to a communication network 114 (e.g., using network protocols such as IEEE 802.3 and / or 802.11); and / or a memory interface 216 configured to receive, communicate with, and / or connect to a plurality of machine-readable memory devices (e.g., storage device 224, removable storage device, etc.).
[0049] In various aspects, network interface 214 operably connects device 200 to communication network 114, which can include direct interconnection, the Internet, a local area network (“LAN”), a metropolitan area network (“MAN”), a wide area network (“WAN”), a wired or Ethernet connection, a wireless connection, and similar types of communication networks, including combinations thereof. In some embodiments, one or more user servers 220 and / or cloud-based services 222 may be connected to device 200 via communication network 114 and network interface 214.
[0050] Memory 204 can be embodied differently in one or more forms of machine-accessible and machine-readable memory. In some examples, memory 204 includes storage device 224, which includes one or more types of memory. For example, storage device 224 can include, but is not limited to, non-transitory storage media, disk storage devices, optical disk storage devices, arrays of storage devices, solid-state memory devices, etc., including combinations thereof.
[0051] Generally, memory 204 is configured to store data / information 240 and instructions 230, which, when executed by one or more processors 202, cause device 200 to perform one or more tasks. In an embodiment, memory 204 can include a participation situation analyzer 226, which includes a collection of programs and / or database components and / or data. Depending on the specific implementation, participation situation analyzer 226 can include software components, hardware components, and / or some combination of both hardware and software components.
[0052] In a particular embodiment, participation situation analyzer 226 and / or one or more individual software packages can be stored in local storage device 224. In other examples, participation situation analyzer 226 and / or one or more individual software packages can be loaded onto and / or updated from remote server 220 via communication network 114.
[0053] For example, referring to Figure 3, the engagement analyzer 226 can include, but is not limited to, instructions 230 having a gaze analysis component 231, a natural language processing ("NLP") component 232, a mapping component 233, an engagement component 234, and / or a feedback component 236. These components can be incorporated into system 102, loaded from system 102, loaded onto system 102, or otherwise operably available to or obtained from system 102. Similarly, device 200 or portions thereof can be incorporated into system 200, loaded from system 200, loaded onto system 200, or otherwise operably available to or obtained from system 200. For example, although program components can be stored in local storage device 224, they can also be loaded and / or stored in other memories (e.g., a remote cloud storage facility accessible via a communication network (e.g., communication network 114)).
[0054] The gaze analysis component 231 can be a stored program component executed by at least one processor (e.g., one or more processors 202 of device 200). In an embodiment, the gaze analysis component 231 can receive gaze data 241 as input and analyze the gaze data 241 to determine one or more gaze tracking measurements 242 associated with at least one participant 104A, 104B. In certain embodiments, these gaze tracking measurements 242 can include, but are not limited to, fixation count, regression fixation count, fixation duration, amplitude, saccade peak velocity, blink rate or interblink interval, blink amplitude, and blink duration, phasic pupil diameter, tonic pupil diameter, etc. Each of these measurements is summarized in Table 1 below.
[0055] Table 1
[0056]
[0057] In a particular embodiment, the gaze analysis component 231 can be configured to establish a gaze pattern 243 of at least one participant 104A, 104B during a communication session 100 based on one or more gaze tracking measurements 242. In an embodiment, the gaze pattern 243 can include a time-coded sequence of gaze tracking measurements 242, which can be compared to a time-coded gaze focus on one or more predefined targets (e.g., visual targets 118A, 118B, 120A, 120B) visible to, for example, (one or more) participants 104A, 104B. For example, refer to Figure 4A and Figure 4B, a focus pattern 402 of one or more predefined objectives of the communication session 100 is illustrated in a first time window. A gaze pattern 404 of participants (e.g., participants 104A, 104B) measured according to the present disclosure is illustrated in the same time window. In an embodiment, the gaze pattern 404 can be stored as gaze pattern data 243.
[0058] Thus, in an embodiment, the gaze pattern data 243 and / or the gaze pattern 404 of each participant 104A, 104B can be analyzed relative to one or more predefined objectives 118A, 118B, 120A, 120B to determine the degree of participation of the participants 104A, 104B. In some embodiments, the participation analyzer 226 can include a participation component 234 configured to analyze the gaze pattern data 243 and / or the gaze pattern 404 of one or more participants 104A, 104B relative to one or more predefined objectives 118A, 118B, 120A, 120B to determine the degree of participation of the participants 104A, 104B. In a particular embodiment, the participation component 234 is configured to analyze the gaze pattern data 243 and / or the gaze pattern 404 of one or more participants 104A, 104B relative to the focus pattern 402 of one or more predefined objectives 118A, 118B, 120A, 120B to determine the degree of participation of the participants 104A, 104B.
[0059] In an embodiment, the participation component 234 can be a stored program component executed by at least one processor (e.g., one or more processors 202 of the device 200). In particular, the participation component 234 can receive gaze tracking measurements 242, gaze pattern data 243, and / or focus pattern data 244 as inputs. Then, the participation component 234 can generate a degree of participation 245 of one or more participants 104A, 104B based thereon. That is, the participation component 234 can analyze the gaze tracking measurements 242, gaze pattern data 243, and / or focus pattern data 244 of at least one participant 104A, 104B of the communication session 100 to determine at least one measurement of the degree of participation 245 of the participants 104A, 104B.
[0060] In a particular embodiment, for example, the degree of participation 245 of the participants 104A, 104B can be evaluated based on a focus difference (e.g., Δ f ocal) between the gaze positions of the participants 104A, 104B and the gaze foci associated with one or more predefined objectives 118A, 118B, 120A, 120B. For example, the focus difference (e.g., Δ f ocal) can be calculated as follows:
[0061]
[0062] Where (x1, y1, z1) are the coordinates of the fixation position, and (x2, y2, z2) are the coordinates of the gaze focus for a predefined target.
[0063] Reference Figure 5A and Figure 5B and, a schematic diagram showing how the degree of engagement 245 of a participant 104 viewing the visual component 108 of a illustrated viewing communication session (e.g., communication session 100) can be determined in accordance with aspects of the present disclosure. As shown, the gaze tracking device 106 can be used to track the gaze 502 of the participant 104 within the field of view 124 when the participant 104 views the visual component 108 of the communication session. The visual component 108 can include one or more predefined targets 504, and based on the information captured by the gaze tracking device 106, it can be determined that the gaze 502 of the participant 104 is directed at a first target 506 rather than a second target 508. In an embodiment, the second target 508 can be the intended focus, so the gaze 502 of the participant 104 indicates a loss of focus. As Figure 5B shown, the coordinates of the fixation position (e.g., at the target 506) and the coordinates of the gaze focus (e.g., at the target 508) can be determined by the system of the present disclosure (e.g., systems 102A, 102B), and then the focus difference can be calculated.
[0064] In an embodiment, the degree of engagement 245 of the participants 104A, 104B can be evaluated as a weighted or unweighted average of the focus differences calculated over a period of time (e.g., over the entire duration of the communication session), as follows:
[0065]
[0066] Where E session is a measure of the degree of engagement of the participant in the communication session or a portion thereof, N is the number of samples of Δ focal calculated in the session or a portion thereof, and ∑Δ focal is the sum of Δ f ocal at N time points.
[0067] In accordance with the present disclosure, focus data 244 (i.e., gaze focus and focus patterns associated with one or more predefined targets) can be determined and / or defined according to a predefined target definition rule 246. In some embodiments, the target definition rule 246 can include rules for tracking a pointing device (e.g., a virtual laser pointer, a mouse cursor, etc.).
[0068] In other embodiments, the target definition rule 246 includes rules for analyzing the auditory components of the communication session 100 to identify one or more predefined targets and define the (one or more) points of gaze. For example, in some embodiments, the natural language processing component 232 can be used to process the auditory components of the communication session 100 and use a domain dictionary or ontology to extract domain concepts. The domain concepts identified by the NLP component 232 can then be mapped by the mapping component 233 to the visual elements of the communication session 100. For example, in Figure 1 the example of, the host or presenter 110 is sharing a visual component 108 that includes a depiction of a heart 118 and a pair of lungs 120. In the auditory component, the presenter 110 may refer to the "heart", the NLP component 232 can identify the "heart", and the mapping component 233 can identify the position of the heart 118 relative to the visual component 108 when speaking.
[0069] In an embodiment, the mapping component 233 can include a named entity recognition (NER) engine configured to associate aspects or concepts in a portion of the communication session 100 (e.g., the auditory component) with another portion of the communication session 100 (e.g., the visual component). In some embodiments, the mapping component 233 can include a machine learning algorithm (e.g., a convolutional neural network) to detect the targets identified by the NLP component 232. In a particular embodiment, the mapping component 233 can include a YOLO algorithm configured to detect the targets identified by the NLP component 232.
[0070] In other words, the host 110 can convey one or more cues associated with at least one predefined target to the participants 104A, 104B in an auditory manner (e.g., by stating "heart" out loud), a visual manner, or both, where the points of gaze of the participants are tracked relative to the participants' responses to this set of cues. In some embodiments, the responses of the participants can be received in a verbal manner (i.e., in an auditory form). In this way, a time-coded identification of the focus information 244 can be generated and used to compare with the gaze tracking measurements 242 and / or gaze patterns 243 of one or more participants 104A, 104B.
[0071] In an embodiment, the natural language processing component 232 can be a stored program component executed by at least one processor (e.g., one or more processors 202 of device 200). Those skilled in the art will understand that natural language processing refers to a branch of artificial intelligence that combines computational linguistics (i.e., rule-based modeling of human language) with statistical, machine learning, and deep learning models in order to process human language in the form of text or speech data. Similarly, the mapping component 233 can be a stored program component executed by at least one processor (e.g., one or more processors 202 of device 200). In an embodiment, the mapping component 233 includes artificial intelligence algorithms that use image processing, convolutional neural networks, and machine learning models to identify visible objects in two-dimensional or three-dimensional space.
[0072] In an embodiment, the participation analyzer 226 can further include a feedback component 234 configured to provide feedback related to the communication session 100. For example, the feedback component 234 can be configured to communicate with the host or presenter 110 of the communication session 100 and / or one or more participants 104A, 104B of the communication session 100, and provide feedback to the presenter 110 or the participants 104A, 104B. In some embodiments, the feedback can include an indication of poor participation.
[0073] In additional embodiments, the feedback component 234 can be configured to adjust one or more attributes 247 of the communication session 100 based on the determined participation level 245. For example, the feedback component 234 can modify certain attributes 247 of the display visible to the participants 104A, 104B, provide an alert to the host 110 or the participants 104A, 104B, etc. In an embodiment, adjusting the attributes 247 of the communication session 100 can include emphasizing certain information (either visually, auditorily, or a combination thereof), or repeating and / or reproducing certain information (either visually, auditorily, or a combination thereof).
[0074] In some embodiments, the feedback component 234 can provide feedback customized for the participation level 245 determined for each participant 104A, 104B. That is, in some embodiments, the first participant 104A can receive different feedback from the second participant 104B based on different evaluations of the participation levels of the participants 104A, 104B.
[0075] In certain embodiments, the feedback includes at least communicating the determined level of engagement for at least one of the participants 104A, 104B. In an embodiment, the determined level of engagement for at least one of the participants 104A, 104B can be communicated in real time or substantially in real time (with minimal latency), or can be communicated at a predefined interval (e.g., intermittently, pausing, or periodically).
[0076] The memory 204 of the device 200 can further include an operating system component 228. The operating system component 228 can be an executable program component that facilitates the operation of the device 200. Generally, the operating system component 228 is configured to facilitate access to I / O, network, and storage interfaces, and can communicate with other components of the device 200.
[0077] Now turning to Figure 6 , Figure 6 FIG. illustrates a method 600 for assessing user engagement during a communication session 100 of participants 104A, 104B related to at least one predefined goal 118A, 118B, 120A, 120B, in accordance with aspects of the present disclosure. In an embodiment, the method 600 includes: in step 610, defining a gaze focus associated with at least one predefined goal; in step 620, establishing a gaze pattern of a participant during the communication session by tracking the gaze focus relative to at least one predefined goal; in step 630, determining the level of engagement of the participant based at least in part on analyzing data associated with the gaze pattern of the participant; and in step 640, adjusting at least one attribute of the communication session based on the determined level of engagement. In an embodiment, the method 600 can further include communicating feedback (e.g., the determined level of engagement) to the participants 104A, 104B or the presenter 110, as described above.
[0078] In an embodiment, step 640 of the method 600 includes automatically adjusting at least one attribute of the communication session 100 based on the determined level of engagement. For example, in some embodiments, step 640 can include: repeating at least a portion of the communication session 100 if the level of engagement of one or more of the participants 104A, 104B drops below a predefined threshold. In additional embodiments, step 640 can include: changing the content or delivery of the communication session once the level of engagement of one or more of the participants 104A, 104B drops below a predefined threshold.
[0079] In another embodiment, method 600 includes communicating a set of cues associated with at least one predefined goal to a participant, wherein the gaze focus of the participant is tracked relative to the participant's response to the set of cues. In some embodiments, the set of cues is associated with the manipulation of at least one predefined goal. As described above, method 500 can include mapping the relationship between the set of cues and at least one predefined goal (e.g., using natural language processing and named entity recognition, etc.).
[0080] As described herein, systems 102A, 102B, gaze tracking devices 106A, 106B, and method 600 can be used in a two-dimensional or three-dimensional virtual space. That is, in some embodiments, communication session 100 can be conducted in a two-dimensional environment (e.g., on a computer monitor, tablet, phone screen, etc.), but can also be conducted in a three-dimensional environment (e.g., a virtual reality environment).
[0081] Accordingly, a virtual reality system 700 is also provided herein, the virtual reality system 700 including: (i) a head-mounted device 702 that covers at least a portion of the field of view of a participant 704; (ii) a server 720 that includes a memory for storing data (e.g., memory 204) related to defining a gaze focus associated with one or more predefined goals; and (iii) one or more processors (e.g., processor 202) that communicate with the head-mounted device 702 and the server 720, wherein the one or more processors are configured to: establish a gaze pattern (e.g., gaze pattern 243, etc.) of the participant 704 during a communication session by tracking gaze-related data (e.g., gaze focus 244, etc.) relative to at least one predefined goal (e.g., goals 118A, 118B, 120A, 120B); determine the level of participation (e.g., participation level 245) of the participant 704 at least in part based on analyzing data associated with the gaze pattern of the participant 704; and adjust at least one attribute of the communication session visible to the participant 704 based on the determined level of participation. In such an embodiment, the communication session can be presented to the participant 704 in the form of a three-dimensional virtual reality environment 708 via the head-mounted device 702.
[0082] It should be understood that all combinations of the foregoing concepts and additional concepts discussed in greater detail below (assuming these concepts are not mutually inconsistent) are contemplated as part of the inventive subject matter disclosed herein. In particular, all combinations of the claimed subject matter appearing at the end of this disclosure are contemplated as part of the inventive subject matter disclosed herein. It should also be understood that the terms explicitly employed herein, as well as those that may appear in any incorporated by reference disclosure, should be accorded a meaning most consistent with the particular concepts disclosed herein.
[0083] All definitions defined and used in this document shall be understood to control dictionary definitions, definitions in the documents incorporated by reference, and / or the ordinary meaning of the defined terms.
[0084] Unless otherwise expressly stated, the words "a" and "an" as used in this specification and the claims shall be understood to mean "at least one".
[0085] As used herein in the specification and the claims, the phrase "and / or" shall be understood to mean "either or both" of the elements so combined, i.e., elements that exist conjunctively in some cases and disjunctively in other cases. Multiple elements listed together with "and / or" shall be construed in the same manner, i.e., "one or more" of the elements so combined. Other elements may optionally exist in addition to the elements specifically identified by the "and / or" clause, whether related or unrelated to those specifically identified.
[0086] As used herein in the specification and the claims, "or" shall be understood to have the same meaning as "and / or" defined above. For example, when separating items in a list, "or" or "and / or" shall be interpreted inclusively, i.e., including at least one element, but also including more than one element or a list of elements, and optionally including additional unlisted items. Only terms that expressly indicate the contrary (e.g., "only one of... " or "exactly one of... ") or when the term "consisting of... " is used in a claim will these terms mean including exactly one element of a multiple element or list of elements. In general, the term "or" as used herein will only be interpreted as indicating an exclusive alternative (i.e., "one or the other but not both") when preceded by an exclusive term (e.g., "any one", "one of... ", "only one of... " or "exactly one of... ").
[0087] As used herein in the specification and the claims, the phrase "at least one" with respect to a list of one or more elements shall be understood to mean at least one element selected from any one or more of the elements in the list of elements, but not necessarily including at least one element of each specifically listed element within the list of elements, and not excluding any combination of the elements in the list of elements. This definition also allows for the possible optional existence of elements in addition to the specifically identified elements referred to by the phrase "at least one" within the list of elements, whether related or unrelated to those specifically identified.
[0088] Although terms such as "first", "second", "third", etc. as used herein may be used in this document to describe various elements or components, these elements or components should not be limited by these terms. These terms are only used to distinguish one element or component from another. Thus, without departing from the teachings of the inventive concept, the first element or component discussed below may also be referred to as the second element or component.
[0089] As used herein, reference numerals followed by letters ("A", "B", "C", etc.) are used to assist in identifying elements or features similar to those having the same basic reference numeral in different embodiments, while facilitating further discussion regarding specific features in individual embodiments. It should be understood that, unless otherwise specified, features or elements having a basic reference numeral with an attached letter are generally arranged and operate as described for that element or feature sharing the basic reference numeral.
[0090] Unless otherwise specified, when an element or component is referred to as "connected to", "coupled to", or "adjacent to" another element or component, it should be understood that the element or component can be directly connected or coupled to the other element or component, or there may be intermediate elements or components. That is, these terms and similar terms encompass situations where one or more intermediate elements or components may be employed to connect two elements or components. However, when an element or component is referred to as "directly connected" to another element or component, this only encompasses the situation where the two elements or components are connected to each other without any intermediate or intervening element or component.
[0091] In the claims and in the above specification, all transitional phrases such as "comprising", "including", "carrying", "having", "containing", "involving", "holding", "covering", etc. should be understood to be open-ended (i.e., meaning including but not limited to). Only the transitional phrases "consisting of" and "consisting essentially of" should be closed or semi-closed transitional phrases, respectively.
[0092] It should also be understood that, unless explicitly indicated to the contrary, in any method claimed herein that includes more than one step or act, the order of the steps or acts of the method is not necessarily limited to the order in which the steps or acts of the method are recited.
[0093] The above examples of the described subject matter can be implemented in any of a variety of ways. For example, some aspects can be implemented using hardware, software, or a combination thereof. When any aspect is implemented at least in part in software, the software code can be executed on any suitable processor or collection of processors, whether provided in a single device or computer or distributed among multiple devices / computers.
[0094] The present disclosure can be implemented as a system, a method, and / or a computer program product at any possible technical detail integration level. The computer program product can include a computer-readable storage medium (or media) having computer-readable program instructions thereon for causing a processor to execute aspects of the present disclosure.
[0095] A computer-readable storage medium can be a tangible device that is capable of storing and retaining instructions for use by an instruction execution device. A computer-readable storage medium can be, by way of example and not limitation, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of the computer-readable storage medium includes the following: a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanically encoded device (such as a punched card or raised structures in a groove having instructions recorded thereon), and any suitable combination of the foregoing. As used herein, "computer-readable storage medium" should not be construed to be a transient signal per se (such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (such as an optical pulse traveling through an optical fiber cable), or an electrical signal transmitted through a wire).
[0096] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to a corresponding computing / processing device, or downloaded to an external computer or external storage device via a network (such as the Internet, a local area network, a wide area network, and / or a wireless network). The network can include copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium within the respective computing / processing device.
[0097] The computer-readable program instructions for performing the operations of the present disclosure can be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state-setting data, configuration data for integrated circuits, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages (e.g., Smalltalk, C++, etc.) and procedural programming languages (e.g., the "C" programming language or similar programming languages). The computer-readable program instructions can be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network connection, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (e.g., using an Internet service provider through the Internet). In some examples, an electronic circuit, including, for example, a programmable logic circuit, a field-programmable gate array (FPGA), or a programmable logic array (PLA), can be personalized by utilizing the state information of the computer-readable program instructions and then execute the computer-readable program instructions to perform aspects of the present disclosure.
[0098] Aspects of the present disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to examples of the present disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0099] The computer-readable program instructions can be provided to a processor of a special-purpose computer or other programmable data processing apparatus to produce a machine, such that the instructions executed via the processor of the computer or other programmable data processing apparatus create a module for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium, which can direct a computer, a programmable data processing apparatus, and / or other devices to operate in a particular manner, such that the computer-readable storage medium in which the instructions are stored includes a manufacture, the manufacture including instructions for implementing aspects of the flowchart and / or block diagram or the functions / actions specified in the block.
[0100] The computer-readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other devices to cause a series of operational steps to be performed on the computer, other programmable apparatus, or other devices, thereby generating a computer-implemented process, and further causing functions / actions specified in one or more blocks of the flowchart and / or block diagram to be implemented on the computer, other programmable apparatus, or other devices.
[0101] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various examples of the present disclosure. In this regard, each block in the flowchart or block diagram can represent a module, segment, or portion of instructions that includes one or more executable instructions for implementing the specified (one or more) logical functions. In some alternative implementations, the functions noted in the blocks can occur in a different order than shown in the figures. For example, two blocks shown in succession can in fact be executed substantially simultaneously, or the blocks can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block of the block diagrams and / or flowchart illustrations, and combinations of blocks in the block diagrams and / or flowchart illustrations, can be implemented by a system based on dedicated hardware that performs the specified functions or actions or a combination of dedicated hardware and computer instructions.
[0102] Other implementations are within the scope of the appended claims and other claims that the applicant can enjoy.
[0103] Although several inventive embodiments have been described and illustrated herein, those of ordinary skill in the art will readily envision various other modules and / or structures for performing the functions and / or obtaining the results and / or one or more of the advantages described herein, and each of these variations and / or modifications is considered to be within the scope of the inventive embodiments described herein. More generally, those skilled in the art will readily understand that all parameters, dimensions, materials, and configurations described herein are exemplary, and the actual parameters, dimensions, materials, and / or configurations will depend on one or more specific applications of the teachings of the present invention. Those skilled in the art will recognize, or be able to determine using only routine experimentation, many equivalents to the specific inventive embodiments described herein. Accordingly, it should be understood that the foregoing embodiments are presented by way of example only, and that the inventive embodiments can be practiced in a different manner than specifically described and claimed within the scope of the appended claims and their equivalents. The inventive embodiments of the present disclosure relate to each individual feature, system, article, material, kit, and / or method described herein. Additionally, any combination of two or more such features, systems, articles, materials, kits, and / or methods (if such features, systems, articles, materials, kits, and / or methods are not mutually inconsistent) is included within the inventive scope of the present disclosure.
Claims
1. A method for evaluating user engagement during a communication session of a participant related to at least one predefined goal, the method comprising: Defining a gaze focus associated with the at least one predefined goal; Establishing a gaze pattern of the participant during the communication session by tracking the gaze focus relative to the at least one predefined goal; Determining the degree of engagement of the participant based at least in part on analyzing data associated with the gaze pattern of the participant; And Adjusting at least one attribute of the communication session based on the determined degree of engagement.
2. The method according to claim 1, further comprising: Communicating the degree of engagement to at least the participant.
3. The method according to claim 2, wherein The degree of engagement is communicated substantially in real time.
4. The method according to claim 2, wherein The degree of engagement is communicated at substantially a preset interval.
5. The method according to claim 1, wherein, The communication session is conducted in a virtual reality environment.
6. The method according to claim 1, wherein The gaze pattern of the participant during the communication session is established by: tracking the gaze focus relative to the manipulation of the at least one predefined goal.
7. The method according to claim 1, further comprising: Communicating a set of cues associated with the at least one predefined goal to the participant, wherein the gaze focus of the participant is tracked relative to the participant's response to the set of cues.
8. The method according to claim 7, wherein, The set of cues is associated with the manipulation of the at least one predefined goal.
9. The method according to claim 7, further comprising: Mapping the relationship between the set of cues and the at least one predefined goal, wherein the mapping is performed using natural language processing and named entity recognition.
10. The method according to claim 7, wherein The response of the participant includes verbal feedback.
11. The method according to claim 1, wherein The at least one attribute of the communication session is automatically adjusted based on the determined degree of engagement.
12. The method according to claim 1, wherein The at least one attribute of the communication session includes the visual aspect of the communication session, the auditory aspect of the communication session, or a combination thereof.
13. The method according to claim 1, wherein, Adjusting the at least one attribute of the communication session based on the determined degree of engagement includes: Repeating at least a part of the communication session after the degree of engagement is below a predetermined threshold.
14. The method according to claim 1, wherein Adjusting the at least one attribute of the communication session based on the determined degree of engagement includes: Changing the content or delivery of the communication session after the degree of engagement drops below a predetermined threshold.
15. A system for evaluating user engagement during a communication session of a participant related to at least one predefined goal, the system comprising: A server including a memory for storing data related to defining a gaze focus associated with the at least one predefined goal; A camera for capturing the gaze of the participant during the communication session; And One or more processors communicating with the server and the camera, wherein the one or more processors are configured to: Establish a gaze pattern of the participant during the communication session by tracking the gaze focus relative to the at least one predefined goal; Determine the degree of participation of the participant based at least in part on analyzing data associated with the gaze pattern of the participant; and Adjust at least one attribute of the communication session based on the determined degree of participation.
16. The system according to claim 15, wherein, The one or more processors are further configured to: Communicate the degree of participation to at least the participant substantially in real time.
17. The system according to claim 15, wherein, The one or more processors are further configured to: Communicate a set of cues associated with the at least one predefined goal to the participant, wherein the gaze focus of the participant is tracked relative to the participant's response to the set of cues.
18. The system according to claim 17, wherein, The set of cues is associated with the manipulation of the at least one predefined goal.
19. A virtual reality system for evaluating user participation during a communication session of a participant related to at least one predefined goal, the system comprising: A head-mounted device that covers at least a portion of the field of view; A server that includes a memory for storing data related to defining a gaze focus associated with the at least one predefined goal; and One or more processors that communicate with the head-mounted device and the server, wherein the one or more processors are configured to: Establish a gaze pattern of the participant during the communication session by tracking the gaze focus relative to the at least one predefined goal; Determine the degree of participation of the participant based at least in part on analyzing data associated with the gaze pattern of the participant; and Adjust at least one attribute of the communication session of the participant based on the determined degree of participation.
20. The virtual reality system according to claim 19, wherein The at least one attribute of the communication session is automatically adjusted based on the determined degree of participation.