Real-time estimation of user engagement levels and other factors using sensors.

A sensor-based system estimates user engagement and cognitive load in real-time, allowing for dynamic adjustments to media content, addressing the limitations of post-consumption evaluations in existing methods.

JP2026513174APending Publication Date: 2026-04-23DOLBY LABORATORIES LICENSING CORP
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
DOLBY LABORATORIES LICENSING CORP
Filing Date
2024-03-19
Publication Date
2026-04-23

AI Technical Summary

Technical Problem

Existing methods for estimating user engagement with media content do not consider real-time user reactions during consumption, relying instead on post-consumption evaluations.

Method used

Implementing a system that uses sensors such as cameras, eye trackers, and ambient light sensors to measure pupil dilation, gaze direction, and ambient illuminance to estimate user engagement, cognitive load, and arousal levels in real-time, allowing for dynamic adjustments to media content.

Benefits of technology

Enables real-time adaptation of media content based on user engagement, cognitive load, and arousal levels, enhancing user experience by providing personalized and dynamic content modifications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026513174000001_ABST
    Figure 2026513174000001_ABST
Patent Text Reader

Abstract

Methods, systems, and devices for estimating pupil dilation or constriction of a person viewing displayed media content, resulting from cognitive load, arousal, or engagement. Some embodiments include: obtaining ambient illuminance data corresponding to ambient light illuminance in the vicinity of a person; obtaining display screen luminance data associated with content luminance values ​​of the displayed media content and brightness of the display screen; obtaining instantaneous pupil size data corresponding to the size of one or more pupils of the person's pupils; estimating photo-induced pupil dilation or constriction caused by ambient illuminance and brightness of the display screen viewed by the person; and estimating pupil dilation or constriction caused by at least one of engagement, arousal, or cognitive load experienced by the person, at least in part on the pupil size data and the photo-induced pupil dilation or constriction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] [Cross - Reference to Related Applications] This application claims priority to U.S. Provisional Application No. 63 / 491,276, filed Mar. 20, 2023, which is incorporated herein by reference.

[0002] [Technical Field] The present disclosure relates to devices, systems, and methods for estimating user engagement levels and related factors based on signals from one or more sensors, and to responses to such estimated factors.

Background Art

[0003] Some methods, devices, and systems for estimating user engagement, such as user engagement with advertising content, are known. Existing devices, systems, and methods can provide benefits in some situations, but improved devices, systems, and methods are desirable.

Summary of the Invention

Means for Solving the Problems

[0004] At least some aspects of the present disclosure may be implemented via one or more methods. In some examples, the methods may be implemented, at least in part, by a control system and / or via instructions (e.g., software) stored on one or more non - transient media. Some methods include estimating dilation or constriction of the pupils of a person viewing displayed media content due to cognitive load, arousal, or engagement.

[0005] Some such methods may include, by a control system, obtaining ambient illuminance data corresponding to the illuminance of ambient light in the vicinity of a person. Some such methods may include, by a control system, obtaining instantaneous pupil size data corresponding to the size of one or more pupils of the person's pupils. According to some examples, pupil size data may be obtained from a camera or eye tracker. Some such methods may include, by a control system, at least in part, the step of estimating photo-induced pupil dilation or constriction caused by ambient illuminance and the brightness of the display screen viewed by the person, based on ambient illuminance data and display screen brightness data. Some such methods may include, by a control system, the step of estimating pupil dilation or constriction caused by at least one of engagement, arousal, or cognitive load experienced by the person, at least in part, based on pupil size data and photo-induced pupil dilation or constriction.

[0006] In some examples, estimating photo-induced pupillary dilation or constriction may involve applying a pupillary model to ambient light intensity and the brightness of the display screen being viewed by the person. According to some examples, the pupillary model may include personalized pupillary model parameters based on one or more measured responses of the person's pupils to brightness and light intensity. In some examples, determining the personalized pupillary model parameters may involve estimating photo-induced pupillary dilation or constriction according to the pupillary model, measuring instantaneous pupil size, and determining the estimation error based on the difference between the measured instantaneous pupil size and the photo-induced pupillary dilation or constriction estimated according to the pupillary model. According to some examples, the pupillary model may also be at least partially based on the cube of the cosine of the visual angle centered at the foveal position. In some examples, the pupillary model may also be at least partially based on the weighting of light wavelengths according to the relative luminous efficiency function of vision.

[0007] Some methods may involve the control system acquiring gaze direction data. In some examples, content luminance data may be based at least partially on gaze direction data.

[0008] Several methods may involve a control system estimating content time intervals corresponding to estimated pupil dilation caused by engagement or cognitive load. In some examples, the content corresponding to the content time interval may include video content, audio content, or a combination thereof. In some examples, estimating content time intervals corresponding to pupil dilation caused by engagement or cognitive load may involve applying a time shift corresponding to the pupil dilation latency period. In some examples, the time shift may be in the range of 1 to 3 seconds.

[0009] Some methods may involve estimating engagement levels, arousal levels, cognitive load levels, or combinations thereof, at least in part, based on estimated pupil dilation or constriction caused by engagement, arousal, or cognitive load. Some such methods may involve outputting analytical data based on the estimated engagement levels, estimated arousal levels, estimated cognitive load levels, or combinations thereof.

[0010] Some methods may include the step of modifying one or more aspects of media content in response to estimated pupil dilation or constriction caused by at least one of engagement, arousal, or cognitive load. In some examples, one or more aspects of media content may be modified after the estimated time of pupil dilation or constriction.

[0011] In some examples, the displayed media content may be part of a video game. Modifying one or more aspects of the media content may include changing the difficulty level of a video game, generating one or more personalized game experiences based on engagement levels, tracking player cognitive challenges for a competitive game, or a combination thereof. In some examples, generating one or more personalized game experiences based on engagement levels may include environmental modifications, aesthetic modifications, animation modifications, game mechanic modifications, or a combination thereof.

[0012] In some examples, the displayed media content may be part of an online learning course. In some examples, modifying one or more aspects of an online learning course may include changing the amount of information provided in the online learning course, changing the amount of time spent in at least part of the online learning course, changing the difficulty level of at least part of the online learning course, or a combination thereof.

[0013] In some examples, modifying one or more aspects of media content may include modifying the identifiability of a graphical object. In some examples, the graphical object may correspond to a person or a topic. In some examples, modifying the identifiability of a graphical object may include changing the camera angle, modifying the amount of time the graphical object is displayed, modifying the size at which the graphical object is displayed, or a combination thereof.

[0014] In some examples, modifying one or more aspects of media content may include modifying one or more aspects of audio content. According to some examples, modifying one or more aspects of audio content may include adaptively controlling the audio enhancement process. In some examples, modifying one or more aspects of audio content may include modifying one or more spatial properties of audio content. According to some examples, modifying one or more spatial properties of audio content may include rendering at least one audio object in a location different from the location in which at least one audio object would otherwise be rendered.

[0015] Some or all of the operations, functions, and / or methods described herein may be executed by one or more devices in accordance with instructions (e.g., software) stored on one or more non-temporary media. Such non-temporary media may include, but are not limited to, memory devices such as random-access memory (RAM) devices and read-only memory (ROM) devices, as described herein. Thus, some innovative aspects of the subject matter described herein may be implemented via one or more non-temporary media on which software is stored.

[0016] At least some aspects of this disclosure may be implemented through an apparatus. For example, one or more devices (e.g., a system comprising one or more devices) may be capable of performing at least partially the methods disclosed herein. In some implementations, the apparatus is or includes an audio processing system having an interface system and a control system. The control system may include one or more general-purpose single or multi-chip processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs) or other programmable logic devices, discrete gates or transistor logic, discrete hardware components, or a combination thereof. The control system may be configured to perform some or all of the methods disclosed herein.

[0017] Details of one or more implementations of the subject matter described herein are given in the accompanying drawings and the following description. Other features, embodiments, and advantages will become apparent from the description, drawings, and claims. Note that the relative dimensions in the following drawings may not be drawn to scale. [Brief explanation of the drawing]

[0018] Similar reference numbers and names in various drawings refer to the same elements. [Figure 1A] This block diagram shows examples of components of a device capable of implementing various aspects of this disclosure. [Figure 1B] This is a system diagram showing an environment that includes examples of system components capable of implementing various aspects of this disclosure. [Figure 2] This example shows the relationship between sensor type, sensor output, and derived user-related metrics. [Figure 3] Examples of factors that affect human pupil size are shown. [Figure 4] This shows an example of a block that could be involved in developing a pupil model that corresponds to an individual's measured pupillary response. [Figure 5] Shows exemplary blocks of a process for estimating pupil dilation or constriction based on experience. [Figure 6] It is a flowchart outlining an example of the disclosed method. **DETAILED DESCRIPTION OF THE INVENTION**

[0019] Currently, we spend a lot of time-consuming media content, including but not limited to audiovisual content, interaction with media content, or combinations thereof. (For the sake of brevity and convenience, both consuming media content and interacting with media content may be referred to herein as "consuming" media content. Consuming audiovisual content can include watching TV shows, watching movies, playing games, participating in video conferences, participating in online learning lectures, etc. Thus, movies, online games, video games, video conferences, online learning lectures, etc. may be referred to herein as types of audiovisual content. Other types of media content may include audio rather than video, such as podcasts, streamed music, etc.

[0020] Conventionally implemented techniques for estimating user engagement with media content such as movies and TV shows do not consider how a person reacts while the person is in the process of consuming the media content. Instead, a person's impression may be evaluated according to the person's evaluation of the content after the user has consumed the content, for example, after the person has finished watching a movie or an episode of a TV show, after the user has played an online game, etc.

[0021] It would be beneficial to estimate one or more states of a person while the person is in the process of consuming media content. Such states may include or may include user engagement, cognitive load, attention, interest, etc.

[0022] The various disclosed examples overcome limitations of conventionally implemented approaches for estimating user engagement. Some such examples include using one or more cameras, eye trackers, ambient light sensors, microphones, wearable sensors, or combinations thereof. Some such examples include measuring a person's level of engagement, heart rate, cognitive load, attention, interest, etc. while the person is consuming media content by watching TV, playing games, participating in a telecommunications experience (such as a video conference, video seminar, etc.), listening to a podcast, etc. According to some examples, it may include changing one or more aspects of the media content in response to the estimated engagement, arousal, cognitive load, etc.

[0023] FIG. 1A is a block diagram showing an example of components of an apparatus capable of implementing various aspects of the present disclosure. Similar to other figures provided herein, the types, numbers, and arrangements of the elements shown in FIG. 1A are provided merely as examples. Other embodiments may include more, fewer, and / or different types, numbers, and arrangements of elements. According to some examples, the apparatus 100 may be configured to execute at least some of the methods disclosed herein. In some implementations, the apparatus 100 may be or include one or more components of a workstation, one or more components of a home entertainment system, etc. For example, the apparatus 100 may be a laptop computer, a tablet device, a mobile device (such as a cellular phone), an augmented reality (AR) wearable, a virtual reality (VR) wearable, an automotive subsystem (such as an infotainment system, a driver assistance or safety system, etc.), a game system or console, a smart home hub, a television, or another type of device.

[0024] In some alternative implementations, device 100 may be a server or include a server. In some such examples, device 100 may be an encoder or include an encoder. In some examples, device 100 may be a decoder or include a decoder. Thus, in some examples, device 100 may be a device configured for use in an environment such as a home environment, while in other examples, device 100 may be a device configured for use in the "cloud," such as a server.

[0025] In some examples, the device 100 may be, or include, an orchestration device configured to provide control signals to one or more other devices. In some examples, the control signals may be provided by the orchestration device to adjust the manner of displayed video content, audio playback, or a combination thereof. In some examples, the device 100 may be configured to change one or more manners of media content currently being provided by one or more devices in the environment in response to estimated user engagement, estimated user arousal, or estimated user cognitive load. Several examples are disclosed herein.

[0026] In this example, the device 100 includes an interface system 105 and a control system 110. In some implementations, the interface system 105 may be configured to communicate with one or more other devices in the environment. In some examples, the environment may be a home environment. In other examples, the environment may be a different type of environment, such as an office environment, a car environment, a train environment, a street or sidewalk environment, a park environment, or an entertainment environment (e.g., a theater, a performance venue, a theme park, a VR experience room, an electronic game arena). In some implementations, the interface system 105 may be configured to exchange control information and related data with other devices in the environment. In some examples, the control information and related data may relate to one or more software applications that the device 100 is running.

[0027] The interface system 105 may, in some implementations, be configured to receive or provide a content stream. In some examples, the content stream may include video data and audio data corresponding to the video data. The audio data may include, but is not limited to, audio signals. In some cases, the audio data may include channel data and / or spatial data such as spatial metadata. The metadata is provided, for example, by what is referred to herein as an “encoder”.

[0028] The interface system 105 may include one or more network interfaces and / or one or more external device interfaces (such as one or more Universal Serial Bus (USB) interfaces). Depending on the implementation, the interface system 105 may include one or more wireless interfaces. The interface system 105 may include one or more devices for implementing the user interface, such as one or more microphones, one or more speakers, a display system, a touch sensor system, a gesture sensor system, or a combination thereof. Thus, although some such devices are shown separately in Figure 1A, such devices may correspond to embodiments of the interface system 105 in some examples.

[0029] In some examples, the interface system 105 may include one or more interfaces between the control system 110 and a memory system, such as the optional memory system 115 shown in Figure 1A. Alternatively or additionally, the control system 110 may include a memory system in some examples. In some implementations, the interface system 105 may be configured to receive input from one or more microphones in the environment.

[0030] The control system 110 may include, for example, a general-purpose single or multi-chip processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or a combination thereof.

[0031] In some implementations, the control system 110 may reside in two or more devices. For example, in some implementations, part of the control system 110 may reside in a device within one of the environments referred to herein, while another part of the control system 110 may reside in a device outside the environment, such as a server, a game console, or a mobile device (such as a smartphone or tablet computer). In other examples, part of the control system 110 may reside in a device within one of the environments shown herein, while another part of the control system 110 may reside in one or more other devices within the environment. For example, control system functions may be shared by an orchestration device (such as one that may be referred to herein as a smart home hub) and one or more other devices within the environment. In other examples, part of the control system 110 may reside in a device implementing cloud-based services, such as a server, while another part of the control system 110 may reside in another device implementing cloud-based services, such as another server or a memory device. The interface system 105 may also reside in two or more devices in some examples.

[0032] In some implementations, the control system 110 may be configured to perform at least partially the methods disclosed herein. According to some examples, the control system 110 may be configured to acquire ambient illuminance data corresponding to the illuminance of ambient light in the vicinity of a person. For example, the control system 110 may be configured to acquire ambient illuminance data from one or more light sensors of the device 100, or from one or more light sensors of another device near a person, such as within 0.5 meters of a person, within 1 meter of a person, or within 2 meters of a person. According to some examples, the control system 110 may be configured to acquire display screen luminance data associated with the luminance value of the displayed media content. As used herein, the term “display screen luminance” encompasses both the brightness of the display screen (e.g., the display screen brightness corresponding to the display device settings) and the relative luminance of the displayed content itself, such as the portion of the displayed content that a person is currently looking at.

[0033] In some examples, the control system 110 may be configured to acquire instantaneous pupil size data corresponding to one or more pupil sizes of a person's pupils. “Instantaneous pupil size data” may include, for example, data from one or more cameras, one or more eye-tracking devices, etc., corresponding to one or more sizes of a person's pupils. According to some examples, the control system 110 may be configured to estimate photo-induced pupil dilation or constriction caused by ambient light intensity and the brightness of the display screen viewed by the person, at least partially based on ambient light data and display screen brightness data. In some examples, the control system 110 may be configured to estimate pupil dilation or constriction caused by at least one of the engagement, arousal, or cognitive load experienced by the person, at least partially based on pupil size data and photo-induced pupil dilation or constriction. Some examples of these estimation processes are described below.

[0034] Some or all of the methods described herein may be executed by one or more devices in accordance with instructions (e.g., software) stored on one or more non-temporary media. Such non-temporary media may include, but are not limited to, memory devices such as random access memory (RAM) devices and read-only memory (ROM) devices, as described herein. One or more non-temporary media may reside, for example, in an optional memory system 115 and / or control system 110 shown in Figure 1A. Thus, various inventive aspects of the subject matter described herein may be implemented in one or more non-temporary media on which software is stored. The software may include, for example, instructions for controlling at least one device to execute some or all of the methods disclosed herein. The software may be executable by one or more components of a control system, such as the control system 110 in Figure 1A.

[0035] In some examples, the device 100 may include an optional microphone system 120, as shown in Figure 1A. The optional microphone system 120 may include one or more microphones. According to some examples, the optional microphone system 120 may include an array of microphones. In some examples, the array of microphones may be configured to determine, for example, direction of arrival (DOA) and / or time of arrival (TOA) information according to instructions from a control system 110. In some cases, the array of microphones may be configured for receiver beamforming, for example, according to instructions from a control system 110. In some implementations, one or more of the microphones may be part of or associated with another device, such as a speaker in a speaker system, a smart audio device, etc. In some examples, the device 100 may not include a microphone system 120. However, in some such implementations, the device 100 may nevertheless be configured to receive microphone data about one or more microphones in the environment via an interface system 110. In some such implementations, the cloud-based implementation of the device 100 may be configured to receive microphone data or data corresponding to microphone data from one or more microphones in the environment via the interface system 110.

[0036] In some implementations, the device 100 may include an optional loudspeaker system 125, as shown in Figure 1A. The optional loudspeaker system 125 may include one or more loudspeakers, which may also be referred to herein as “speakers” or more generally as “audio playback transducers.” In some examples (e.g., cloud-based implementations), the device 100 may not include the loudspeaker system 125.

[0037] In some embodiments, the device 100 may include an optional sensor system 130, as shown in Figure 1A. The optional sensor system 130 may include one or more touch sensors, gesture sensors, motion detectors, cameras, eye-tracking devices, or a combination thereof. In some implementations, one or more cameras may include one or more standalone cameras. In some examples, one or more cameras, eye trackers, etc., of the optional sensor system 130 may be present in a television, mobile phone, smart speaker, laptop, game console or system, or a combination thereof. In some examples, the device 100 may not include the sensor system 130. However, in some such implementations, the device 100 may nevertheless be configured to receive sensor data from one or more sensors (cameras, eye trackers, monitors with cameras, etc.) present in or on other devices in the environment via the interface system 110.

[0038] In some embodiments, the apparatus 100 may include an optional display system 135, as shown in Figure 1A. The optional display system 135 may include one or more displays, such as one or more light-emitting diode (LED) displays. In some examples, the optional display system 135 may include one or more organic light-emitting diode (OLED) displays. In some examples, the optional display system 135 may be an automotive subsystem (e.g., an infotainment system, driver assistance or safety system) or another type of device, which may include one or more displays of a television, laptop, mobile device, or smart audio device. In some examples where the apparatus 100 includes a display system 135, the sensor system 130 may include a touch sensor system and / or gesture sensor system adjacent to one or more displays of the display system 135. According to some such implementations, the control system 110 may be configured to control the display system 135 to present one or more graphical user interfaces (GUIs).

[0039] According to some such examples, device 100 may be or include a smart audio device such as a smart speaker. In some such implementations, device 100 may be or include a wake word detector. For example, device 100 may be configured to implement (at least partially) a virtual assistant.

[0040] This is a system diagram showing an environment including examples of system components that can implement various embodiments of the present disclosure. As with other diagrams provided herein, the types, number, and arrangement of elements shown in Figure 1B are provided merely as examples. Other embodiments may include more, fewer, and / or different types, numbers, and arrangements of elements. In this example, the environment 140 includes a system 145 and one or more people 150, which may also be referred to herein as “users” 150. According to this example, the system 145 includes one or more televisions (TVs) 155, one or more laptop computers 160, and one or more mobile phones (“cellphones”) 165, each of which is an example of the apparatus 100 in Figure 1A. In some examples, one or more of the TVs 155, laptop computers 160, or mobile phones 165 may include, for example, a camera deployed as a camera-embedded monitor. Depending on the particular implementation, the system 145 may be one or more components of a home entertainment system, one or more components of an office workstation, etc., or may include them.

[0041] In some examples, one or more users 150 may be using a video conferencing application while looking at (for example, sitting in front of) a laptop 160, TV 155, computer monitor screen, mobile phone 165, etc. In other examples, one or more users 150 may be watching a movie, watching a TV show, listening to music, or playing a video game.

[0042] In this example, system 145 includes an instance of sensor system 130, which is described with reference to Figure 1A. Sensor system 130 may have elements of various types, numbers, and arrangements depending on the particular implementation. In some examples, sensor system 130 may include one or more cameras directed at one or more of the users 150 and configured to acquire pupil size data relating to one or more of the users 102. In some examples, sensor system 130 may include one or more eye-tracking devices configured to track the gaze of one or more of the users 150 and, in some cases, to acquire pupil size data relating to one or more of the users 102. In some examples, sensor system 130 may include one or more ambient light sensors, which may also be referred to herein as brightness sensors. In some examples, sensor system 130 may include one or more microphones. In some examples, the sensors of sensor system 130 may be located in or on multiple locations in the environment 140, and in other examples, the sensors of sensor system 130 may be located in or on a single device in the environment 140.

[0043] The sensors in the sensor system 130 can acquire information from one or more users 150, such as gaze information (e.g., gaze direction information), pupil size information, facial expression information, posture information, information about the presence or absence of one or more users 150, the location where one or more users 150 are sitting, the number of users 150 in the environment 140, and the heart rate of one or more users 150. If the sensor system 130 includes at least one microphone, ambient noise in the environment 140 and speech from one or more users 150 can be detected. If the sensor system 130 includes at least one brightness sensor, the luminosity of the environment 140 can be measured.

[0044] In this example, the state estimation module 175 is configured to estimate one or more user states, environmental states, etc., based at least in part on sensor data from the sensor system 130. In some examples, an optional sensor fusion module 170 may be configured to process and analyze the sensor data acquired by the sensor system 130 using one or more sensor fusion algorithms, if present, and provide the sensor fusion data to the state estimation module 175. For example, a simple linear regression algorithm can be used to map the input sensor data to levels of user engagement and state. More complex machine learning models, such as support vector machines, can be used for the same purpose. Multilayer perceptrons and deep neural networks can also be used. In some examples, the user states estimated by the state estimation module 175 may include one or more metrics describing the user's mental state, emotional state, physiological state, or a combination thereof. For example, the state estimation module 175 may estimate the user's emotions, feelings, engagement, preferences, focus of attention, heart rate, or a combination thereof. In some examples, the state estimation module 175 may determine or estimate the presence of users in the environment 140, identify users, estimate the location of users in the environment 140, estimate the number of users in the environment 140, or a combination thereof. According to some examples, the state estimation module 175 may estimate environmental conditions such as ambient light, noise level, background chatter (utterances by one or more users 150), or a combination thereof. In some examples, the state estimation module 175 may determine or estimate, at least in part, how many people are present in the environment 140, their IDs, how far they are from a display screen presenting content, or a combination thereof, based on microphone signals.

[0045] In this example, the state estimation module 175 (and, if present, the sensor fusion module 170) is implemented by an instance of the control system 110 described with reference to Figure 1A. In some examples, the control system 110 may reside within a device in the environment 140. As described in the description of Figure 1A, in some implementations, the control system 110 may reside within two or more devices. For example, in some implementations, part of the control system 110 may reside within a device in the environment 140, and another part of the control system 110 may reside within a device outside the environment, such as a server, a game console, and a mobile device (e.g., a smartphone or tablet computer). In other examples, part of the control system 110 may reside within a device in the environment 140, and another part of the control system 110 may reside within one or more other devices in the environment 140. For example, control system functions may be shared by an orchestration device (such as what may be referred to herein as a smart home hub) and one or more other devices in the environment. In other examples, at least part of the control system 110 may reside on a device implementing cloud-based services, such as a server, and another part of the control system 110 may reside on another device implementing cloud-based services, such as another server or memory device.

[0046] The state estimation of the state estimation module 175 can be used in different ways depending on the specific implementation of the system 145. In some examples, the state estimation of the state estimation module 175 can be used to dynamically modify and adapt the user experience, as indicated by arrow 180. In one such example, a game can be made relatively more or relatively less challenging based on the user's state, such as the user's emotional response. In another example, a movie or television program can be dynamically modified based on the user's state, for example, in an attempt to induce a higher level of engagement. Additional examples are provided below.

[0047] Alternatively, or additionally, the state estimates from state estimation module 175 may be used to generate analysis 185. For example, analysis 185 may be generated during a marketing call to quantify user engagement or general sentiment (positive or negative) during the call. In some examples, the state estimates from state estimation module 175 may be used to provide information for understanding the overall outcome of the call. Additional examples are provided below.

[0048] Figure 2 shows an example of the relationship between sensor type, sensor output, and derived user-related metrics. As with other figures provided herein, the type, number, and arrangement of elements shown in Figure 2 are provided for illustrative purposes only. Other embodiments may include more, fewer, and / or different types, numbers, and arrangements of elements.

[0049] In this example, system 145 includes an instance of sensor system 130, which is described with reference to Figure 1A. Sensor system 130 shown in Figure 2 may be present in, for example, environment 140 in Figure 1B or a similar environment. In this example, sensor system 130 includes an eye tracker 205, a camera 210, a microphone 215, and an ambient light sensor 220. In some examples, a single device (such as a TV, computer monitor, laptop computer, mobile phone, or game console) may comprise the entire sensor system 130 shown in Figure 2. Alternatively, or additionally, one or more other devices present in the same environment may comprise an instance of sensor system 130.

[0050] In this example, the eye tracker 205 is configured to collect gaze and pupil size information. In some examples, the eye tracker 205 may also be configured to provide information about the presence or absence of one or more users in at least a portion of the environment, such as in front of the display screen currently showing media content. In some examples, the eye tracker 205 may be configured to provide information about the distance of one or more users from the display screen currently showing media content. In some examples, the state estimation module 175 in Figure 1B may use gaze information to quantify the user's attention focus and user preferences (estimated by, for example, the amount of time the user spends focusing on the displayed media content, the amount of time the user spends focusing on a particular object, person, or area of ​​interest in the displayed media content, etc.). Gaze can also be used as an input system. For example, gaze information may be used to interact with one or more buttons in a graphical user interface (GUI), or to control the camera view in a game (for example, by reorienting a virtual camera).

[0051] Pupil size data can be acquired via the eye tracker 205, by the camera 210, or a combination thereof. Pupil size can be used to quantify factors such as user cognitive load and user engagement. Several examples of estimating factors such as user engagement and cognitive load from pupil size, pupil dilation, and pupil constriction are described below.

[0052] Camera feeds (which may include still images, videos, or a combination thereof) acquired from camera 210 can be used to quantify the user's presence, emotions, body posture, etc. In some examples, the state estimation module 175 estimates the user's heart rate based on a video of the user's face, a relevant method disclosed in S. Sanyal and K. Nundy, Algorithm for Monitoring Heart Rate and Respiratory Rate from Video of a User's Face (IEEE Journal of Translational Engineering in Health and Medicine 2018;6:2700111), which is incorporated herein by reference. In other examples, heart rate may be estimated by a wearable device such as a wristwatch, fitness tracking device, or another type of wearable health monitoring device.

[0053] In some examples, the state estimation module 175 may be configured to quantify, based on the camera feed, how many people are in front of the camera and their positions in the environment. In some examples, the state estimation module 175 may be configured to identify one or more users according to the camera feed, for example, by implementing a facial recognition algorithm. User identification is related to providing a personalized and tailored media consumption experience.

[0054] The camera feed may also be used to estimate pupil size and gaze direction. Therefore, it is not essential that the sensor system 130 includes an eye tracker. In some examples, the state estimation module 175 may be configured to estimate user distance, user engagement, cognitive load, or a combination thereof, according to pupil size and gaze direction information obtained from the camera feed.

[0055] The microphone signal from microphone 215 can be used to measure the noise level in the environment, the presence or absence of user utterances such as background chatter, etc. In some examples, the state estimation module 175 may be configured to enhance the audio experience based on such audio information, for example, by increasing or decreasing the volume of audio corresponding to media content presented in the environment according to ambient noise in the environment, or by increasing or decreasing the volume of dialogue in a media content stream. The voice information recorded by the microphone can also be used as input to quantify the user's state and engagement. For example, the volume and emotion of the voice inferred from the voice information can be used to understand the user's response to the experience, stress, and engagement levels.

[0056] The ambient light sensor 220 may be used to measure the illuminance of ambient light near one or more people in the environment. Such ambient light data may be used, for example, to establish a “baseline” pupil size corresponding to the ambient light illuminance when a person is not looking at or reacting to displayed content. Alternatively or additionally, such ambient light data may be used (for example, by the state estimation module 175 in Figure 1B) to compensate for pupil dilation or constriction caused by changes in ambient light illuminance.

[0057] Figure 3 shows examples of factors that influence human pupil size. The examples shown in Figure 3 relate to factors that directly contribute to pupil dilation and constriction in humans viewing content on a screen. Screen brightness and content brightness are components of display screen brightness (also referred to herein as “screen brightness”) and are the primary contributors to pupil size response. Ambient illumination also affects pupil size. Engagement, cognitive load, and arousal are other factors that influence pupil dilation and constriction. Factors such as engagement, cognitive load, and arousal may be referred herein to as the user’s “experience-based physiological response,” or simply the user’s “experience-based response.” It is an assumption underlying some disclosed implementations that, during the time a user is consuming media content, the emotional response underlying the user’s experience-based physiological response corresponds, primarily or entirely, to the user’s emotional response to the media content being consumed.

[0058] According to some disclosed examples, pupillary response can be considered the sum of three main factors: screen brightness, ambient light, and the user's physiological response to the media content being consumed. Some aspects of this disclosure include separating pupillary response caused by ambient light and screen brightness from pupillary response caused by the user's physiological response based on user experience.

[0059] The extent to which screen brightness, ambient light, and experience-based physiological responses contribute to pupil dilation and constriction can vary substantially from person to person. Furthermore, the extent to which caffeine, alcohol, or other substance consumption, as well as fatigue, affect pupil dilation or constriction (assuming constant stimuli) can also vary substantially from person to person. Factors such as caffeine intake are referred to here as "user personal state." User personal state factors may be considered to establish a baseline for an individual's pupil size in relation to the time and place of media content consumption.

[0060] In some examples, it is assumed that the user's personal state factors and ambient light remain constant throughout the duration of the user's session consuming media content. In other words, such examples include the assumption that ambient light and the user's personal state factors remain constant during video conferencing sessions, movie viewing sessions, or gaming experiences.

[0061] According to some disclosed examples, if ambient light and the user's personal state factors remain constant during a user's session consuming media content, it is assumed that a person's pupillary response is based solely on changes in display screen brightness and experience-based pupillary response. Accordingly, some disclosed examples include estimating experience-based pupillary response by subtracting a pupillary response estimated based on changes in display screen brightness. Some such examples include obtaining data on an individual person's personalized pupillary response based on changes in display screen brightness, for example, through a process described herein with reference to Figure 4, or by a similar process.

[0062] In other implementations, it is not assumed that ambient light is constant during a user's session or experience consuming media content. According to some such implementations, ambient light data from an ambient light sensor may be used (for example, by the state estimation module 175 in Figure 1B) to compensate for pupil dilation or constriction caused by changes in ambient light intensity. Some such examples include obtaining data on an individual's personalized pupillary response based on changes in ambient light intensity and using this data to compensate for pupil dilation or constriction caused by changes in ambient light intensity.

[0063] Some examples disclosed involve developing what is referred to herein as a “pupil model,” which may be a model of the dilation of a particular person’s light-induced pupil, or, for example, a model of the dilation or constriction of a particular person’s light-induced pupil caused by the brightness of a display screen that the person is looking at, the illuminance of ambient light, or a combination thereof.

[0064] Figure 4 shows an example of blocks that may be involved in developing a pupil model corresponding to an individual's measured pupillary response. The pupil model may include personalized pupil model parameters based on one or more measured responses of the person's pupil to brightness, one or more measured responses of the person's pupil to illuminance, or a combination thereof. The pupil model may be based on current pupil size measurements, sometimes referred to herein as “instantaneous” pupil size, in response to the brightness of the display screen being viewed by the person, in response to the illuminance of ambient light, or a combination thereof.

[0065] In the example shown in Figure 4, an instance of the control system 110 in Figure 1 is configured to implement a display screen brightness estimation module 410 and a pupil model determination module 415. In this example, the pupil model determination module 415 is configured to determine personalized pupil model parameters 417. In this example, determining personalized pupil model parameters 417 includes an iterative process of estimating photo-induced pupil dilation or constriction according to the current pupil model parameters, obtaining instantaneous pupil size measurements, determining an estimation error based on the difference between the measured instantaneous pupil size and the photo-induced pupil dilation or constriction estimated according to the pupil model parameters, and adjusting the current pupil model parameters according to the estimation error.

[0066] In the example shown in Figure 4, the ambient light sensor 220 (also referred to herein as a brightness sensor) is configured to determine the ambient light intensity in the vicinity of the person whose pupil response is being evaluated (e.g., within 0.5 meters of the person, within 1 meter of the person, within 2 meters of the person, etc.) and to provide the ambient light intensity data to the pupil model determination module 415.

[0067] The appropriate selection of calibration content 405 is crucial for obtaining optimal pupil model parameters. In some cases, calibration content 405 may be identified to allow stimuli ranging from black to white, in other words, from the minimum possible content brightness to the maximum possible content brightness.

[0068] In this example, pupil size measurement is performed using an eye tracker 205 configured to provide instantaneous pupil size data to a control system 110. In some alternative examples, instantaneous pupil size measurement may be performed using a camera instead of, or in addition to, the eye tracker 205. In this example, the eye tracker 205 also determines the direction of the person's gaze and provides gaze information to a display screen brightness estimation module 410. In some examples, gaze information indicates at least whether the person is looking at the display screen on which the calibration content 405 is presented. In some examples, gaze information may also indicate which part of the display screen the person is looking at. As described elsewhere in this specification, the display screen brightness data may correspond to both the content brightness of the displayed media content, in this example, the displayed calibration content 405, and the brightness of the display screen on which the calibration content 405 is presented.

[0069] Therefore, in some implementations, the display screen brightness estimation module 410 may be configured to determine the content brightness of calibration content displayed in a specific area of ​​the display screen based on gaze information indicating which part of the display screen a person is looking at, and based on calibration content information. In some examples, the brightness of the display screen may correspond to the brightness setting of the display screen. In some examples, the display screen brightness estimation module 410 is configured to determine display screen brightness data 412 based on the content brightness value of the displayed calibration content and based on the brightness of the display screen.

[0070] According to some examples, the display screen brightness estimation module 410 may be configured to determine the display screen brightness data 412 based at least partially on the "Y channel" portion of the calibration content information. The "Y channel" is related to the YUV color model, which takes human perception into account. The Y channel correlates approximately with perceived intensity, while the U and V channels provide color information. Regardless of how it is determined, the display screen brightness estimation module 410 is configured to provide the display screen brightness data 412 to the pupil model determination module 415.

[0071] The pupil model determination module 415 may use different models to establish the relationship between input luminosity and output pupil size, depending on the specific implementation. Some such models are partially based on the assumption that ambient light generally has less influence than display screen brightness. This is generally true, for example, when viewing in a dimly lit room. The display's field of view (FOV) is an important factor. In one simple example, the pupil model determination module 415 may apply a model that assumes pupil size is controlled by specific regions around the fovea (e.g., a 4-degree region, a 6-degree region, an 8-degree region, a 10-degree region, etc.). The foveal position may be determined, for example, according to data from the eye tracker 205. According to some such examples, the model may also be based on the assumption that pupil size is controlled only by specific components, such as the green component of visible light. In other examples, the pupil model determination module 415 may apply a more nuanced model. Some such models involve using the cube of the cosine of the visual angle centered on the foveal position. Some such models may also include weighting wavelengths by a relative luminous efficiency function of vision, such as the photopic luminous efficiency function established by the International Commission on Illumination (CIE). Some such models may take into account both the effect of display screen brightness and the effect of ambient illumination, such as light reflected from the walls surrounding the display. In some examples, the pupil model determination module 415 may apply relatively more sophisticated models, such as models based on neural networks, such as recurrent neural networks (RNNs), which are trained on datasets of large pupil sizes in response to display screen brightness, ambient illumination, or both. According to some such examples, at least a part of the control system 110, such as the pupil model determination module 415, may be configured to implement a neural network. In some such examples, determining the error metric may include applying a loss function used to train the neural network. According to some such examples, the estimated error may be, or correspond to, the loss function gradient determined by applying the loss function.

[0072] Figure 5 shows an exemplary block of the method for the process of estimating experience-based pupil dilation or constriction according to the present invention. As described elsewhere in this specification, “experience-based pupil dilation or constriction” and “pupil dilation or constriction” refer to pupil dilation or constriction that a person experiences in response to an experience, such as an experience of consuming media content, which is caused by engagement, arousal, cognitive load, etc.

[0073] In the example shown in Figure 5, an instance of the control system 110 in Figure 1 is configured to implement a display screen brightness estimation module 410 and a retinal illumination estimation module 515. In this example, the control system 110 is configured to estimate photo-induced pupil dilation or constriction by applying a pupil model 510 that is at least partially based on personalized pupil model parameters 417. In some examples, the personalized pupil model parameters 417 may be predetermined according to one or more of the processes described with reference to Figure 4, or a similar process.

[0074] In the example shown in Figure 5, the ambient light sensor 220 (also referred to herein as a brightness sensor) is configured to determine the ambient light intensity in the vicinity of a person whose pupillary response is being evaluated (e.g., within 0.5 meters of the person, within 1 meter of the person, within 2 meters of the person, etc.) and to provide ambient light intensity data to the retinal illuminance estimation module 515.

[0075] In this example, pupil size measurement is performed using an eye tracker 205 configured to provide instantaneous pupil size data to a control system 110. In some alternative examples, instantaneous pupil size measurement may be performed using a camera instead of, or in addition to, the eye tracker 205. In this example, the eye tracker 205 also determines the direction of the person's gaze and provides the gaze information to a display screen brightness estimation module 410.

[0076] In this example, the display screen brightness estimation module 410 is configured to determine the display screen brightness data 412 according to one or more of the methods described with reference to Figure 4. One minor difference is that instead of receiving calibration content information corresponding to calibration content 405 as input, the display screen brightness estimation module 410 receives media content information corresponding to the audiovisual media content 505 currently being viewed by the person whose pupil size is being evaluated.

[0077] In some examples, the control system 110 may be configured to estimate empirically-based pupillary dilation or constriction by subtracting the estimated photo-evoked pupillary dilation or constriction, estimated by the control system 110 according to the pupil model 510 and personalized pupil model parameters 417, from the pupillary dilation or constriction corresponding to instantaneous pupillary dilation or constriction, such as measured by the eye tracker 205 or measured by the camera. In some examples, the control system 110 may be configured to estimate empirically-based pupillary dilation or constriction by subtracting the estimated photo-evoked pupillary size, estimated by the control system according to the pupil model and personalized pupil model parameters, from instantaneous pupillary dilation or constriction, such as measured by the eye tracker 205 or measured by the camera. This is valid if we assume that the overall pupillary dilation (or constriction) is the result of adding (or subtracting) the change in empirically-based pupillary dilation (or constriction) and the change in photo-evoked pupillary dilation (or constriction). The same reasoning can be applied if we assume that two factors (photocatalytic pupillary change and experience-based pupillary change) contribute to the overall pupillary size change through the multiplication of their effects. In that case, experience-based pupillary dilation (or constriction) can be obtained by dividing the overall pupillary size change by the estimated photocatalytic pupillary change.

[0078] As indicated by the “Time Synchronization” arrow in Figure 5, there is generally a time lag between the displayed content time interval that triggers experience-based pupil dilation or constriction and the actual experience-based pupil dilation or constriction itself. This time may be referred to herein as the “pupil dilation latency period” and in some examples may range from 1 to 3 seconds. Thus, some disclosed examples may involve estimating at least the start of the content time interval corresponding to the pupil dilation triggered by engagement or cognitive load by applying a time shift corresponding to the pupil dilation latency period. For example, if experience-based pupil dilation or constriction is detected to have started at time T, some such examples may include subtracting a time shift ranging from 1 to 3 seconds from time T to estimate the start of the time interval in which the displayed content produced experience-based pupil dilation or constriction. In some examples, the end of the time interval in which the displayed content produced experience-based pupil dilation or constriction may be determined by subtracting a time shift from the time in which experience-based pupil dilation or constriction ceases to occur.

[0079] Information regarding pupil dilation or constriction based on estimated experience can be used for many different purposes, depending on the specific implementation. Some examples are listed below.

[0080] Real-time analysis Some examples include generating analytics about one or more users having media content-based experiences such as watching movies, watching TV shows, video conferencing, or playing games. Examples of analyses that may be generated include estimates of emotions / mood, attention, focus of attention, engagement, gaze, and heart rate. Some examples include generating analytics about the presence or absence of users, the number of users participating in the experience, user identification information, user location, and whether or not there is background chatter.

[0081] In some examples, analysis can be generated for a single individual and provided (e.g., displayed) in real time. Alternatively, or additionally, analysis can be provided at the end of a single individual's media consumption experience. In some examples, analysis can represent cumulative information, such as information about multiple instances of media consumption by a single individual.

[0082] Alternatively, or additionally, analysis may be generated for groups of two or more people, for example, when two or more people consume the same media content. In some examples, the analysis data may be anonymized, resulting in only group analysis information being provided. Anonymized group analysis may be appropriate for, for example, a person's presentation to a large group, or for a musician's live streaming to a crowd. Other appropriate examples may include generating post-conference analysis for sales or marketing applications.

[0083] video conferencing In the case of video conferencing experiences, analysis can be used to improve the experience itself, for example, by following a feedback loop. For example, eye-tracking information can be used to modify the user interface (UI) and increase engagement. In a particular video conferencing experience, analysis can enable participants (such as presenters) to understand the level of engagement of individual participants, all participants in a group, or combinations thereof. For example, the UI, presented materials, graphics, animations, etc., can be modified to encourage participants with low engagement levels to pay attention to the presented materials and participate in the discussion. In other embodiments, one or more aspects of a video conference (such as windows corresponding to participants' videos) can be modified to reinforce and highlight people with relatively high levels of engagement.

[0084] game In some examples, analysis can be generated while one or more people are playing a video game. In some game-related examples, analysis can quantify player performance, player engagement levels, player frustration levels, player enjoyment levels, and player behavior. In some examples, analysis can be used for user training, improving user game performance, etc. According to some examples, game difficulty and design can be dynamically adjusted based on analysis.

[0085] Dialogue Improvement In some examples, the analysis can be used to quantify the level of user challenge in understanding dialogue. Some such examples may involve adaptively controlling one or more dialogue enhancement features. Dialogue enhancement may include, for example, boosting the level of speech audio separately from the level of background content without increasing the overall loudness of the audio scene. According to some examples, the analysis related to dialogue enhancement may include emotion / affect quantification, cognitive load estimated from pupil size, posture, facial expressions, or a combination thereof.

[0086] Personalized video content for TV and streaming By using sensors to understand the viewer's experience in real time, there is untapped potential to personalize and customize the viewer's experience to suit their interests, preferences, and mood. In some examples, video content could be automatically edited and enhanced to provide the best experience for a particular person at a particular time. Some such examples may include using content metadata to specify options and intents.

[0087] In the case of sports content, relevant analytics regarding viewer attention and engagement may include which players viewers are most interested in, which camera angles they are most interested in, or both. Using multiple camera feeds, such analytics can inform switching / editing algorithms to show users more of the players they are interested in. Even with a single camera feed, cropping can be personalized using intelligent zoom or dynamic reframing techniques. In some cases, custom overlays may be created and presented on the video content to show statistics or graphics highlighting players of interest. According to some examples, existing cutscenes and graphics may also be customized in this way.

[0088] Similar methods may be applied to concert streams or pre-recorded concerts, such as switching between camera angles, zooming in on specific musicians of interest to the user, personalizing the audio mix to focus on a particular vocalist or instrumentalist of interest, or a combination of these.

[0089] There are many reality television programs with numerous actors or participants, and these programs often broadcast only a small fraction of the interviews and filmed content. By measuring user attention and engagement in accordance with analytics such as those disclosed herein, intelligent editing systems can track which characters, actors, or participants users are most interested in, and then edit the content to show more of those interviews or content, including in some cases interviews or content that are not presented to most other viewers or listeners.

[0090] Online learning As several examples show, one or more types of analysis disclosed can be applied either in real time or after the experience to improve the online learning experience. Online learning suffers from inherent challenges that make it difficult for teachers to assess student engagement and understanding, and difficult for students to stay focused and engaged. However, online courses also offer unprecedented opportunities to personalize the educational experience. Because the content is virtual and typically experienced in isolation, there is great potential to customize the content to maximize the learning benefit for individual students.

[0091] For example, estimated cognitive load can be used to categorize how much effort an individual is putting into retaining or learning something. In some such cases, such analysis can be used to provide a more personalized learning experience. High cognitive load in the context of the learning process is not necessarily a negative factor. Some students seek out tasks that cause high cognitive load, while others may quickly lose motivation and even decide not to complete the instructed course. By measuring and categorizing different types of cognitive load signals, insights can be provided into whether learners are motivated, unengaged, not retaining information, losing motivation, or being easily distracted.

[0092] Some implementations may include measuring complementary signals from the webcam, such as eye gaze, posture, and facial expressions (e.g., frown lines, scowls), which can provide information about the student's experience. In some examples, UI metrics such as mouse movements and button click timing can extend to other types of analysis.

[0093] In some examples, analysis can be aggregated and delivered to teachers or presenters in real time, enabling them to have a more interactive understanding of their audience's engagement, comprehension, and other factors. Such examples allow for more interactive teaching and presentations.

[0094] In some cases, analytics can be provided as feedback to educational software platforms used for online learning. In some cases, analytics can be used to intelligently personalize content to optimize engagement, retention, and conceptual understanding. In some cases, such modifications can be near real-time, for example, increasing or decreasing the time spent on a particular topic, slowing or accelerating lectures, or providing additional context or information on a particular topic. In some cases, different balances of learning modules, difficulty levels, or content length can be suggested based on analytical data, such as analytical data collected across multiple sessions.

[0095] According to some examples, at least some of the collected metrics can be presented directly to the user on a regular basis. Such metrics can, for example, provide students with more insight into their learning patterns and guide them in a positive way. Some students enjoy having data on their learning progress and can use it for motivation or positive reinforcement.

[0096] Figure 6 is a flowchart outlining an example of the disclosed method. The blocks of Method 600, as with other methods described herein, are not necessarily performed in the order shown. According to some examples, one or more blocks may be performed in parallel. Furthermore, some similar methods may include more or fewer blocks than those illustrated and / or described. In this example, Method 600 includes estimating pupil dilation or constriction of a person viewing displayed media content, resulting from cognitive load, arousal, engagement, or a combination thereof.

[0097] Method 600 may be implemented by an apparatus or system such as the apparatus 100 shown in Figure 1A and described above. In some examples, apparatus 100 includes at least the control system 110 shown in Figure 5 and described above. In some examples, blocks of Method 600 may be implemented by one or more devices in an audio environment, for example, by an audio system controller (which may be referred to herein as a smart home hub), or by another component of the audio system such as a television, a television control module, a laptop computer, a game console or system, or a mobile device (such as a cellular phone). However, in some implementations, at least some blocks of Method 600 may be implemented by one or more devices configured to implement a cloud-based service, such as one or more servers.

[0098] In this example, block 605 includes the control system acquiring ambient illuminance data corresponding to the illuminance of ambient light near a person. In some examples, the control system may acquire ambient illuminance data from an optical sensor, such as the ambient light sensor 220 disclosed herein.

[0099] In this example, block 610 includes obtaining display screen brightness data, which is associated with the content brightness value of the displayed media content and with the brightness of the display screen, via the control system. In some examples, block 610 may include receiving display screen brightness data from the display screen brightness estimation module 410 in Figure 5.

[0100] In this example, block 615 includes obtaining instantaneous pupil size data corresponding to the pupil sizes of one or more of a person's pupils via a control system. According to some examples, block 615 may include receiving pupil size data from the eye tracker 205 in Figure 5. In other examples, block 615 may include receiving pupil size data from a camera.

[0101] In this example, block 620 includes, by a control system and at least in part, estimating photo-induced pupil dilation or constriction caused by ambient light illuminance and the luminance of the display screen viewed by a person, based on ambient light data and display screen luminance data. In some examples, block 620 may include applying a pupil model and personalized pupil model parameters, as illustrated with reference to Figure 5. Personalized pupil model parameters may, in some examples, be based on one or more measured responses of a person's pupil to luminance and illuminance, as illustrated with reference to Figure 4. Determining personalized pupil model parameters may include estimating photo-induced pupil dilation or constriction according to the pupil model, measuring instantaneous pupil size, and determining the estimation error based on the difference between the measured instantaneous pupil size and the photo-induced pupil dilation or constriction estimated according to the pupil model (e.g., as described above with reference to Figure 4). In some examples, the pupil model may be at least in part based on the cube of the cosine of the visual angle centered on the foveal position. In some examples, the pupil model may be at least partially based on a weighting of light wavelengths according to a relative luminous efficiency function of vision. In some examples, block 620 may include estimating light-induced pupillary dilation or constriction by applying the pupil model to ambient light illuminance and the brightness of the display screen being viewed by the person in order to generate an estimate of light-induced pupillary dilation or constriction.

[0102] In this example, block 625 includes the control system estimating pupil dilation or constriction caused by at least one of the engagement, arousal, or cognitive load experienced by a person, based at least in part on pupil size data and photo-evoked pupil dilation or constriction. In some examples, block 625 may include estimating experience-based pupil dilation or constriction by subtracting the estimated photo-evoked pupil dilation or constriction from the pupil dilation or constriction corresponding to instantaneous pupil size, such as measured by the eye tracker 205 or the camera. In some examples, block 625 may include estimating experience-based pupil dilation or constriction by subtracting the photo-evoked pupil size estimated from instantaneous pupil size, such as measured by the eye tracker 205 or the camera.

[0103] In some examples, method 600 may include acquiring gaze direction data by a control system. In some such examples, content luminance data may be based at least in part on gaze direction data.

[0104] In some examples, Method 600 may include a control system estimating content time intervals corresponding to estimated pupil dilation caused by engagement or cognitive load. The content corresponding to the content time interval may include video content, audio content, or a combination thereof. In some examples, estimating content time intervals corresponding to pupil dilation caused by engagement or cognitive load may include applying a time shift corresponding to the pupil dilation latency period. In some examples, the time shift may be in the range of 1 to 3 seconds. In other examples, the time shift may be a longer or shorter time interval. In some examples, the pupil dilation latency period of a person may be predetermined as part of a calibration process, for example, as described with reference to Figure 4.

[0105] In some examples, Method 600 may include estimating engagement levels, arousal levels, cognitive load levels, or combinations thereof, at least in part on estimated pupil dilation or constriction caused by engagement, arousal, or cognitive load. According to some examples, a person's actual engagement levels, arousal levels, cognitive load levels, or combinations thereof may be predetermined as part of a calibration process that includes, for example, obtaining feedback from the person regarding the actual levels of engagement, arousal, cognitive load, or combinations thereof, and may be pre-correlated with estimated pupil dilation or constriction caused by engagement, arousal, or cognitive load. According to some examples, Method 600 may include outputting analytical data based on estimated engagement levels, estimated arousal levels, estimated cognitive load levels, or combinations thereof.

[0106] In some examples, method 600 may include modifying one or more aspects of media content in response to estimated pupil dilation or constriction caused by at least one of engagement, arousal, or cognitive load. In some examples, one or more aspects of media content may be modified after a time in which pupil dilation or constriction is estimated.

[0107] In some examples, the displayed media content may be part of a video game. Modifying one or more aspects of the media content may include changing the difficulty level of a video game, generating one or more personalized game experiences (e.g., based on engagement levels), tracking player cognitive challenges for a competitive game, or a combination thereof. In some examples, generating one or more personalized game experiences may include environmental modifications, aesthetic modifications, animation modifications, game mechanic modifications, or a combination thereof.

[0108] In some examples, whether in a game context or another context (such as online learning, video conferencing, or movie viewing), aesthetic modifications may involve modifying the visual properties (e.g., contrast, sharpness, average brightness level, color, hue, tone, saturation, brightness, transparency, or other visual properties) of displayed graphical elements (e.g., displayed graphical elements in a game). In some examples, aesthetic modifications may involve adding or removing graphical elements from multiple displayed graphical elements. In some examples, aesthetic modifications may involve modifying the acoustic properties of audio associated with displayed graphical elements (e.g., changing character voices / accents / languages, changing volume, changing background music, changing notification sounds, such as notification sounds associated with displayed chat messages).

[0109] According to some examples, the displayed media content may be part of an online learning course. Modifying one or more aspects of the media content may include modifying one or more aspects of the online learning course. For example, modifying one or more aspects of the online learning course may include modifying the amount of information provided in the online learning course, modifying the amount of time spent on at least one part of the online learning course, modifying the difficulty level of at least one part of the online learning course, or a combination of these.

[0110] In some examples, altering one or more aspects of media content may include modifying the identifiability of a graphical object. A graphical object may, for example, correspond to a person or a topic. Altering the identifiability of a graphical object may include, for example, changing the camera angle, changing the field of view, changing the amount of time the graphical object is displayed, changing the size at which the graphical object is displayed, or a combination thereof.

[0111] In some examples, modifying one or more aspects of media content may include modifying one or more aspects of audio content. In some examples, modifying one or more aspects of audio content may include adaptively controlling audio enhancement processes, such as dialogue enhancement processes. In some examples, modifying one or more aspects of audio content may include modifying one or more spatial properties of audio content. In some examples, modifying one or more spatial properties of audio content may include rendering at least one audio object in a location different from where it would otherwise be rendered.

[0112] Some aspects of this disclosure include a system or device configured (e.g., programmed) to perform one or more examples of the disclosed method, and a tangible computer-readable medium (e.g., a disk) for storing code for performing one or more examples of the disclosed method or its steps. For example, some disclosed systems are or include a programmable general-purpose processor, digital signal processor, or microprocessor, programmed in software or firmware and / or configured to perform any of a variety of operations on data, including embodiments of the disclosed method or its steps. Such a general-purpose processor may be or include a computer system including an input device, memory, and a processing subsystem programmed (and / or otherwise configured) to perform one or more examples of the disclosed method (or its steps) in response to asserted data.

[0113] Some embodiments may be implemented as a configurable (e.g., programmable) digital signal processor (DSP) configured (e.g., programmed and otherwise configured) to perform necessary processing on an audio signal(s), including the execution of one or more examples of the disclosed methods. Alternatively, embodiments of the disclosed system (or its elements) may be implemented as a general-purpose processor (e.g., a personal computer (PC) or other computer system or microprocessor, which may include input devices and memory) programmed in software or firmware and / or otherwise configured to perform any of a variety of operations, including one or more examples of the disclosed methods. Or, elements of some embodiments of the system of the present invention may be implemented as a general-purpose processor or DSP configured (e.g., programmed) to execute one or more examples of the disclosed methods, and the system also includes other elements (e.g., one or more loudspeakers and / or one or more microphones). The general-purpose processor configured to execute one or more examples of the disclosed methods may be coupled to input devices (e.g., a mouse and / or keyboard), memory, and display devices.

[0114] Another aspect of the Disclosure is a computer-readable medium (e.g., a disk or other tangible storage medium) that stores code (e.g., a coder operable to execute) for performing one or more examples of the disclosed method or steps thereof.

[0115] While specific embodiments and applications of the Disclosure are described herein, it will be apparent to those skilled in the art that many modifications are possible to the embodiments and applications described herein without departing from the scope of the Disclosure described herein and claimed herein. While specific forms of the Disclosure are shown and described, it should be understood that the Disclosure should not be limited to the specific embodiments or methods described herein.

[0116] Various aspects of this disclosure can be understood from the following listed exemplary embodiments (EEE).

[0117] EEE1. A method for estimating pupil dilation or constriction in a person viewing displayed media content, due to cognitive load, arousal, or engagement, The control system acquires ambient illuminance data corresponding to the illuminance of the surrounding light near a person, The control system acquires the content brightness value of the displayed media content and display screen brightness data associated with the brightness of the display screen, The control system acquires instantaneous pupil size data corresponding to the pupil sizes of one or more of the person's pupils, The control system estimates, at least partially, the dilation or constriction of the pupil caused by the illuminance of the ambient light and the brightness of the display screen viewed by the person, based on the ambient illuminance data and the display screen brightness data. A method comprising using the control system to estimate pupil dilation or constriction caused by at least one of engagement, arousal, or cognitive load experienced by the person, based at least in part on the pupil size data and the photo-induced pupil dilation or constriction.

[0118] EEE2. The method according to EEE1, wherein estimating the dilation or constriction of the light-induced pupil comprises applying a pupil model to the illuminance of the ambient light and the brightness of the display screen viewed by the person.

[0119] EEE3. The method according to EEE2, wherein the pupil model includes personalized pupil model parameters based on one or more measured responses of the person's pupils to brightness and illuminance.

[0120] EEE4. The method of EEE3, wherein determining the personalized pupil model parameters includes estimating the dilation or constriction of the photo-evoked pupil according to the pupil model, measuring the instantaneous pupil size, and determining the estimation error based on the difference between the measured instantaneous pupil size and the estimated dilation or constriction of the photo-evoked pupil according to the pupil model.

[0121] EEE5. The method according to EEE2, wherein the pupil model is at least partially based on the cube of the cosine of the visual angle centered on the foveal position.

[0122] EEE6. The method according to EEE5, wherein the pupil model is also at least partially based on the weighting of light wavelengths according to the relative luminous efficiency function of vision.

[0123] EEE7. The pupil size data is obtained from a camera or eye tracker, according to the method described in any one of EEE1 to EEE6.

[0124] EEE8. The method according to any one of EEE1 to EEE7, further comprising acquiring gaze direction data by the control system, wherein the content brightness data is at least partially based on the gaze direction data.

[0125] EEE9. The method according to any one of EEE1 to EEE8, further comprising the control system estimating a content time interval corresponding to an estimated pupil dilation caused by engagement or cognitive load.

[0126] EEE10. Content corresponding to a content time interval includes video content, audio content, or a combination thereof, as described in EEE9.

[0127] EEE11. Estimating content time intervals corresponding to pupil dilation caused by engagement or cognitive load, including applying a time shift corresponding to the pupil dilation latency period, as described in EEE9 or EEE10.

[0128] EEE12. The method according to EEE11, wherein the time shift is in the range of 1 to 3 seconds.

[0129] EEE13. The method according to any one of EEE1 to EEE12, further comprising estimating an engagement level, an arousal level, a cognitive load level, or a combination thereof, at least in part on an estimated pupillary dilation or constriction caused by engagement, arousal, or cognitive load.

[0130] EEE14. The method according to EEE13, further comprising outputting analytical data based on an estimated engagement level, an estimated arousal level, an estimated cognitive load level, or a combination thereof.

[0131] EEE15. The method of any one of EEE1 to EEE14, further comprising altering one or more aspects of media content in response to an estimated pupillary dilation or constriction caused by at least one of engagement, arousal, or cognitive load.

[0132] EEE16. The method according to EEE15, wherein one or more aspects of the media content are changed after a time in which pupil dilation or constriction is estimated.

[0133] EEE17. The method according to EEE15 or EEE16, wherein the displayed media content is part of a video game, and modifying one or more aspects of the media content includes at least one of modifying the difficulty level of the video game, generating one or more personalized game experiences based on engagement levels, or tracking the cognitive challenges of a player for a competitive game.

[0134] EEE18. The method described in EEE17, which involves generating one or more personalized game experiences based on engagement levels, including at least one of environmental modifications, aesthetic modifications, animation modifications, or game mechanic modifications.

[0135] EEE19. The methods of EEE15, wherein the displayed media content is part of an online learning course, and changing one or more aspects of the media content includes changing one or more aspects of the online learning course.

[0136] EEE20. The method according to EEE19, wherein modifying one or more aspects of the online learning course includes one or more of the following: modifying the amount of information provided in the online learning course, modifying the amount of time spent in at least a portion of the online learning course, or modifying the difficulty level of at least a portion of the online learning course.

[0137] EEE21. The method of EEE15, which involves modifying one or more aspects of the media content, including modifying the identifiability of a graphical object.

[0138] EEE22. The graphical object corresponds to a person or topic, as described in EEE21.

[0139] EEE23. The method according to EEE21 or EEE22, wherein modifying the identifiability of the graphical object includes changing the camera angle, modifying the time the graphical object is displayed, modifying the size at which the graphical object is displayed, or a combination thereof.

[0140] EEE24. Modifying one or more aspects of media content, including modifying one or more aspects of audio content, as described in any one of EEE15 to EEE23.

[0141] EEE25. The method according to EEE24, wherein modifying one or more aspects of the audio content includes adaptively controlling the audio enhancement process.

[0142] EEE26. A method of EEE24 or EEE25 that modifies one or more aspects of audio content, which includes modifying one or more spatial characteristics of audio content.

[0143] EEE27. A method of EEE26 that modifies one or more spatial properties of audio content, which includes rendering at least one audio object at a different position than where at least one audio object would otherwise be rendered.

[0144] EEE28. An apparatus configured to perform the method described in any one of the items EEE1 to EEE27.

[0145] EEE29. A system configured to perform the actions described in any one of the items from EEE1 to EEE27.

Claims

1. A method for estimating pupil dilation or constriction in a person viewing displayed media content, due to cognitive load, arousal, or engagement, The control system acquires ambient illuminance data corresponding to the illuminance of the surrounding light near a person, The control system acquires the content brightness value of the displayed media content and display screen brightness data associated with the brightness of the display screen, The control system acquires instantaneous pupil size data corresponding to the pupil size of one or more of the person's pupils, The control system estimates, at least partially, the dilation or constriction of the pupil caused by the ambient light illuminance and the brightness of the display screen viewed by the person, based on the ambient illuminance data and the display screen brightness data. A method comprising: using the control system to estimate pupil dilation or constriction caused by at least one of engagement, arousal, or cognitive load experienced by the person, at least in part on the pupil size data and the photo-induced pupil dilation or constriction.

2. The method according to claim 1, wherein estimating the dilation or constriction of the light-induced pupil comprises applying a pupil model to the illuminance of the ambient light and the brightness of the display screen viewed by the person.

3. The method according to claim 2, wherein the pupil model includes personalized pupil model parameters based on the measured response of one or more of the person's pupils to brightness and illuminance.

4. The method according to claim 3, wherein determining the personalized pupil model parameters includes estimating the dilation or constriction of the photo-evoked pupil according to the pupil model, measuring the instantaneous pupil size, and determining the estimation error based on the difference between the measured instantaneous pupil size and the estimated dilation or constriction of the photo-evoked pupil according to the pupil model.

5. The method according to claim 2, wherein the pupil model is at least partially based on the cube of the cosine of the visual angle centered on the foveal position.

6. The method according to claim 5, wherein the pupil model is also at least partially based on the weighting of light wavelengths according to the relative luminous efficiency function of vision.

7. The method according to claim 1, wherein the pupil size data is obtained from a camera or eye tracker.

8. The method according to claim 1, further comprising acquiring gaze direction data by the control system, wherein the content brightness data is at least partially based on the gaze direction data.

9. The method according to claim 1, further comprising the control system estimating a content time interval corresponding to an estimated pupil dilation caused by engagement or cognitive load.

10. The method according to claim 9, wherein the content corresponding to the content time interval comprises video content, audio content, or a combination thereof.

11. The method according to claim 9, wherein estimating the content time interval corresponding to the pupil dilation caused by engagement or cognitive load includes applying a time shift corresponding to the pupil dilation latency period.

12. The method according to claim 11, wherein the time shift is within the range of 1 to 3 seconds.

13. The method according to claim 1, further comprising estimating an engagement level, an arousal level, a cognitive load level, or a combination thereof, at least in part on estimated pupil dilation or constriction caused by engagement, arousal, or cognitive load.

14. The method according to claim 13, further comprising outputting analytical data based on an estimated engagement level, an estimated arousal level, an estimated cognitive load level, or a combination thereof.

15. The method according to claim 1, further comprising modifying one or more aspects of the media content in response to the estimated pupil dilation or constriction caused by at least one of engagement, arousal, or cognitive load.

16. The method according to claim 15, wherein one or more aspects of the media content are changed after a time in which pupil dilation or constriction is estimated.

17. The displayed media content is part of a video game, and changing one or more of the media content is, Changing the difficulty level of the aforementioned video game, To generate one or more personalized game experiences based on engagement levels, The method according to claim 15, comprising at least one of tracking the cognitive tasks of a player for a competitive game.

18. The method according to claim 17, wherein generating one or more personalized game experiences based on engagement levels includes at least one of environmental modifications, aesthetic modifications, animation modifications, or game mechanic modifications.

19. The method according to claim 15, wherein the displayed media content is part of an online learning course, and changing one or more aspects of the media content includes changing one or more aspects of the online learning course.

20. The method according to claim 19, wherein changing one or more aspects of the online learning course includes one or more of changing the amount of information provided in the online learning course, changing the amount of time spent in at least a portion of the online learning course, or changing the difficulty level of at least a portion of the online learning course.

21. The method according to claim 15, wherein changing one or more aspects of the media content includes modifying the identifiability of a graphical object.

22. The method according to claim 21, wherein the graphical object corresponds to a person or a topic.

23. The method according to claim 21, wherein modifying the identifiability of the graphical object includes changing the camera angle, modifying the time the graphical object is displayed, modifying the size at which the graphical object is displayed, or a combination thereof.

24. The method according to claim 15, wherein changing one or more aspects of the media content includes changing one or more aspects of the audio content.

25. The method according to claim 24, wherein modifying one or more aspects of the audio content includes adaptively controlling the audio enhancement process.

26. The method according to claim 24, wherein changing one or more aspects of the audio content includes changing one or more spatial characteristics of the audio content.

27. The method according to claim 26, wherein changing the one or more spatial properties of the audio content includes rendering at least one audio object in a location different from the location in which the at least one audio object would otherwise be rendered.

28. It is an electronic device, One or more processors, A system comprising: a memory for storing one or more programs configured to be executed by one or more processors, wherein the one or more programs include instructions for carrying out the method according to any one of claims 1 to 27.

29. A non-temporary computer-readable medium for storing one or more programs configured to be executed by one or more processors of an electronic device, wherein the one or more programs include instructions for performing the method according to any one of claims 1 to 27.