System and method for monitoring and quantification of human attention during media consumption
Patent Information
- Application Number
- PCT/IB2026/000189
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-02-20
- Filing Date
- 2026-02-20
- Publication Date
- 2026-08-27
Smart Images

Figure IB2026000189_27082026_PF_FP_ABST
Abstract
Description
SYSTEM AND METHOD FOR MONITORING AND QUANTIFICATION OF HUMAN ATTENTION DURING MEDIA CONSUMPTION
[0001] This application claims the benefit of U.S. Provisional Patent Application No. 63 / 761,159; filed: February 20, 2025 and entitled "SYSTEM AND METHOD FOR MONITORING AND QUANTIFICATION OF HUMAN ATTENTION DURING MEDIA CONSUMPTION," which is incorporated herein by reference in its entirety including any appendices to the provisional patent application.FIELD
[0002] This disclosure relates generally to monitoring and quantifying human emotion and attention, and more specifically to systems and methods for monitoring and assessing a viewer’s emotional response and attention when exposed to media content.BACKGROUND
[0003] The entertainment and advertising industries are also heavily reliant on subjective ways of analyzing the audience’s experience, such as through post-hoc self-reporting and audience surveys, in evaluating the effectiveness of content such as advertisements and films in achieving a desired objective.SUMMARY
[0004] Emotional valence generally refers to the extent to which an emotional response is positive or negative. For example, certain negative emotions (e.g., sadness, stress, or fear) are considered to have a negative valence, while certain positive emotions (e.g., happiness or excitedness) are said to have a positive valence. Attention may be passive or active. Passive attention generally refers to the extent of a viewer’s sensory perception of content. As opposed to passive attention, active attention includes a degree of cognitive and / or emotional engagement with the perceived content in addition to perception by the senses. Current approaches to evaluating the manner and extent to which media such as advertisements and films achieve their desired effect generally involve measuring passive attention, but not active attention or valence.
[0005] Eye-tracking and facial coding are examples of technologies that are used to monitor passive attention. Eye-tracking technologies monitor the viewer’s gaze at different locations on a screen. Facial coding technologies identify when a viewer is looking at a screen based on an estimation of their head position.
[0006] Accordingly, current objective approaches to measuring audience attention simply measure passive attention, for example, whether, and where, a viewer is looking at a screen. Current approaches such as eye-tracking and facial coding fail to comprehensively characterize the viewer’s emotional reaction. Other methods of measuring audience attention are subjective, relying on post-hoc self-reporting. This leads to inaccuracies in the attention data caused by individual biases and memory limitations. Accordingly, there exists a need for an objective way to measure and analyze a viewer’s active attention and emotional response to media content in order to more accurately evaluate whether the content achieves its desired effect on the audience.
[0007] Provided herein are systems and methods of measuring and analyzing a viewer’s attention and emotional response to perceived content. In some examples, a method may involve collecting physiological data in multiple modalities, such as facial expression data (e.g., a distance from the eyes and / or a sensor on a wearable eyewear device to the tops of a viewer’ s cheekbones) and sensor data such as electrodermal activity (EDA) data and / or photoplethysmogram (PPG) data. By collecting data in multiple modalities, a more robust and accurate representation of a user’s emotional response and attentiveness can be ascertained.
[0008] A method may include collecting baseline physiological data from the viewer by displaying content that has been pre-selected as being likely to elicit a particular emotional response. For example, a pre-selected collection of video clips that are intended to be “funny” or “sad” may be shown to the viewer in order to measure the viewer’s baseline positive and negative valence responses. The extent of a viewer’s attention (hereinafter “activation”) may be determined from the baseline valences and / or from the sensor data. By determining baseline valences and activations for each viewer, differences between individuals in terms of emotional sensitivity and / or attentiveness can be accurately and objectively accounted for.
[0009] Once the pre-selected content has been shown to the user for the purposes of collecting the physiological baseline, content for which the viewer’s response is desired to be assessed (hereinafter “content of interest”) may be displayed to the viewer. Data representing the viewer’s physiological response may be collected in real time while the content of interest is being displayed. For different temporal windows of the content (for example, every few seconds), the viewer’s valence and activation may be determined from the physiological response data, taking into account the baseline valence and activation. A method may further include generating an outputting a dynamic visualization representing the viewer’s valence and activation for the temporal window, and the visualization may be updated in real time to include valence and activation temporal windows throughout the entire content. The visualization may be generated and updated on the level of an individual viewer, a small subset of an audience, or an entire audience. Individual and group-level analytics can be generated based on the physiological responses to the content of interest.
[0010] By providing objective ways to measure valence and audience attention, the systems and methods of the present disclosure can provide content creators and producers with more accurate metrics to quantify the effectiveness of media in achieving its desired effect on an audience.
[0011] In some embodiments, a method of characterizing a viewer’s physiological response to perceived content includes: displaying, on a display device, pre-selected content associated with predetermined valence and activation levels; while displaying the pre-selected digital content, collecting first physiological response data from the viewer; determining, based on the first physiological response data, a baseline valence and baseline activation of the viewer; displaying, on the display device, content of interest to the viewer; and while displaying the content of interest: collecting second physiological response data from the viewer; determining a valence and activation of the viewer over a temporal window of the content of interest based on the second physiological response data, the baseline valence, and the baseline activation; and generating and outputting a dynamic visualization representing the viewer’s valence and activation for the temporal window of the content of interest.
[0012] In some embodiments, collecting first physiological response data from the viewer comprises collecting physiological data using a camera, a wearable device, or both.
[0013] In some embodiments, collecting first physiological response data from the viewer comprises collecting physiological data from the viewer using a wearable device, and wherein determining, based on the first physiological response data, a baseline valence and baseline activation of the viewer comprises processing wearable device data by performing one or more of filtering, artifact detection, peak detection, or deconvolution on the wearable device data.
[0014] In some embodiments the first physiological response data comprises one or more of facial expression data, electrodermal activity data, or heart rate data.
[0015] In some embodiments, determining, based on the first physiological response data, a baseline valence of the viewer comprises: normalizing facial expression data for groups of positive and negative facial expressions; determining an average positive valence and an average negative valence based on the normalized facial expression data; and subtracting the average positive valence from the average negative valence.
[0016] In some embodiments, determining, based on the first physiological response data, a baseline activation of the viewer comprises normalizing facial expression data for groups of facial expressions indicative of attentiveness and for groups of facial expressions indicative of lack of attentiveness; determining an average maximum valence from the normalized facial expression data for the group of facial expressions indicative of attentiveness; determining an average minimum valence from the normalized facial expression data for the group of facial expressions indicative of lack of attentiveness; and determining the baseline activation based on subtracting the average maximum valence from the average minimum valence.
[0017] In some embodiments, the dynamic visualization comprises a plot having a point for each temporal window of the content of interest, the point having a first coordinate location representing valence and a second coordinate location representing activation.
[0018] In some embodiments, the dynamic visualization comprises a plurality of zones, each of the plurality of zones representing a degree of the viewer’s immersion with the content of interest.
[0019] In some embodiments, the dynamic visualization is outputted to a backend user.
[0020] In some embodiments, the method further comprises collecting the first and second physiological response data from a plurality of viewers in an audience; and generating analytics characterizing the audience’s physiological response to the content of interest.
[0021] In some embodiments, the method comprises outputting the analytics to a backend user.
[0022] In some embodiments, a system is provided, the system including one or more processors; and memory storing computer program code executable by the one or more processors to cause the system to: display, on a display device, pre-selected content associated with predetermined valence and activation levels; while displaying the pre-selected digital content, collect first physiological response data from the viewer; determine, based on the first physiological response data, a baseline valence and baseline activation of the viewer; display, on the display device, content of interest to the viewer; and while displaying the content of interest: collect second physiological response data from the viewer; determine a valence and activation of the viewer over a temporal window of the content of interest based on the second physiological response data, the baseline valence, and the baseline activation; and generate and output a dynamic visualization representing the viewer’s valence and activation for the temporal window of the content of interest.
[0023] In some embodiments, a non-transitory computer readable storage medium storing one or more programs is provided, the one or more programs comprising instructions, which, when executed by a system comprising one or more processors and memory, cause the system to: display, on a display device, pre-selected content associated with predetermined valence and activation levels; while displaying the pre-selected digital content, collect first physiological response data from the viewer; determine, based on the first physiological response data, a baseline valence and baseline activation of the viewer; display, on the display device, content of interest to the viewer; and while displaying the content of interest: collect second physiological response data from the viewer; determine a valence and activation of the viewer over a temporal window of the content of interest based on the second physiological response data, the baseline valence, and the baselineactivation; and generate and output a dynamic visualization representing the viewer’s valence and activation for the temporal window of the content of interest.
[0024] In some embodiments, any one or more of the characteristics of any one or more of the systems, methods, and / or computer-readable storage mediums recited above may be combined, in whole or in part, with one another and / or with any other features or characteristics described elsewhere herein.
[0025] As used herein, a “valence” refers to a value representing the extent of a person’s emotional response. A valence that is “positive” corresponds to one or more positive emotions, including: astonished, delighted, excited, happy, pleased, content, satisfied, relaxed, and calm. A valence that is “negative” corresponds to one or more negative emotions, including: “angry, afraid, annoyed, distressed, frustrated, depressed, sad, and bored.
[0026] As used herein, an ’’activation” refers to a value representing the degree of a viewer’ s active attention. A low activation corresponds to a low degree of active attention, while a high activation corresponds to a high degree of active attention.BRIEF DESCRIPTION OF THE FIGURES
[0027] The patent or application file contains at least one drawing executed in color. Copies of this patent or patent application publication with color drawing(s) will be provided by the Office upon request and payment of the necessary fee.
[0028] A better understanding of the features and advantages of the present disclosure will be obtained by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the disclosure are utilized, and the accompanying drawings of which:
[0029] FIG. 1 illustrates an exemplary system, according to some embodiments.
[0030] FIG. 2A illustrates an exemplary method, according to some embodiments.
[0031] FIG. 2B illustrates an exemplary method continuing from FIG. 2A, according to some embodiments.
[0032] FIG. 3 illustrates an exemplary baselining protocol, according to some embodiments.
[0033] FIG. 4 illustrates an invite page of a user interface, according to some embodiments.
[0034] FIG. 5 illustrates a user dashboard of a user interface, according to some embodiments.
[0035] FIGs. 6 through 8 illustrate screening preview pages of the user interface, according to some embodiments.
[0036] FIGs. 9-11 illustrate a dynamic visualization page of a user interface, according to some embodiments.
[0037] FIG. 12 illustrates an exemplary dynamic visualization, according to some embodiments.
[0038] FIG. 13 illustrates an exemplary dynamic visualization, according to some embodiments.
[0039] FIG. 14 illustrates an exemplary method for determining a normalized valence, according to some embodiments.
[0040] FIG. 15 illustrates an exemplary method for determining a normalized activation, according to some embodiments
[0041] FIG. 16 illustrates exemplary signal processing sequences, according to some embodiments.
[0042] FIG. 17 illustrates a computer, according to some embodiments.DETAILED DESCRIPTION
[0043] Provided are systems and methods of characterizing a viewer’s physiological response to perceived content. In some examples, a method may include displaying pre-selected content associated with predetermined valence and activation levels to a viewer of a display device. The pre-selected content may be used to elicit first physiological response data from the viewer, which may be used to determine a baseline valence and baseline activation of the viewer for assessing the viewer’s physiological response to content of interest. The physiological response data, including valence and activation, may be obtained from one or more cameras positioned on or built into the display device (e.g., a webcam of a viewer’s laptop), one or more cameras positioned within a screening room or movie theater, and / or from a wearable device such as glasses or goggles or a wearable EDA or PPG sensor.
[0044] In some examples, content of interest may be displayed on the display device. The content may be of interest to the viewer (for example, the viewer may be able to select the content thatthey wish to view). In some examples, the content may additionally or alternatively be of interest to a content creator, media producer, advertising agency, or anyone else who may be interested in evaluating the efficacy of the content in achieving a desired effect on a viewer. While displaying the content of interest, the method may include collecting second physiological response data from the viewer, and the viewer’s valence and activation can be determined over a temporal window (e g., at regular time intervals throughout the content, or at specific time intervals such as during the beginning, climax, and / or ending of the content). The viewer’s valence and activation can be determined based on the baseline valence and activation, such that the valence and activation are normalized to the unique physiology of the individual viewer.
[0045] A dynamic visualization representing the viewer’s valence and activation for the temporal window of the content of interest may be generated and outputted to the viewer and / or the person or entity who is interested in analyzing the viewer’s physiological response. The dynamic visualization may include, for example, a plot with valence on a first axis and activation on a second axis. Optionally, the dynamic visualization may be divided into a plurality of zones, with each zone representing a degree of the viewer’ s immersion (e.g., combined valence and activation). Analytics may be generated from the dynamic visualization at the level of the individual viewer as well as an audience, providing objective insights into the emotional response and attentiveness of one or more viewers at various points throughout the content of interest.
[0046] In the following description of the various embodiments, it is to be understood that the singular forms ”a,” ”an,” and the” used in the following description are intended to include the plural forms as well, unless the context clearly indicates otherwise. It is also to be understood that the term “and / or” as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items. It is further to be understood that the terms “includes, ’’including,” ’’comprises,” and / or ’’comprising,” when used herein, specify the presence of stated features, integers, steps, operations, elements, components, and / or units but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, units, and / or groups thereof.
[0047] Certain aspects of the present disclosure include process steps and instructions described herein in the form of an algorithm. It should be noted that the process steps and instructions of the present disclosure could be embodied in software, firmware, or hardware and, when embodied in software, could be downloaded to reside on and be operated from different platforms used by a variety of operating systems. Unless specifically stated otherwise as apparent from the following discussion, it is appreciated that, throughout the description, discussions utilizing terms such as “processing,” “computing,” “calculating,” “determining,” “displaying,” “generating” or the like, refer to the action and processes of a computer system, or similar electronic computing device, that manipulates and transforms data represented as physical (electronic) quantities within the computer system memories or registers or other such information storage, transmission, or display devices.
[0048] The present disclosure in some embodiments also relates to a device for performing the operations herein. This device may be specially constructed for the required purposes, or it may comprise a general-purpose computer selectively activated or reconfigured by a computer program stored in the computer. Such a computer program may be stored in a non-transitory, computer readable storage medium, such as, but not limited to, any type of disk, including floppy disks, USB flash drives, external hard drives, optical disks, CD-ROMs, magnetic-optical disks, read-only memories (ROMs), random access memories (RAMs), EPROMs, EEPROMs, magnetic or optical cards, application specific integrated circuits (ASICs), or any type of media suitable for storing electronic instructions, and each connected to a computer system bus. Furthermore, the computing systems referred to in the specification may include a single processor or may be architectures employing multiple processor designs, such as for performing different functions or for increased computing capability. Suitable processors include central processing units (CPUs), graphical processing units (GPUs), field programmable gate arrays (FPGAs), and ASICs.
[0049] The methods, devices, and systems described herein are not inherently related to any particular computer or other apparatus. Various general-purpose systems may also be used with programs in accordance with the teachings herein, or it may prove convenient to construct a morespecialized apparatus to perform the required method steps. The structure for a variety of these systems will appear from the description below. In addition, the present invention is not described with reference to any particular programming language. It will be appreciated that a variety of programming languages may be used to implement the teachings of the present disclosure as described herein.Exemplary System and Method for Monitoring and Quantifying Valence and Activation
[0050] FIG. 1 illustrates a system 100 for monitoring and quantifying the valence and activation of a content viewer, according to one or more examples of the present disclosure. System 100 may include a display device 102, on which media content may be displayed for viewing. Display device 102 may include a screen for displaying outputs and may include one or more processors for receiving and / or processing inputs. Display device 102 may be connected to, and / or be in communication with, one or more devices over a network for receiving inputs. Display device 102 may be, for example, a desktop or laptop computer, a television, a movie theater screen, a projector, and / or a mobile device such as a smartphone or a tablet. Display device 102 may optionally have a touchscreen for facilitating touch-based interactions with a user interface, as will be described.
[0051] System 100 may include one or more monitoring devices 104. Monitoring devices may include, for example, cameras and / or wearable devices that may have one or more sensors for collecting physiological data. Monitoring devices 104 may include, for example, cameras, sensors configured to measure a distance from the eyes to the tops of the wearer’s cheekbones, and / or sensors for measuring electrodermal activity (EDA) and / or heart rate (PPG). In some examples, the monitoring devices may be wearable devices such as glasses, bracelets, watches, or rings that may have one or more built-in sensors. In some examples, the monitoring devices may include devices such as smart watches or rings that are already owned by the viewer, and other components of system 100 may be configured to be compatible with these devices. In some examples, a user may be administered one or more monitoring devices for use in the methods described herein. A camera such as a webcam built in or connected to the viewer’s computer or placed around thecontent screening area may also be used to collect physiological data such as facial expression data from the viewer.
[0052] System 100 may include an immersion analysis platform 106. Immersion analysis platform 106 may include a frontend application (e.g., a user interface) and a backend infrastructure. Immersion analysis platform 106 may generate and manage data requests and transmissions from the frontend application to one or more components of the backend infrastructure. Immersion analysis platform 106 may be configured to perform various data processing techniques on data received from monitoring devices 104, for example, through one or more software modules that include data processing algorithms embodied in code. Immersion analysis platform 106 may be configured to calculate the valence and activation (attention) of a viewer as a normalized baseline and in real-time while the viewer perceives content of interest while the monitoring devices 104 collect physiological data. Immersion analysis platform 106 may be configured to perform these and other functionalities using code programmed in any suitable programming language, including but not limited to C, C#, C++, Python, Rust, Lua, and / or JavaScript. In some examples, the backend infrastructure of immersion analysis platform 106 may be implemented on a cloud computing service such as Amazon Web Services, Microsoft Azure, or Google Cloud Platform.
[0053] System 100 may include a data store 108. Data store 108 may be used to store raw data collected by monitoring devices 104, as well as data that is processed and / or generated by immersion analysis platform 106. For example, data store 108 may store baseline valence and activation data, demographic data on a viewer or audience, instances of dynamic visualizations generated while the viewer views content of interest, and analytics that are generated based on the viewer’s reaction to the content of interest. Data store 108 may also store the pre-selected content that is to be shown to the viewer to determine the baseline valence and activation levels. In some examples, data store 108 may be a multi -model database such as Azure Cosmos DB, Amazon DynamoDB, ArangoDB, or MongoDB Atlas.
[0054] System 100 may include a backend user device 110. In some examples, the backend user device 110 may be used by a backend user such as a content creator or representative of the entitycreating the content of interest, such as a film producer, director, advertising agency representative, independent filmmaker, or any other person interested in studying the effectiveness of the content of interest in achieving a desired reaction from a viewer. In some examples, the immersion analysis platform 106 and backend user device 110 may communicate via an application programming interface (API), such that immersion analysis platform 106 may receive inputs from and / or provide outputs to backend user device 110. For example, immersion analysis platform 106 may be configured to receive inputs from backend user device 110 that include requests for analytics to be generated on the viewer’s valence and activation at various temporal windows in the content of interest. These analytics may be generated by immersion analysis platform 106 at the level of the individual viewer or a larger viewing audience and may be outputted to backend user device 110. In some examples, immersion analysis platform 106 may be configured to generate a dynamic visualization of the viewer’s valence and activation as the viewer views content of interest, and the dynamic visualization may be displayed on backend user device 110. Like display device 102, backend user device 110 may include a display and one or more processors. Backend user device 110 may be, for example, a desktop or laptop computer and / or a mobile device such as a smartphone or a tablet.
[0055] FIGs. 2A and 2B illustrate an exemplary method 200 for characterizing a viewer’s physiological response to perceived content. One or more steps of method 200 may be implemented by one or more components of system 100. At step 202, method 200 may involve collecting baseline physiological data from the viewer. The baseline physiological data may include, for example, baseline facial expression data (e.g., data representing the viewer’s resting face or neutral expression) as well as minimum and maximum values of valence, which may be determined in a variety of ways.
[0056] For example, FIG. 3 illustrates sub-steps of step 202 of method 200 that can be performed to obtain baseline physiological data from the viewer. At step 202a, the viewer may be displayed a prompt on display device 202, for example, on a user interface or on the movie theater screen or television within the viewing area. The prompt may instruct the viewer to hold a neutral facialexpression for a fixed time period (e.g., seconds or minutes). While the viewer maintains a neutral facial expression for a fixed time period, a monitoring device I 04 (e.g., a camera or wearable device) may perceive the viewer’s face. Collecting neutral facial expression data in this manner can, in turn, make subsequent determinations of valence and / or activation more accurate, since unique differences in the resting positions of facial features between different individuals may be accounted for. As will be described, a minimum valence may be determined from the physiological data collected at step 202a, which may be used to normalize the valence and activation determined from physiological data collected from the viewer as they view content of interest.
[0057] At step 202b, the viewer may be shown emotional content that is intended to elicit an emotional response while one or more monitoring devices 104 perceive the viewer. For example, the viewer may be shown one or more images, a short film, a video, and / or a compilation of video clips of a predetermined duration (e.g., five or ten minutes) that are generally associated with positive and / or negative emotional responses so that relative maximum values of valence can be determined for the viewer. For example, the viewer may be shown a video clip that has been perceived by a general viewing audience as very ’’happy” or ’’funny” in order to elicit a maximum positive valence from the viewer, while a viewer may be shown a video clip that has been perceived by a general viewing audience as very ”sad” or ’’scary” to elicit a maximum negative valence from the viewer. Facial expression data, as well as wearable device data (e.g., skin conductance levels, heart rate, heart rate variability, and so on). As will be described, a relative maximum positive and negative valence may be determined from the physiological data collected in this step, which can later be used to normalize the valence and activation of the viewer as they view content of interest.
[0058] At step 202c, the viewer may be shown content that is intended to relax the viewer for a predetermined period of time. Step 202c can serve multiple purposes. First, step 202c can be used to determine a relative minimum value for wearable device data (e.g., skin conductance levels, heart rate, heart rate variability, and so on). Second, step 202c can be used to return the viewer to a neutral emotional state prior to viewing the content of interest, to ensure that their response to the content of interest is not affected by any preexisting emotions or distractions. Accordingly,steps 202a-202c may be performed in order, so that the viewer is placed in a neutral / relaxed state before and after viewing stimulating content to reduce the impact that any external influences may have on determining the valence and activation of the viewer.
[0059] Turning back to FIG. 2A, at step 204, method 200 may include determining baseline valences by normalizing physiological data. Valences can be determined from the physiological data in several ways. For example, one or more machine learning algorithms may be trained to classify facial expression data as belonging to a particular emotion and to assign the emotion an intensity value. A machine learning algorithm may be trained to analyze images or videos of a viewer’s face, for example, and may output a probability as to whether the facial expression data is indicative of a particular emotion (e.g., expressed as a confidence / intensity score from Oto 1). If the machine learning model’s outputted intensity or confidence that the facial expression data is indicative of a particular emotion is above a predefined threshold (e.g., 80%), then the facial expression data may be classified as corresponding to a particular emotion. In some examples, a distance from the tops of the viewer’s cheekbones to their eyes, from the viewer’s eyes to their eyebrows, and / or from the corners of the viewer’s mouth to their eyes may also be used to classify and / or measure valence.
[0060] Based on the physiological data, and optionally the outputs from an emotion analysis machine learning model, valence values may be determined. A “positive” valence may include the emotions “astonished,” “delighted,” “excited,” “happy,” “content,” “pleased,” “satisfied,” and / or “relaxed,” while a “negative” valence may include the emotions ’’afraid,” “angry,” “annoyed,” “distressed,” “frustrated,” “depressed,” “bored,” or “sad.” The positive and negative valence may be normalized to the viewer using minimum-maximum normalization, for example, in accordance with the following:1) Norm (positive)= Norm(astonished, delighted, excited, happy, content, pleased, satisfied, relaxed)= (data point - minimum value) / (maximum value - minimum value)2) Norm(negative)= Norm(afraid, angry, annoyed, distressed, frustrated, depressed, bored, sad)= (data point - minimum value) / (maximum value - minimum value) In these formulas, the ’’data point” may represent the emotional intensity value outputted by the machine learning model, or it may be raw data representing the metric used to determine valence (e.g., the raw data representing the distance from the eyes to the cheekbones at a given time point, if this metric is to be used). The minimum and maximum values used in the above formulas may represent the minimum and maximum positive or negative emotional intensity values, respectively, for the data collected within the step of the baselining protocol (e.g., one of steps 202a-202c).
[0061] The valence calculations can be determined for each identified emotion that was detected from the viewer during each stage of the physiological baselining protocol, or the valence calculations may be determined by grouping the valences of emotions together based on whether the emotions are positive or negative. For example, physiological data may be normalized at the individual emotion level, or emotions can be grouped together and averaged based on whether the valence is ’’positive” (e.g., corresponding to a positive emotion) or ’’negative (e.g., corresponding to a negative emotion). Averages can be taken for the positive normalized valences and the negative normalized valences, and the average negative normalized valence can be subtracted from the average positive normalized valence to determine a representative valence value for the user for that particular stage of the baselining protocol. For example, for step 202a of the physiological baselining protocol, the representative valence calculated in the manner above would correspond to the viewer’s relative minimum valence, while for step 202b, the representative valence calculated in the manner above would correspond to the viewer’s relative maximum valence. An exemplary method for determining a normalized valence for positive and negative groups of emotions is illustrated and summarized in Appendix A and in FIG. 14, but in some examples, the method provided in FIG. 14 may also be performed at the individual emotion level.
[0062] From step 204, a baseline activation (e.g., attentiveness) of the viewer can be determined in several ways. Turning to step 206, one way to determine a baseline activation is to use data collected from one or more devices worn by the user, such as one or more wearable sensors formeasuring electrodermal activity (EDA) and / or heart rate (PPG). Step 206 may involve processing the wearable sensor data using one or more signal processing protocols that are configured in immersion analysis platform 106. For the raw EDA signal, immersion analysis platform 106 may be configured to filter noise, detect and remove artifacts, and perform deconvolution on the signal before using the EDA metrics to calculate physiological activation. For the raw PPG signal, immersion analysis platform 106 may be configured to filter noise from the PPG signal, detect and remove artifacts in the PPG signal, and detect peak heart rates from the PPG signal before using the PPG metrics to calculate physiological activation.
[0063] Processing of the EDA signal may result in several metrics, including (1) mean Skin Conductance Level (SCL) captured in microSiemens, (2) frequency of Skin Conductance Responses (SCR freq), and (3) magnitude of Skin Conductance Response (SCR mag), also captured in microSiemens. Processing of the PPG signal may result in heart rate (beats / second) and heart rate variability measured as a root mean square of successive differences. These metrics may be calculated for each sample point taken as the viewer views content. The baseline activation can be determined by performing minimum / maximum normalization on the wearable sensor data collected at each sampling interval during sub step 202c of step 202. The normalized values may be entered into the following formula to quantify the baseline activation:3) Physiological Activation (SCL + SCR freq+ SCR mag+ HR- HRV)
[0064] In some examples, the baseline activation may be determined from the normalized physiological data, as indicated at step 208. For example, the baseline activation may be determined by normalizing facial expression data of certain emotions that correspond to high and low levels of attention:4) Norm (High)= Norm(angry, annoyed, afraid, astonished, delighted, excited)= (data point\- minimum value) / (maximum value- minimum value)5) Norm (Low)= Norm(bored, calm, relaxed, depressed, satisfied, sleepy)= (data point-minimum value) / (maximum value- minimum value)
[0065] As described with respect to step 204, the “data point” may represent the emotional intensity value outputted by the machine learning model, or it may be from the raw data used as the metric for quantifying valence. The minimum and maximum values used in the above formulas may represent the minimum and maximum values, respectively, for the data collected within one of steps 202a-202c. Once the values for each sample point are normalized, an average of the high attention normalized points and an average of the low attention normalized points can be determined. The baseline activation can be determined by subtracting the average low from the average high. An exemplary method for determining a normalized activation is illustrated and summarized in Appendix A and in FIG. 15.
[0066] Once the baseline valence and activation have been determined from the data collected at steps 202a-202c, at step 210, content of interest may be shown to the viewer. The content of interest may be, for example, a video, film, advertisement, and the like for which the viewer’s emotional and attentive response are to be determined. At step 212, while the content of interest is shown to the viewer, the viewer’s valence and activation may be determined in real time. The valence and activation determined while the viewer views the content of interest can be determined using the same formulas as the baseline valence and activation, but the valence and activation collected while the viewer views the content of interest can be baselined and / or normalized further. For example, the valence and / or activation while the viewer views the content of interest can be baselined by subtracting the baseline valence and / or activation from the valence and / or activation that are observed in real time. The valence and / or activation while the viewer views the content of interest can be normalized through the minimum / maximum normalization formula provided above (e g., data point- minimum value) / (maximum value - minimum value).
[0067] Continuing from FIG. 2A to FIG. 2B, at step 214, the valence and activation determined at each sample point while the viewer views the content of interest can be characterized into a zone of immersion. A “zone of immersion” may be a categorical label used to describe the extent of the viewer’s combined emotional and attentive response to the content. In some examples, a sample point may be assigned to a zone of immersion by plotting the point on a graph that includes thevalence provided on the x-axis and the activation included on they-axis. The graph may be divided into a plurality of zones based on a number of sample points that are to be collected and / or a number of viewers in the viewing audience. Each zone may be labeled according to a relative degree of immersion. For example, as shown in FIG. 13, a graph may be divided into four zones of immersion, labeled “passive,” “engaged,” “engrossed,” or “immersed.” While the appearance of the graph may vary depending on the needs of a backend user, in an example such as FIG. 13 where the zones of immersion are concentric semicircles centered around an origin, the amount of “immersion” can be determined as the distance from the sample point to the origin (e.g.,) / (xA2 + yA2). For example, in FIG. 13, a distance to the origin from 0-0.25 may correspond to the “passive” zone, a distance from 0.25-0.5 may correspond to an “engaged” zone, a distance from 0.5-0.75 may correspond to an “engrossed” zone, and a distance from 0.75-1 may correspond to an “immersed” zone. In this manner, each sample point can be assigned to a zone of immersion at step 214. At step 216, the plot of the zones of immersion can be generated and / or updated by the immersion analysis platform as new data is received. At step 218, analytics may be generated from the zones of immersion plot based on one or more inputs received from a backend user. For example, a backend user such as a film producer may be interested in viewing analytics on the valences and / or activations associated with a target subset of the audience such as a particular age demographic or gender. A backend user may request these analytics at step 218, and these analytics can be generated by immersion analysis platform 106 and displayed to the backend user at step 220.Exemplary User Interface for Monitoring and Quantifying Valence and Activation
[0068] Turning to FIG. 4, FIG. 4 illustrates an invite page of a user interface. The invite page may be used by a backend user to invite one or more viewers to a viewing experience, or by a frontend user (e g., the content viewer) to enroll in a screening of content.
[0069] FIG. 5 illustrates a user dashboard of a user interface. In some examples, the user dashboard may include one or more surveys for the frontend user (e.g., the content viewer) to fill out so thatdemographic information can be collected on the viewer for the purposes of generating viewer and / or audience-level analytics.
[0070] FIGs. 6 through 8 illustrate screening preview pages of the user interface. The screening launch pages may display information to the viewer to inform their viewing experience. For example, the screening preview page may display an overview of the screening process, as shown in FIG. 6, and / or a prompt or request for permission to access one or more monitoring devices that are perceiving the viewer, as shown in FIGs. 7 and 8.
[0071] FIGs. 9-11 illustrate a dynamic visualization page of a user interface. In some examples, the dynamic visualization page may be displayed to a backend user of system 100 while the viewer views the content of interest. On the dynamic visualization pages, one or more dynamic visualizations may be displayed. For example, a graph representing the zones of immersion for each sample point may be shown in a first location on the dynamic visualization page, while a plot of the magnitude and direction of the viewer’s valence over different temporal windows of the content of interest can be shown at another location on the dynamic visualization page.
[0072] FIG. 12 illustrates an example of a dynamic visualization of a viewer’s valence and activation. In this example, the dynamic visualization is a complete circle, with valence plotted on the x-axis and activation plotted on the y-axis in both the positive and negative directions. Points can be plotted in real time as physiological data is collected from the viewer and as valence and activation are calculated using the immersion analysis platform, and the dynamic visualization may be displayed to the backend user.
[0073] FIG. 13 illustrates a dynamic visualization of a viewer’s immersion. As shown, the viewer’s valence may be plotted on the x-axis, while the viewer’s activation may be plotted on the y-axis. In this example, valence can be positive or negative, but the magnitude of the activation is plotted such that the visualization takes a semi-circular shape. The dynamic visualization may include a plurality of zones characterizing each data point collected from the viewer under a relative degree of immersion. In this example, the dynamic visualization includes four concentric semicircular zones of immersion that divide the visualization evenly into four intervals. However,more or less zones of immersion may be configured differently in the dynamic visualization based on preferences of the backend user and / or to better suit the pattern of data points as they are collected from the viewer.
[0074] FIG. 17 depicts a computer 1700, in accordance with various embodiments. One or more components of computer 1700 may make up one or more components of system 100 as described herein. For example, computer 1700 can be a host computer connected to a network (e.g., a computer hosting immersion analysis platform 106). Computer 1700 can be a client computer or a server. As shown in FIG. 17, computer 1700 can be any suitable type of microprocessor-based device, such as a personal computer, workstation, server, videogame console, or handheld computing device, such as a phone or tablet. The computer can include, for example, one or more of processor 1701, computer input device 1702, output device 1703, storage 1704, and communication device 1705. Computer input device 1702 can generally correspond to those described above and can either be connectable or integrated with the computer.
[0075] Computer input device 1702 can be any suitable device that provides input, such as a touch screen or monitor, keyboard, mouse, or voice-recognition device. Output device 1703 can be any suitable device that provides output, such as a touch screen, monitor, printer, disk drive, or speaker.
[0076] Storage 1704 can be any suitable device that provides storage, such as an electrical, magnetic, or optical memory, including a RAM, cache, hard drive, CD-ROM drive, tape drive, removable storage disk, or other non-transitory computer readable medium. Storage 1804 can include one storage device or more than one storage device. As used herein, the terms storage, memory, and / or storage medium / media may refer to singular and / or plural devices which may store data and / or code / instructions individually, redundantly, and / or in cooperation with one another, for example in a local and / or cloud storage environment. Communication device 1705 can include any suitable device capable of transmitting and receiving signals over a network, such as a network interface chip or card. The components of the computer can be connected in any suitable manner, such as via a physical bus or wirelessly. Storage 1704 can be a non-transitory computer-readable storage medium comprising one or more programs, which, when executed byone or more processors, such as processor 1701, cause the one or more processors to execute methods described herein.
[0077] Software 1706, which can be stored in storage 1704 and executed by processor 1801, can include, for example, the programming that embodies the functionality of the present disclosure (e.g., as embodied in the systems, computers, servers, and / or devices as described above). In some embodiments, software 1706 can be implemented and executed on a combination of servers such as application servers and database servers.
[0078] Software 1706, or part thereof, can also be stored and / or transported within any computer-readable storage medium for use by or in connection with an instruction execution system, apparatus, or device, such as those described above, that can fetch and execute instructions associated with the software from the instruction execution system, apparatus, or device. In the context of this disclosure, a computer-readable storage medium can be any medium, such as storage 1704, that can contain or store programming for use by or in connection with an instruction execution system, apparatus, or device.
[0079] Software 1706 can also be propagated within any transport medium for use by or in connection with an instruction execution system, apparatus, or device, such as those described above, that can fetch and execute instructions associated with the software from the instruction execution system, apparatus, or device. In the context of this disclosure, a transport medium can be any medium that can communicate, propagate, or transport programming for use by or in connection with an instruction execution system, apparatus, or device. The transport-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, or infrared wired or wireless propagation medium.
[0080] Computer 1700 may be connected to a network, which can be any suitable type of interconnected communication system. The network can implement any suitable communications protocol and can be secured by any suitable security protocol. The network can comprise network links of any suitable arrangement that can implement the transmission and reception of networksignals, such as wireless network connections, T1 or T3 lines, cable networks, DSL, or telephone lines.
[0081] Computer 1700 can implement any operating system suitable for operating the network. Software can be written in any suitable programming language, such as C, C++, Java, or Python. In various embodiments, application software embodying the functionality of the present disclosure can be deployed in different configurations, such as in a client / server arrangement or through a web browser as a Web-based application or Web service, for example.Appendix AThe protocol for the measurement of immersion is constructed as linear process that incorporates algorithmic modules with a defined method for data collection.The Onboarding protocol presents participants with clear instructions to check that: (1) webcam has a clear view of the face, and (2) that the wearable sensor is correctly fitted. The Neutral Face Protocol is the first part of our personalisation process, participants are required to hold a neutral facial expression for 15 seconds; this provides minimum values for facial emotional expression and allows the protocol to control for individual differences in facial morphology. The Emotional Reactivity Protocol requires participants to view a curated selection of film clips for 6 minutes, the resulting data yields maximum values for facial emotions and phy siology. The final protocol is called the Physiological Baselining Protocol and requires participants to watch a relaxing film clip for 4 minutes.The resulting data set from the protocol are combined in order to calculate levels of Activation and Emotional Valence. All data are normalised to the individual prior to the calculation of Activation and Emotional Valence. Specifically, data are subjected to amin-max normalisation protocol using this formula: normalised data = (raw data -minimum value) / (maximum value - minimum value). All emotional and physiological data has been normalised from this point onwards using the maximum and minimum values gathered during the three protocols described earlier.To calculate the level of Emotional Valence, we calculate scores for Positive and Negative valence based on the normalised emotional facial expression data.To calculate the level of Activation, we can use one of the two following approaches. The first one utilises data from the normalised facial expression data.Tn order to generate these data, raw data from the facial expression sensor and the protocols is stored in a database called Cosmos. The calculations are created in a bespoke form using SQL via Synapse Link.The second method for calculating the level of Activation utilises data from the wearable device. This device yields two signals: electrodermal activity (EDA), and a photoplethysmogram (PPG). Both signals will be subjected to processing protocols that have been designed and tested in-house and encapsulated in scripts written in Python. As before, data will be stored in a Cosmos database and shared with Python code modules. This processing takes the form of data 'cleaning', artifact correction and generation of metrics. The stages for each protocol are provided below.The EDA will yield three distinct metrics: (1) mean Skin Conductance Level (SCL) captured in microSiemens, (2) frequency of Skin Conductance Responses (SCR freq), and (3) magnitude of Skin Conductance Response (SCR mag), also captured in microSiemens. The PPG will deliver two metrics, which are heart rate (HR) (in beats per seconds) and heart rate variability (HRV) expressed as Root Mean Square of Successive Differences (RMSSD).These metrics are calculated for each time epoch when content is viewed and will be either baselined (i.e., current value - value calculated during the Physiological Baselining Protocol) or subjected to min-max normalisation using minimum values from the Physiological Baselining Protocol and maximum values from Emotional Reactivity Protocol.The calculated values for Emotional Valence and Activation are combined in a two-dimensional space that is configured as shown below. Activation is the vertical axis and Emotional Valence is the horizontal axis. The two measures combine to describe31the level of immersion experienced by individual(s) who are viewing content. This two-dimensional space is divided into four zones of immersion that are defined by the combination of Activation and Valence.Each epoch of data collection will place the experience in one of the four zones of immersion as shown in the figure below. Metrics of immersion are subsequently generated based on: (1) number of epochs recorded in each zone of immersion for each individual when viewing media, and (2) number of epochs recorded in each zone of immersion for each individual that is positively valenced and number that is negatively valenced. The same metrics can be scaled-up to characterise an audience or sub-group of an audience.By way of another example, if the physiological data collected from the viewer during step 202b indicated the following emotions and intensities:Emotion IntensityHappy 0.7Relaxed 0.4Astonished 0.5Frustrated 0.6Sad 0.8Angry 0.9Then the normalized positive and negative valence values would be:Norm (positive)= Norm(astonished, happy,relaxed)= Astonished=(0.5 - 0.4) / (0.8 - 0.4)=.25Happy=(O. 7-0.4) / (0.7-0.4)=1.0Relaxed=(0.4-0.4) / (0.7-0.4)= 0Average= 0.42Norm (negative)= Norm(frustrated, sad, angry)Frustrated= (0.6-0.6) / (0.9-0.6)= 0Sad= (0.8-0.6) / (0.9-0.6)=0.67Angry= (0.9-0.6) / (0.9-0.6)=lAverage=0.56The relative maximum valence of the viewer during step 202b would be (0.42)-(0.56)= -0.14 (a valence of 0.14 in the negative direction, indicating that an overall negative valence was elicited from the user during step 202b).
Claims
CLAIMS1. A method of characterizing a viewer’s physiological response to perceived content, comprising:displaying, on a display device, pre-selected content associated with predetermined valence and activation levels;while displaying the pre-selected content, collecting first physiological response data from the viewer;determining, based on the first physiological response data, a baseline valence and baseline activation of the viewer;displaying, on the display device, content of interest to the viewer; and while displaying the content of interest:collecting second physiological response data from the viewer;determining a valence and activation of the viewer over a temporal window of the content of interest based on the second physiological response data, the baseline valence, and the baseline activation; and generating and outputting a dynamic visualization representing the viewer’s valence and activation for the temporal window of the content of interest.
2. The method of claim 1, wherein collecting first physiological response data from the viewer comprises collecting physiological data using a camera, a wearable device, or both.
3. The method of claim 1, wherein collecting first physiological response data from the viewer comprises collecting physiological data from the viewer using a wearable device, and wherein determining, based on the first physiological response data, a baseline valence and baseline activation of the viewer comprises processing wearable device data by performing one or more of filtering, artifact detection, peak detection, or deconvolution on the wearable device data.
234. The method of claim 1, wherein the first physiological response data comprises one or more of facial expression data, electrodermal activity data, or heart rate data.
5. The method of claim 1, wherein determining, based on the first physiological response data, the baseline valence of the viewer comprises:normalizing facial expression data for groups of positive and negative facial expressions; determining an average positive valence and an average negative valence based on the normalized facial expression data; andsubtracting the average positive valence from the average negative valence.
6. The method of claim 1, wherein determining, based on the first physiological response data, the baseline valence and baseline activation of the viewer comprises determining a baseline activation by:normalizing facial expression data for groups of facial expressions indicative of attentiveness and for groups of facial expressions indicative of lack of attentiveness;determining an average maximum valence from the normalized facial expression data for the group of facial expressions indicative of attentiveness;determining an average minimum valence from the normalized facial expression data for the group of facial expressions indicative of lack of attentiveness; and determining the baseline activation based on subtracting the average maximum valence from the average minimum valence.
7. The method of claim 1, wherein the dynamic visualization comprises a plot having a point for each temporal window of the content of interest, the point having a first coordinate location representing valence and a second coordinate location representing activation.
8. The method of claim 1, wherein the dynamic visualization comprises a plurality of zones, each of the plurality of zones representing a degree of the viewer’s immersion with the content of interest.
9. The method of claim 1 , wherein the dynamic visualization is outputted to a backend user.
10. The method of claim 1, further comprising:collecting the first and second physiological response data from a plurality of viewers in an audience; andgenerating analytics characterizing the audience’s physiological response to the content of interest.
11. The method of claim 10, further comprising outputting the analytics to a backend user.
12. A system comprising:one or more processors; andmemory storing computer program code executable by the one or more processors to cause the system to:display, on a display device, pre-selected content associated with predetermined valence and activation levels;while displaying the pre-selected content, collect first physiological response data from the viewer;determine, based on the first physiological response data, a baseline valence and baseline activation of the viewer;display, on the display device, content of interest to the viewer; and while displaying the content of interest:collect second physiological response data from the viewer; determine a valence and activation of the viewer over a temporal window of the content of interest based on the second physiological response data, the baseline valence, and the baseline activation; andgenerate and output a dynamic visualization representing the viewer’s valence and activation for the temporal window of the content of interest.
13. The system of claim 12, wherein collecting first physiological response data from the viewer comprises collecting physiological data using a camera, a wearable device, or both.
14. The system of claim 12, wherein collecting first physiological response data from the viewer comprises collecting physiological data from the viewer using a wearable device, and wherein determining, based on the first physiological response data, a baseline valence and baseline activation of the viewer comprises processing wearable device data by performing one or more of filtering, artifact detection, peak detection, or deconvolution on the wearable device data.
15. The system of claim 12, wherein the first physiological response data comprises one or more of facial expression data, electrodermal activity data, or heart rate data.
16. The system of claim 12, wherein determining, based on the first physiological response data, a baseline valence of the viewer comprises:normalizing facial expression data for groups of positive and negative facial expressions; determining an average positive valence and an average negative valence based on the normalized facial expression data; andsubtracting the average positive valence from the average negative valence.2617. The system of claim 12, wherein determining, based on the first physiological response data, a baseline valence and baseline activation of the viewer comprises determining a baseline activation by:normalizing facial expression data for groups of facial expressions indicative of attentiveness and for groups of facial expressions indicative of lack of attentiveness;determining an average maximum valence from the normalized facial expression data for the group of facial expressions indicative of attentiveness;determining an average minimum valence from the normalized facial expression data for the group of facial expressions indicative of lack of attentiveness; and determining the baseline activation based on subtracting the average maximum valence from the average minimum valence.
18. The system of claim 12, wherein the dynamic visualization comprises a plot having a point for each temporal window of the content of interest, the point having a first coordinate location representing valence and a second coordinate location representing activation.
19. The system of claim 12, wherein the dynamic visualization comprises a plurality of zones, each of the plurality of zones representing a degree of the viewer’s immersion with the content of interest.
20. The system of claim 12, wherein the dynamic visualization is outputted to a backend user.
21. The system of claim 12, wherein the system is further caused to:collect the first and second physiological response data from a plurality of viewers in an audience; andgenerate analytics characterizing the audience’s physiological response to the content of interest.2722. A non-transitory computer readable storage medium storing one or more programs, the one or more programs comprising instructions, which, when executed by a system comprising one or more processors and memory, cause the system to:display, on a display device, pre-selected content associated with predetermined valence and activation levels;while displaying the pre-selected digital content, collect first physiological response data from the viewer;determine, based on the first physiological response data, a baseline valence and baseline activation of the viewer;display, on the display device, content of interest to the viewer; and while displaying the content of interest:collect second physiological response data from the viewer; determine a valence and activation of the viewer over a temporal window of the content of interest based on the second physiological response data, the baseline valence, and the baseline activation; and generate and output a dynamic visualization representing the viewer’s valence and activation for the temporal window of the content of interest.