Attention tracking using sensors

By using sensors and AI technology to detect user reactions in real time and adjust content presentation, this technology solves the problem of not being able to track user attention in real time, thus improving user experience and interactivity.

CN121533028APending Publication Date: 2026-02-13DOLBY LABORATORIES LICENSING CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480045491.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-03-21
Filing Date
2024-07-01
Publication Date
2026-02-13

AI Technical Summary

Technical Problem

Existing user attention estimation methods cannot provide real-time information and cannot track user attention and interaction in real time during content consumption, resulting in a poor user experience.

Method used

By using a variety of sensors, such as cameras, microphones, and eye trackers, combined with deep neural networks and AI technology, the system can detect user reactions in real time and adjust content presentation accordingly, including volume, audio rendering position, and personalized delivery of advertising content.

Benefits of technology

It enables real-time user attention tracking, improves user experience, and allows for real-time adjustments to content presentation based on user feedback, enhancing interactivity and personalized services.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121533028A_ABST
    Figure CN121533028A_ABST
Patent Text Reader

Abstract

Some disclosed methods involve obtaining sensor data from a sensor system during presentation of content and estimating a user response event based on the sensor data. Some disclosed methods involve generating a user attention analysis based at least in part on an estimated user response event corresponding to an estimated user attention for a content interval of a content presentation. Some disclosed methods involve causing a content presentation to be altered based at least in part on a user attention analysis, and causing the altered content presentation to be provided.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application claims priority to International Patent Application No. PCT / CN2023 / 105695, filed July 4, 2023; U.S. Provisional Application No. 63 / 513,318, filed July 12, 2023; and U.S. Provisional Application No. 63 / 568,378, filed March 21, 2024, each of which is incorporated herein by reference in its entirety. Technical Field

[0002] This disclosure relates to apparatus, systems, and methods for estimating user attention levels and related factors based on signals from one or more sensors, and for responding to such estimation of user attention levels. Background Technology

[0003] Several methods, devices, and systems are known for estimating user attention (such as user attention to advertising content). Previously implemented methods for estimating user attention to media content involve assessing a person's rating of the content after they have consumed it (e.g., after watching a movie or an episode of a TV show, or after playing an online game). While existing devices, systems, and methods can offer benefits in some situations, improvements to these devices, systems, and methods are still desirable. Summary of the Invention

[0004] At least some aspects of this disclosure can be implemented via one or more methods. In some instances, the methods may be implemented at least in part by a control system and / or via instructions (e.g., software) stored on one or more non-transitory media. Some disclosed methods involve acquiring sensor data from a sensor system during content presentation and estimating user response events based on said sensor data. Some disclosed methods involve generating user attention analysis based at least in part on estimated user response events corresponding to estimated user attention to a content interval of content presentation. Some disclosed methods involve modifying content presentation based at least in part on user attention analysis and providing the modified content presentation.

[0005] Some or all of the operations, functions, and / or methods described herein can be executed by one or more devices according to instructions (e.g., software) stored on one or more non-transitory media. Such non-transitory media may include memory devices as described herein, including but not limited to random access memory (RAM) devices, read-only memory (ROM) devices, etc. Therefore, some innovative aspects of the subject matter described herein can be implemented via one or more non-transitory media on which software is stored.

[0006] At least some aspects of this disclosure can be implemented via one or more devices. For example, one or more devices (e.g., a system including one or more devices) may be able to perform at least partially the methods disclosed herein. In some embodiments, the apparatus is or includes an audio processing system having an interface system and a control system. The control system may include one or more general-purpose single-chip or multi-chip processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic, discrete hardware components, or combinations thereof. The control system may be configured to implement some or all of the methods disclosed herein.

[0007] In some examples, a system may include a host unit, a speaker system, a sensor system, and a control system. The control system may include one or more device analytics engines and a user attention analytics engine. The one or more device analytics engines may be configured to estimate user response events based on sensor data received from the sensor system. The user attention analytics engine may be configured to generate user attention analysis based at least in part on the estimated user response events received from the one or more device analytics engines. The user attention analysis may correspond to the estimated user attention for content regions presented via the host unit and speaker system.

[0008] The control system can be configured to modify the content presentation based at least in part on user attention analysis, and the modified content presentation is provided by a host unit, a speaker system, or a host unit and a speaker system.

[0009] In some examples, the system may also include an interface system configured to provide communication between the control system and one or more other devices via a network. In some examples, the modified content presentation may be, or may include, modified content received from one or more other devices via a network. According to some examples, modifying the content presentation may involve sending user attention analysis from a user attention analysis engine via the interface system and receiving modified content in response to the user attention analysis.

[0010] According to some examples, altering content presentation can involve personalizing or enhancing the presentation. In some examples, personalizing or enhancing content presentation can involve changing one or more of the following: audio playback volume, audio rendering position, one or more other audio characteristics, or a combination thereof.

[0011] In some examples, the host unit may be or may include a television. According to some such examples, personalizing or enhancing content presentation may involve changing one or more television display characteristics. In some examples, the host unit may be or may include a digital media adapter.

[0012] Based on some examples, personalizing or enhancing content presentation can involve changing the storyline, adding characters or other story elements, changing the time period involving the characters, changing the time period dedicated to another aspect of the content presentation, or a combination thereof.

[0013] In some examples, personalizing or enhancing content presentation may involve providing personalized advertising content. In some such examples, providing personalized advertising content may involve providing advertising content corresponding to estimated user attention to one or more content segments involving one or more products or services.

[0014] Based on some examples, personalizing or enhancing content presentation can involve providing or changing the laughter soundtrack.

[0015] In some examples, the host unit, speaker system, and sensor system may be located in the first environment. According to some such examples, the control system may be further configured to modify the content presentation based at least in part on sensor data corresponding to one or more other environments, estimated user response events, user attention analysis, or a combination thereof.

[0016] Based on some examples, the control system can be further configured to pause or replay the content.

[0017] In some examples, the sensor system may include one or more cameras. According to some such examples, the sensor data may include camera data.

[0018] According to some examples, the sensor system may include one or more microphones. According to some such examples, the sensor data may include microphone data. In some such examples, the control system may be further configured to implement an echo management system to mitigate the effects of audio played back by the speaker system and detected by one or more microphones.

[0019] In some examples, one or more first parts of the control system may be deployed in a first environment, and a second part of the control system may be deployed in a second environment. In some such examples, one or more first parts of the control system may be configured to implement one or more device analysis engines, and the second part of the control system may be configured to implement a user attention analysis engine.

[0020] Details of one or more embodiments of the subject matter described in this specification are set forth in the following figures and description. Other features, aspects, and advantages will become apparent from the description, figures, and claims. Note that the relative dimensions in the following figures may not be drawn to scale. Attached Figure Description

[0021] In the various figures, similar reference numerals and names indicate similar elements.

[0022] Figure 1A This is a block diagram illustrating examples of components of an apparatus capable of implementing various aspects of this disclosure.

[0023] Figure 1B An example environment is shown that includes components capable of implementing various aspects of this disclosure.

[0024] Figure 1C Another environment is shown, which includes examples of components capable of implementing various aspects of this disclosure.

[0025] Figure 2 An example component of an Attention Tracking System (ATS) is shown.

[0026] Figure 3A The components of an ATS residing in a playback environment are shown according to an example.

[0027] Figure 3B The diagram shows components of an ATS residing in the cloud, based on an example.

[0028] Figure 4 The components of a neural network capable of performing real-time acoustic event detection are shown, based on an example.

[0029] Figure 5 The diagram shows components of a device analytics engine configured to perform real-time attitude estimation, based on an example.

[0030] Figure 6 The components of an ATS according to another example are shown.

[0031] Figure 7 , Figure 8 , Figure 9 and Figure 10 An example of providing feedback from ATS during a video conference is shown.

[0032] Figure 11 An example of presenting feedback from ATS after a video conference is shown.

[0033] Figure 12 This is a flowchart outlining an example of the disclosed methods. Detailed Implementation

[0034] We currently spend a significant amount of time consuming media content (including but not limited to audiovisual content), interacting with media content, or a combination of both. (For simplicity and convenience, both consuming media content and interacting with media content can be referred to as "consuming" media content in this article.) Consuming media content can involve watching TV programs or movies, watching or listening to advertisements, listening to music or podcasts, playing games, video conferencing, participating in online learning courses, etc. Therefore, movies, online games, video games, video conferencing, advertisements, online learning courses, podcasts, streaming music, etc., can be referred to as types of media content in this article.

[0035] Previous methods for estimating user attention to media content such as movies and television programs did not consider how people reacted while consuming the content. Instead, user perception could be assessed based on their ratings of the content after consumption (e.g., after watching a movie or an episode of a television program, or after playing an online game). Current systems typically track what content a user selected, where they stopped the content, and which parts of the content were replayed. These metrics lack granular information about how consumers interact with content while consuming it. Current systems often do not know whether a user is present at any given time. For this reason, current user attention estimation methods typically fail to provide any real-time information. Such real-time information, in addition to real-time feedback on content replay, allows for more detailed analysis, which can improve the user experience.

[0036] It is beneficial to estimate one or more states of a person while they are consuming media content. These states may include or involve factors such as user attention, cognitive load, and level of interest. In addition to real-time feedback on content playback, this real-time information allows for more detailed analysis, which can improve the user experience.

[0037] Various publicly available examples overcome the limitations of previously implemented methods for estimating user attention. Some publicly available technologies and systems utilize available sensors to detect user responses or lack thereof in real time. Some such examples involve using one or more cameras, eye trackers, ambient light sensors, microphones, wearable sensors, or combinations thereof. Other examples involve measuring a person's engagement level, heart rate, cognitive load, attention, interest, etc., as they consume media content by watching television, playing games, participating in remote communication experiences (such as video conferencing, video seminars, etc.), listening to podcasts, etc. Recent advances in AI, such as automatic speech recognition (ASR), emotion recognition, and gaze tracking, have made such attention tracking systems possible. Furthermore, smart devices and various human-oriented sensors have become commonplace in our lives. For example, our smart speakers, phones, and televisions (TVs) have microphones, game consoles have cameras, and smartwatches have electrodermal responses. Combining some or all of these technologies allows for the implementation of enhanced user attention tracking systems.

[0038] Furthermore, the presence of multiple playback devices can lead to a richer playback experience, which may involve audio playback from satellite smart speakers, haptic feedback from mobile phones, and more. Coordinating multiple devices to play back audio can make it difficult to detect audio events using microphones due to echo. Therefore, some publicly available technologies and systems include descriptions of various methods for performing echo management.

[0039] In addition to receiving more data from sensors, advancements in deep neural networks (DNNs) have allowed for a deeper understanding of the meaning behind the data collected by these sensors. DNNs are now capable of performing tasks such as ASR (Aspect-Responsive Character Recognition), emotion recognition, and gaze tracking. The output of such models can serve as an indicator of real-time attention.

[0040] Some publicly available examples involve using real-time sensor data to adjust content presentation for users in real time. Some such examples may involve changing one or more aspects of media content in response to estimated attention, arousal, cognitive load, etc. Some publicly available examples involve aggregating real-time user attention information at both the user and content levels to provide insights, respectively, about future content options for users and content changes for future consumers.

[0041] Based on some examples, attention tracking systems consistent with the proposed technique can form the following feedback loop: - Content can be played back from one or more devices; - Users have a certain level of attention to the presented content; - Sensors detect indications of the user's attention; - The attention analysis engine determines the user's attention level in real time; and - Content presentation is adjusted in real time to improve the user experience.

[0042] In some examples, user attention analytics can be aggregated over time to determine how people interact with content presentations and users' affinity for different aspects of those presentations. These user attention analytics can still be used in conjunction with traditional attention detection methods.

[0043] An example application of real-time attention detection is a virtual laughter soundtrack. Whenever a user finds the content humorous, the virtual laughter soundtrack laughs along with the user. Short-term uses of real-time user feedback can include customized content selection and targeted advertising. In some examples, long-term analytics can be used to adjust content presentation to appeal to a wider audience, improve targeted user advertising, and so on.

[0044] What does attention mean? Throughout this disclosure, user attention, engagement, response, and reaction are used interchangeably. In some embodiments of the proposed techniques and systems, user response can refer to any form of attention to content, such as audible responses, body posture, gestures, heart rate, wearing content-related clothing or accessories, etc. Attention can take many forms, such as binary (e.g., a user saying "yes"), spectrum (e.g., excitement, loudness, leaning forward), or open-ended (e.g., the topic of discussion, multidimensional embedding). Attention can infer certain information related to the content presentation or objects within the content presentation. On the other hand, attention to non-content-related information can correspond to a low level of engagement with the content presentation.

[0045] Based on some examples, the attention to be detected can be in a short list such as that specified by any combination of user, content, content provider, user device, etc. (e.g., "wow", "ah", "red", "blue", slouching, leaning forward, left hand raised, right hand raised). It should be understood that such a short list is not required. A short list of possible responses (if provided by content rendering) can be reached through the metadata stream of content rendering.

[0046] Figure 1A This is a block diagram illustrating examples of components of an apparatus capable of implementing various aspects of this disclosure. As with the other figures provided herein, Figure 1AThe types, quantities, and arrangements of elements shown are provided by way of example only. Other implementations may include more, fewer, and / or different types, quantities, and arrangements of elements. According to some examples, device 150 may be configured to perform at least some of the methods disclosed herein. In some implementations, device 150 may be or may include one or more components of a workstation, one or more components of a home entertainment system, etc. For example, device 150 may be a laptop computer, tablet device, mobile device (such as a cellular phone), augmented reality (AR) wearable device, virtual reality (VR) wearable device, automotive subsystem (e.g., infotainment system, driver assistance or safety system, etc.), gaming system or console, smart home hub, television, or other types of devices.

[0047] According to some alternative implementations, device 150 may be or may include a server. In some such examples, device 150 may be or may include an encoder. In some examples, device 150 may be or may include a decoder. Thus, in some instances, device 150 may be a device configured for use in an environment such as a home environment; however, in other instances, device 150 may be a device configured for use in the “cloud,” such as a server.

[0048] According to some examples, device 150 may be or may include an orchestration device configured to provide control signals to one or more other devices. In some examples, the control signals may be provided by the orchestration device to coordinate aspects of displayed video content, audio playback, or a combination thereof. In some examples, device 150 may be configured to modify one or more aspects of media content currently provided by one or more devices in the environment in response to estimated user engagement, estimated user arousal, or estimated user cognitive load. Some examples are disclosed herein.

[0049] In this example, device 150 includes an interface system 155 and a control system 160. In some embodiments, the interface system 155 may be configured to communicate with one or more other devices in the environment. In some examples, the environment may be a home environment. In other examples, the environment may be another type of environment, such as an office environment, a car environment, a train environment, a street or sidewalk environment, a park environment, an entertainment environment (e.g., a theater, a performance venue, a theme park, a VR experience room, a video game arena), etc. In some embodiments, the interface system 155 may be configured to exchange control information and associated data with other devices in the environment. In some examples, the control information and associated data may be related to one or more software applications being executed by device 150.

[0050] In some implementations, the interface system 155 may be configured to receive or provide a content stream. In some examples, the content stream may include video data and corresponding audio data. The audio data may include, but is not limited to, audio signals. In some instances, the audio data may include spatial data such as channel data and / or spatial metadata. For example, the metadata may be provided by a device which may be referred to herein as an "encoder".

[0051] Interface system 155 may include one or more network interfaces and / or one or more external device interfaces (such as one or more Universal Serial Bus (USB) interfaces). According to some implementations, interface system 155 may include one or more wireless interfaces. Interface system 155 may include one or more devices for implementing a user interface, such as one or more microphones, one or more speakers, a display system, a touch sensor system, a motion sensor system, or a combination thereof. Therefore, although some such devices... Figure 1A While not explicitly stated, in some examples such a device may correspond to various aspects of interface system 155.

[0052] In some examples, interface system 155 may include control system 160 and memory system (e.g., Figure 1A One or more interfaces are shown between the optional memory system 165. Alternatively or additionally, in some instances, the control system 160 may include a memory system. In some embodiments, the interface system 155 may be configured to receive input from one or more microphones in the environment.

[0053] The control system 160 may include, for example, a general-purpose single-chip or multi-chip processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, discrete hardware components or combinations thereof.

[0054] In some implementations, the control system 160 may reside in more than one device. For example, in some implementations, a portion of the control system 160 may reside in a device within one of the environments described herein, and another portion of the control system 160 may reside in a device outside the environment, such as a server, game console, mobile device (e.g., a smartphone or tablet computer), etc. In other examples, a portion of the control system 160 may reside in a device within one of the environments depicted herein, and another portion of the control system 160 may reside in one or more other devices within the environment. For example, the functionality of the control system may be shared by an orchestration device (e.g., a device that may be referred to herein as a smart home hub) and one or more other devices within the environment. In other examples, a portion of the control system 160 may reside in a device implementing a cloud-based service (e.g., a server), and another portion of the control system 160 may reside in another device implementing a cloud-based service (e.g., another server, storage device, etc.). In some examples, the interface system 155 may also reside in more than one device.

[0055] In some embodiments, the control system 160 may be configured to perform at least partially the methods disclosed herein. Based on some examples, the control system 160 may implement one or more device analysis engines configured to estimate user response events based on sensor data received from a sensor system. In some examples, the control system 160 may implement a user attention analysis engine configured to generate user attention analysis based at least partially on estimated user response events received from one or more device analysis engines. The user attention analysis may correspond to an estimated user attention to a content range presented. In some examples, the control system 160 may be configured to modify content presentation based at least partially on the user attention analysis. According to some examples, the control system 160 may be configured such that the modified content presentation is provided by one or more displays, a speaker system, or one or more displays and a speaker system. Some examples of these components and processes are described below.

[0056] Some or all of the methods described herein can be executed by one or more devices according to instructions (e.g., software) stored on one or more non-transitory media. Such non-transitory media may include memory devices as described herein, including but not limited to random access memory (RAM) devices, read-only memory (ROM) devices, etc. One or more non-transitory media may, for example, reside on... Figure 1AIn the optional memory system 165 and / or control system 160 shown. Therefore, various innovative aspects of the subject matter described in this disclosure can be implemented in one or more non-transitory media on which software is stored. For example, the software may include instructions for controlling at least one device to perform some or all of the methods disclosed herein. For example, the software may be provided by, Figure 1A The control system 160 and other control systems are executed by one or more components.

[0057] In some examples, device 150 may include Figure 1A The optional microphone system 170 is shown. The optional microphone system 170 may include one or more microphones. According to some examples, the optional microphone system 170 may include a microphone array. In some examples, the microphone array may be configured to determine, for example, direction of arrival (DOA) and / or time of arrival (TOA) information based on instructions from the control system 160. In some instances, the microphone array may be configured to perform receive-side beamforming, for example, based on instructions from the control system 160. In some implementations, one or more microphones may be part of or associated with another device, such as a speaker in a speaker system, a smart audio device, etc. In some examples, the device 150 may not include the microphone system 170. However, in some such implementations, the device 150 may still be configured to receive microphone data from one or more microphones in the environment via the interface system 160. In some such implementations, a cloud-based implementation of the device 150 may be configured to receive microphone data or data corresponding to microphone data from one or more microphones in the environment via the interface system 160.

[0058] According to some embodiments, device 150 may include Figure 1A The optional speaker system 175 is shown in the diagram. The optional speaker system 175 may include one or more speakers, which may also be referred to herein as "speakers" or more generally as "audio reproduction transducers". In some examples (e.g., cloud-based implementations), the device 150 may not include the speaker system 175.

[0059] In some embodiments, device 150 may include Figure 1AThe optional sensor system 180 is shown. The optional sensor system 180 may include one or more touch sensors, motion sensors, motion detectors, cameras, eye-tracking devices, or combinations thereof. In some embodiments, the one or more cameras may include one or more standalone cameras. In some examples, one or more cameras, eye trackers, etc., of the optional sensor system 180 may reside in a television, mobile phone, smart speaker, laptop computer, game console or system, or a combination thereof. In some examples, device 150 may not include sensor system 180. However, in some such embodiments, device 150 may still be configured to receive sensor data from one or more sensors (such as cameras, eye trackers, camera-equipped monitors, etc.) residing in or on other devices in the environment via interface system 160. Although microphone system 170 and sensor system 180 are... Figure 1A While shown as a separate component, the microphone system 170 can be referred to as the sensor system 180 and can be considered part of the sensor system.

[0060] In some embodiments, device 150 may include Figure 1A The optional display system 185 is shown. The optional display system 185 may include one or more displays, such as one or more light-emitting diode (LED) displays. In some instances, the optional display system 185 may include one or more organic light-emitting diode (OLED) displays. In some examples, the optional display system 185 may include one or more displays of a television, laptop computer, mobile device, smart audio device, automotive subsystem (e.g., infotainment system, driver assistance or safety system, etc.), or another type of device. In some examples where device 150 includes display system 185, sensor system 180 may include a touch sensor system and / or motion sensor system for proximity to one or more displays of display system 185. According to some such embodiments, control system 160 may be configured to control display system 185 to present one or more graphical user interfaces (GUIs).

[0061] According to some such examples, device 150 may be or may include a smart audio device, such as a smart speaker. In some such implementations, device 150 may be or may include a wake word detector. For example, device 150 may be configured to (at least partially) implement a virtual assistant.

[0062] Figure 1B An example environment is shown, including components capable of implementing various aspects of this disclosure. Similar to the other figures provided herein, Figure 1BThe types, quantities, and arrangements of elements shown are provided as examples only. Other implementations may include more, fewer, and / or different types, quantities, and arrangements of elements.

[0063] In this example, environment 100A includes a host unit 101, which in this example is a television (TV). In some implementations, host unit 101 may be or may include a digital media adapter (DMA), such as an Apple TV™ DMA, Amazon Fire™ DMA, or Roku™ DMA. According to this example, content presentation is provided via host unit 101 and a speaker system including the TV's speakers and satellite speakers 102a and 102b. In this example, the attention levels of one or more of persons 105a, 105b, 105c, 105d, and 105e are detected by a camera 106 on the TV, microphones from satellite speakers 102a and 102b, a microphone from smart sofa 104, and a microphone from smart table 103.

[0064] In this example, the sensors of environment 100A are primarily used to detect auditory and visual feedback that can be detected by camera 106. However, in alternative implementations, the sensors of environment 100A may include additional types of sensors, such as one or more additional cameras, an eye tracker configured to collect gaze and pupil size information, one or more ambient light sensors, one or more thermal sensors, one or more sensors configured to measure skin conductance responses, etc. According to some implementations, one or more cameras in environment 100A (which may include camera 106) may be configured for eye tracker functionality.

[0065] Figure 1B The elements include: 101: Host unit 101, in this example a TV, which provides audiovisual content presentation and uses a microphone to detect user attention; 102a, 102b: Multiple satellite speakers that play back content and use a microphone array to detect attention; 103: A smart table that uses a microphone array to detect attention; 104: A smart sofa that uses a microphone array to detect attention; 105a, 105b, 105c, 105d, 105e: Multiple users concerned with the content in environment 100; and 106: Camera installed on host unit 101.

[0066] Figure 1C Another environment is shown, including examples of components capable of implementing various aspects of this disclosure. Similar to the other figures provided herein, Figure 1C The types, quantities, and arrangements of elements shown are provided as examples only. Other implementations may include more, fewer, and / or different types, quantities, and arrangements of elements.

[0067] In this example, Figure 1C This demonstrates a scenario of implementing an Attention Tracking System (ATS) in a car environment. Therefore, in this example, environment 100B is the car environment. Rear-seat passengers 105h and 105i are paying attention to the content on their respective displays 101c and 101d. Front-seat users 105f and 105g are paying attention to driving, display 101b, and audio content playing from speakers. The audio content may include music, podcasts, navigation directions, etc. The ATS utilizes all sensors in the car to determine the level of attention each user is paying to the content.

[0068] Figure 1C The elements include: 101b: The main display screen inside the vehicle, which includes satellite navigation, a reversing camera, music controls, etc. 101c, 101d: Passenger screens designed to play entertainment content such as movies; 105f, 105g, 105h, 105i: Multiple users interested in content within the vehicle; 106b: A camera facing the exterior of the vehicle, used to detect content such as billboards; 106c: An in-vehicle camera that detects the user's attention; 301d, 301e, 301f, 301g: Multiple microphones that pick up noise in the vehicle, including content playback, audio indicators of user attention, etc.; and 304d, 304e, 304f, 304g: Multiple speakers, for playback content.

[0069] Figure 2 An example component of an attention tracking system (ATS) is shown. Similar to other figures provided in this article, Figure 2 The types, quantities, and arrangements of elements shown are provided as examples only. Other implementations may include more, fewer, and / or different types, quantities, and arrangements of elements.

[0070] According to this example, content presentation 202 (also referred to as content playback 202) is provided via a device in environment 100C, and the user's reaction is then detected by one or more sensors of sensor system 180. In this example, Attention Analytics Engine (AAE) 201 is configured to estimate the user's attention level based on sensor data from one or more sensors. In this example, analysis of how one or more people 105 pay attention to the content presentation can be aggregated over time: here, content analytics and affinity module 203 is configured to estimate the affinity of individual users and groups for different aspects of the content presentation. These groups may correspond to specific demographic characteristics or demographic groups. According to this example, content presentation module 205 is configured to adjust the content presentation in real time based on attention analysis 206 received from AAE 201. Adjustments made by content presentation module 205 may include, for example, content selection, virtual laughter, and cheers. Various additional examples of adjustments that can be made by content presentation module 205 are disclosed herein. The real-time adjustments to the content by the content presentation module 205 and the corresponding user feedback detected by the sensor system 180 can create a feedback loop, thereby providing a better experience for one or more people 105.

[0071] Figure 2 The elements include: 105: One or more people who are watching content replay 202, and their reactions are detected by sensors. 180: A sensor system comprising one or more sensors configured to detect real-time user responses to content presentation 202; 201: Attention Analysis Engine (AAE), which analyzes user attention information for the current content by acquiring data from sensor system 180 (or, in some implementations, from one or more Device Analytics Engines (DAEs) that are generated using measurements from the sensors).

[0072] 202: Playback of content provided by content presentation module 205, which is performed through any combination of speakers, displays, lights, etc.; 203: Content analytics and user affinity module, configured for the aggregation of attention analytics to provide insights into user attention to content playback 202; and 205: Content presentation module, which is configured to provide the content to be played and to adjust the content presentation in real time based on attention analysis 206 provided by AAE 201.

[0073] exist Figure 2In the example shown, the attention tracking system (ATS) 200 includes an AAE 201, a content analysis and user affinity module 203, a sensor system 180, and a content presentation module 205. In this example, the AAE 201, the content analysis and user affinity module 203, and the content presentation module 205 are... Figure 1A Example implementation of control system 160. In some examples, AAE 201, content analysis and user affinity module 203, and content presentation module 205 may be implemented as instructions (e.g., software) stored on one or more non-transitory computer-readable media.

[0074] According to some examples, ATS includes an attention analysis engine 201, sensors, and one or more device analysis engines. Figure 2 (Not shown in the diagram, each device analytics engine corresponds to each sensor or sensor group), all operating in real time. In some such examples, the ATS is configured to: - Collect information from available sensors; - Transmit sensor information through one or more device analytics engines; and - User attention is determined by passing the results of the device analytics engine through the attention analytics engine (201).

[0075] The output of ATS can vary depending on the specific implementation. In some examples, the output of ATS can be any type of attention analysis-related information, content presentation corresponding to the attention analysis-related information, etc. According to some examples, ATS can be implemented via an attention analysis engine 201 and content presentation module 205 residing in a device within a playback environment (e.g., implemented by one of the playback devices in the playback environment), or via an attention analysis engine 201 and content presentation module 205 residing in the "cloud" (e.g., implemented by one or more servers not within the playback environment), etc.

[0076] Figure 3A The diagram illustrates components of an ATS residing in a playback environment, based on an example. Similar to the other figures provided herein, Figure 3A The types, quantities, and arrangements of elements shown are provided as examples only. Other implementations may include more, fewer, and / or different types, quantities, and arrangements of elements.

[0077] exist Figure 3AIn this example, the ATS 200 includes an Attention Analysis Engine (AAE) 201 and a Content Presentation Module 205 residing in a host unit 101, which in this example is a television (TV). Multiple sensors, including microphones 301a of the host unit 101, microphones 301b of satellite speakers 102a and 301c of satellite speakers 102b, provide sensor data to the AAE 201 corresponding to events occurring in the environment 100D. Other implementations may include additional types of sensors, such as one or more cameras, one or more eye trackers configured to collect gaze and pupil size information, one or more ambient light sensors, one or more thermal sensors, one or more sensors configured to measure skin conductance responses, etc. As indicated by point and component number 102c, some implementations may include three or more satellite speakers.

[0078] Microphones 301a to 301c will detect the audio of the content presentation. In this example, echo management modules 302a, 302b, and 302c are configured to suppress the audio of the content presentation, thereby allowing for more reliable detection of sound corresponding to the user's reaction to the content in the signals from microphones 301a to 301c. In this example, content presentation module 205 is configured to send echo reference information 306 to echo management modules 302a, 302b, and 302c. Echo reference information 306 may, for example, contain information about the audio played back by speakers 304a, 304b, and 304c. As a simple example, local echo paths 307a, 307b, and 307c can be canceled using local echo references via a local echo canceller. However, any type of echo management system, such as a distributed acoustic echo canceller, can be used here.

[0079] according to Figure 3AIn the example shown, host unit 101 includes a Device Analysis Engine (DAE) 303a, satellite speaker 102a includes a DAE 303b, and satellite speaker 102b includes a DAE 303c. Here, DAEs 303a, 303b, and 303c are configured to detect user activity from sensor signals, in these examples, microphone signals. The DAE 303 can exist in different implementations, even within the same attention tracking system in some instances. A specific implementation of the DAE 303 can, for example, depend on the sensor type or mode. For instance, some implementations of the DAE 303 can be configured to detect user activity from microphone signals, while other implementations of the DAE 303 can be configured to detect user activity, attention, etc., based on camera signals. Some DAEs can be multimodal, thus receiving and interpreting input from different sensor types. In some examples, a DAE can share sensor input with other DAEs. The output of the DAE 303 can also vary depending on the specific implementation. DAE outputs can include, for example, detected phonemes, emotion type estimates, heart rate, body posture, and latent spatial representations of sensor signals.

[0080] exist Figure 3A In the illustrated implementation, outputs 309a, 309b, and 309c from DAEs 303a, 303b, and 303c, respectively, are fed into AAE 201. Here, AAE 201 is configured to combine information from DAEs 303a to 303c to generate attention analysis. AAE 201 can be configured to use various types of data refinement techniques, such as neural networks, algorithms, etc. For example, AAE 201 can be configured to use natural language processing (NLP) with speech recognition from one or more DAE outputs. In this example, the analysis generated by AAE 201 allows content presentation module 205 to adjust content presentation in real time. Content presentation 310 is then provided to and played from actuators, in this example, the actuators including speakers 304a, 304b, and 304c and a TV display 305. Other examples of actuators that can be used include lighting and haptic feedback devices.

[0081] Figure 3A The elements include: 301a, 301b, 301c: Microphones that pick up sound in the environment 100D, including content playback, audio corresponding to user responses, etc. 302a, 302b, 302c: Echo management modules, which are configured to reduce the playback level of content picked up by the microphone; 303a, 303b, 303c: Device analytics engine, which can be configured to convert sensor readings into probabilities of certain attentional events (such as laughter, panting, or cheering); 304a, 304b, 304c: Speakers in each device that reproduce audio from the content presentation module 205; 305: TV display, configured to display content from content presentation module 205; 306a, 306b, 306c: Echo reference information; and 307a, 307b, 307c: Echo paths between multiple speakers and multiple microphones of a duplex device.

[0082] In these examples, AAE 201, content rendering module 205, echo management module 302a, and DAE 303a are... Figure 1A The control system 160 is implemented in instance 160a, the echo management module 302b and DAE 303b are implemented in another instance 160b of the control system 160, and the echo management module 302c and DAE 303c are implemented in a third instance 160c of the control system 160. In some examples, AAE 201, content presentation module 205, echo management modules 302a to 302c, and DAE 303a to 303c may be implemented as instructions (e.g., software) stored on one or more non-transitory computer-readable media.

[0083] Figure 3B The diagram illustrates components of an ATS residing in the cloud, based on an example. Similar to the other diagrams provided in this article, Figure 3B The types, quantities, and arrangements of elements shown are provided as examples only. Other implementations may include more, fewer, and / or different types, quantities, and arrangements of elements.

[0084] In this example, the ATS 200 includes components from... Figure 1A The control system instance 160d of the control system 160 implements AAE201 and content presentation module 205. The control system instance 160d resides in cloud 308. Cloud 308 may, for example, include one or more servers residing outside environment 100E but configured to communicate via a network with host unit 101 and at least satellite speakers 102a and 102b. Echo reference information 306 and local echo path 307 may be as follows: Figure 3A As shown, but it has already been from Figure 3B The echo management system is removed to reduce the visual complexity of the graph. Those skilled in the art will understand that any type of echo management system can be used here. This implementation is similar to... Figure 3A The two key differences between the implementation methods shown are: - DAE results 309a, 309b, and 309c are sent via the network from DAE 303a, 303b, and 303c to AAE 201 in cloud 308; and - Content presentation module 205 is also implemented in cloud 308 and configured to provide content 310 to devices in environment 100E via the network.

[0085] Acoustic event detection Figure 4 The diagram illustrates components of a neural network capable of performing real-time acoustic event detection, based on an example. Similar to the other figures provided herein, Figure 4 The types, quantities, and arrangements of elements shown are provided as examples only. Other implementations may include more, fewer, and / or different types, quantities, and arrangements of elements.

[0086] Figure 4 An example component of a neural network 400 capable of implementing a device analysis engine 303 for a microphone 301 is shown. In this example, the neural network 400 is implemented by a control system instance 160e. According to this example, a microphone signal 410 from the microphone 301 is passed to a segmentation block 401 to create a time-frequency representation of the audio in the microphone signal 410. The resulting frequency band 412 is passed to two-dimensional convolutional layers 402a and 402b, which are configured for feature extraction and downsampling, respectively. In this example, positional encoding 403 is stacked on the features 411 output by the convolutional layers 402a and 402b, so that a real-time streaming transformer 404 can take temporal information into account. The embedding 414 produced by the transformer is projected (using a fully connected layer 405) onto a desired number of unit scores 406. The unit scores can represent anything related to an acoustic event, such as word units, phonemes, laughter, cheers, gasps, etc. According to this example, the SoftMax module 407 is configured to normalize the unit score 406 into a unit probability 408 representing the posterior probability of an acoustic event.

[0087] Other examples of attention-related sound events that can be detected and used in the proposed techniques and systems include: -Possible sounds indicating input: laughter, screaming, cheering, hissing, crying, sobbing, groaning, making "oh," "ah," "shh" sounds, talking content, cursing, etc.; - Sounds that may indicate a lack of engagement: typing, creaking doors, snoring, footsteps, vacuuming, washing dishes, chopping vegetables, talking about things unrelated to the content, or different content played on a device not connected to the attention system; - Sounds that can be used to indicate attention to a specific type of content, such as: oMovie: Name the actors; o Sports: Name the player or team; o When the content contains music: whistling, clapping, stomping, snapping fingers, singing along with the content, and people making repetitive noises in rhythm with the content; o In children's programs: Children make emotional sounds or respond to "call and answer" prompts; Fitness-related content: snoring, heavy breathing, groaning, panting; - Other noises that can help infer attention to content based on its context: such as silence during dramatic time intervals, mentioning objects, characters, or concepts in a scene, etc.

[0088] Figure 4 The elements include: 400: Example neural network architecture for a real-time audio event detector; 303d: Device analytics engine, configured to detect attention-related events in microphone input data; 401: A segmented block, which is configured to process time-domain input into segmented time-frequency domain information; 402a, 402b: Two-dimensional convolutional layers; 403: Position encoding, which stacks positional information onto feature 411 output by convolutional layers 402a and 402b; 404: Multiple (six in this example) real-time streaming converter layers; 405: Fully connected linear layer, which is used as a projection of the cell score 414 output by the real-time streaming converter layer 404; 406: Represents the unit score for different audio event categories. Unit scores can represent audio events such as laughter, panting, cheering, etc.

[0089] 407: SoftMax module 407, which is configured to normalize the unit score 406 into unit probabilities 408 representing the likelihood of an acoustic event; and 408: The obtained unit probability.

[0090] Visual inspection Figure 5 This illustrates components of a device analytics engine configured to perform real-time pose estimation, based on an example. Similar to other figures provided herein, Figure 5 The types, quantities, and arrangements of elements shown are provided as examples only. Other implementations may include more, fewer, and / or different types, quantities, and arrangements of elements.

[0091] Figure 5An example component of a device analytics engine (DAE) 303e configured to estimate user attention based on visual information is shown. In this example, the DAE 303e is implemented by a control system instance 160f. The DAE 303e can be configured to estimate user attention based on visual information via a range of techniques, depending on the specific implementation. Some examples involve applying machine learning methods, using algorithms implemented in software, etc.

[0092] exist Figure 5 In the example shown, DAE 303e includes a skeleton estimation module 501 and a pose classifier 502. The skeleton estimation module 501 is configured to calculate the positions of the main human skeleton from camera data 510 (including a video feed in this example) and output skeleton information 512. The skeleton estimation module 501 can be implemented using publicly available toolkits such as YOLO-Pose. The pose classifier 502 can be configured to implement any suitable process for mapping the skeleton information to pose probabilities, such as a Gaussian mixture model or a neural network. According to this example, DAE 303e (in this example, pose classifier 502) is configured to output pose probabilities 503. In some examples, DAE 303e may also be configured to estimate the distances to one or more parts of the user's body based on camera data.

[0093] Visual detection can reveal a range of attentional information. Some examples include: - Visual expressions that can indicate active interaction: leaning forward, leaning back, moving in response to events in the content, wearing clothing that symbolizes loyalty to something in the content, etc. - Visual expressions that can indicate negative interactions: a disgusted expression, a gesture of raising the middle finger, etc. - Visual signs that indicate a lack of interaction: a person looking at their phone when it's not being used to present content; a person holding their phone to their ear when it's not being used to present content; a person falling asleep; no one in the room paying attention; no one being present, etc.

[0094] Figure 5 The elements include: 303e: Device analytics engine, which is configured to perform real-time pose estimation based on camera data 510; 501: Skeletal estimation module, which is configured to calculate the position and rotation of the main skeleton of a person from camera data 510 and output skeleton information 512; 502: A pose classifier configured to map skeletal information 512 to pose probabilities 503; and 503: The obtained pose probability.

[0095] Systems and methods for measuring and indicating user attention during video interaction This section describes numerous concepts designed to distill attention metrics from sensor data captured during video consumption (e.g., watching videos such as movies, social media posts, and participating in video conferences) and then present these attention metrics as feedback via a graphical interface. This section includes technologies for providing real-time feedback during meetings and offline feedback through reports.

[0096] attention model This section describes the collection of data via sensors, which are typically included in display systems (accelerometers, IMUs, microphones, etc.) or devices on or around the user (e.g., various sensors on smart devices such as smartwatches). An attention analysis model is used to analyze all inputs from sensors, user interface interactions, and camera image processing.

[0097] Figure 6 The components of an ATS according to another example are shown. As with the other figures provided in this article, Figure 6 The types, quantities, and arrangements of elements shown are provided as examples only. Other implementations may include more, fewer, and / or different types, quantities, and arrangements of elements.

[0098] According to this example, the ATS 200 takes independent measurements from various types of sensors in the sensor system 180 as input and outputs one or more refined attention metrics. Figure 6 In the example shown, ATS 200 includes Device Analysis Engines (DAEs) 303f, 303g, 303h, 303i, and 303j, as well as AAE 201, all implemented by control system instance 160g. According to this example, each of DAEs 303f through 303j receives and processes sensor data from different types of sensors in sensor system 180: DAE 303f receives and processes camera data from one or more cameras, DAE 303g receives and processes motion sensor data 602 from one or more motion sensors, DAE 303h receives and processes wearable sensor data 604 from one or more wearable sensors, DAE 303i receives and processes microphone data 410 from one or more microphones, and DAE 303j receives and processes user input 606 from a user interface. In this example, DAE 303f produces DAE output 309f, DAE 303g produces DAE output 309g, DAE 303h produces DAE output 309h, DAE 303i produces DAE output 309i, and DAE 303j produces DAE output 309j.

[0099] In this example, AAE 201 is configured to implement one or more types of data refinement methods, data combination methods, or both on DAE outputs 309f to 309j. In some examples, data refinement and / or combination can be performed by implementing a fuzzy inference system. Other data refinement and / or combination methods may involve implementing support vector machines or neural networks. According to some examples, AAE 201 can be configured to implement an attention analysis model tuned or trained on ground truth for subjective measurements. In some examples, AAE 201 can be configured to implement an attention analysis model tuned or trained using research training data collected in a research lab environment. Such data may not be available in a consumer environment. Research training data may, for example, include sensor data from a wider array of sensors than those typically used in consumer environments, such as all the sensors disclosed herein. In some examples, research training data may include measurements not typically available to consumers, such as EEG records.

[0100] According to this example, AAE 201 is configured to output an estimated attention score 610 and an individual estimated attention metric 612. In this example, the estimated attention score 610 is a combined attention score that includes the responses of all people currently consuming the content presentation, and the individual estimated attention metric 612 is the individual attention score of each person currently consuming the content presentation.

[0101] Figure 7 , Figure 8 , Figure 9 and Figure 10 An example of providing feedback from ATS during a video conference is shown. Figure 11 An example is shown of feedback being presented from the ATS after a video conference. In some instances, the ATS can, as... Figure 6 As shown, or a similar ATS can be used. According to some examples, an ATS with fewer sensor types and fewer DAEs can be provided based on the input to the AAE 201. Figures 7 to 11 The feedback is shown. Similar to other figures provided in this article, Figures 7 to 11 The types, quantities, and arrangements of elements shown are provided as examples only. Other implementations may include more, fewer, and / or different types, quantities, and arrangements of elements.

[0102] Figure 7 An example of feedback being provided from ATS during a one-to-many video conference presentation is shown. Figure 7Examples of various windows that can be displayed on a presenter's monitor during a video conference presentation are shown. Window 705 displays the presenter's own video, window 710 displays the content being presented, and windows 715a, 715b, and 715c display the videos of non-presentation participants. In this example, window 720 displays a graph indicating three types of ATS feedback: curve 725 indicates estimated arousal, curve 730 indicates estimated attention, and curve 735 indicates estimated cognitive load. According to this example, these three types of ATS feedback are aggregated and based on data from all non-presentation participants.

[0103] Figure 8 An example of feedback being provided from ATS during a many-to-many video conference discussion is shown. Figure 8 An example of a window that can be presented on each participant's monitor during a video conference discussion is shown. Window 805 shows the video of each participant. In this example, each of windows 805 includes an attention score 810, which indicates the estimated attention of that individual participant. In some examples, the ATS (Attention Score) can estimate more than one type of feedback. In some such examples, the attention score 810 can indicate aggregated ATS feedback. In other examples, multiple attention scores 810 can be presented in each of windows 805. In some alternative examples, the attention score 810 may not be presented on each participant's monitor during a video conference discussion. In some such examples, the attention score 810 may only be presented on the monitors of a subset of participants (e.g., one participant's monitor) during a video conference discussion. The subset of participants may include, for example, supervisors, moderators, etc.

[0104] Figure 9 This shows another example of feedback being provided from the ATS during a many-to-many video conference discussion. Figure 9 An example is shown of windows that can be displayed on each participant's monitor or on a subset of participants' monitors during a video conference discussion. Window 905 displays the video of each participant. In this example, window border 910a is shown with the thickest outline, indicating that the current speaker is shown in the corresponding window 905a. According to this example, window borders 910b and 910c are shown with outlines that are thinner than those of window border 910a but thicker than those of the other windows 905, indicating that the participants shown in windows 905b and 905c are currently looking at the current speaker.

[0105] Figure 10 This shows another example of feedback being provided from the ATS during a many-to-many video conference discussion. Figure 10An example is shown of windows that can be displayed on each participant's monitor or on a subset of participants' monitors during a video conference discussion. Window 1005 displays the video of each participant. In this example, window boundary 1010a is shown with the thickest outline, indicating that the current speaker is shown in the corresponding window 1005a. According to this example, window boundaries 1010b and 1010c are shown with outlines that are thinner than those of window boundary 1010a but thicker than those of the other windows 1005, indicating that the ATS estimates that the participants shown in windows 1005b and 1005c have a higher level of attention.

[0106] Figure 11 An example of feedback being provided from ATS following a many-to-many video conference discussion is shown. Figure 11 The table shown represents the ATS feedback for each of the six video conference participants. In this example, row 1105 indicates the percentage of each participant's speaking time, and column 1110 indicates each participant's individual attention score. The inner cells indicate the individual participants' attention to each other and to other stimuli. For example, cell 1115a indicates user 1's attention to user 2, and cell 1115b indicates user 6's attention to user 5. The cells in column 1120 indicate the individual participants' attention to stimuli or events other than those of other video conference participants. For example, cell 1115c indicates user 6's attention to stimuli or events other than those of other video conference participants.

[0107] Attention Feedback Loop Any decision made based on information provided by the ATS can have its effectiveness evaluated using the ATS to detect user responses. In some examples, this can form a closed loop where decisions made using ATS information can improve over time. According to some examples, this closed loop of user attention and responses to user attention can be taken even further, and once a decision is known to have a specific effect on a user, it can be passed on to another user to experience.

[0108] Specific examples: - One or more users are detected laughing while watching a comedy. This laughter is detected by the user's (multiple) ATS-enabled devices. Using this information, additional laughter is added to the comedy (e.g., by the content presentation module 205 according to instructions from AAE 201). The ATS detects that this adds to the user's enjoyment, so the system (e.g., by the content presentation module 205 according to instructions from AAE 201) continues to add laughter as the ATS detects that the user is laughing at the comedy. Over time, the system begins to optimize the types of laughter added that each user finds most amusing.

[0109] - It is detected that one or more users laughed at a comedy ad after their TV program ended. This laughter was detected by the user's (multiple) ATS-enabled devices. Using this information, it is assumed (e.g., by AAE 201) that the user will engage more with the comedy ad after the program ends. This is detected to be true for the user, so the system (e.g., content presentation module 205 according to instructions from AAE 201) continues to play the comedy ad after the TV program ends. Over time, the system begins to optimize the type of comedy ad that each individual user finds most interesting for each user. In some examples, the affinity of an individual user or group for a comedy ad after their TV program can be determined through multiple viewing sessions. In some such examples, this affinity can be determined by content analysis and user affinity block 203, while in other examples, it can be determined by AAE 201.

[0110] - It was detected that most viewers found an unintentionally funny moment in the content more entertaining than the rest of the comedy. This enjoyment was detected by (multiple) ATS-enabled devices used by the user. Using this information, more similar humorous moments were added to the same comedy. It was detected that this increased the overall enjoyment for the user. Over time, the system began to optimize how much of these types of jokes should be placed in the content.

[0111] - It detected that one or more users were laughing at a comedy at home. This laughter was detected by (multiple) ATS-enabled devices of the user. Using this information, additional laughter was added to the comedy. It was detected that this added to the user's enjoyment, so the system began adding laughter when it heard the user laughing at a comedy podcast in their car. Over time, the system began optimizing for each individual user, focusing on which types of locations and media they found most entertaining.

[0112] The option to evaluate the effectiveness of decisions made and optimize decisions based on attention responses applies to all use cases listed in this disclosure.

[0113] Personalization and Enhancement Use Cases Using Attention Feedback Modern personalization systems for content enjoyment rely on information explicitly provided by users through content control systems such as play and pause. For example, current content recommendation systems make decisions based on what content users choose to consume, which parts they replay, and whether they like a particular piece of content. This offers little insight into how users engage with content or whether they were present during replays. Furthermore, slow personalization might take the form of fan feedback, through explicitly leaving comments or ratings about content. This allows content creators to tailor the next piece of content for their audience, but the feedback loop is slow and only represents those users who left comments.

[0114] The possibilities for personalized systems are greatly expanded by using sensors such as microphones, cameras, and accelerometers to utilize richer metrics of user attention in real time. Having such attention-tracking systems allows for the identification of content that interests a user. For example, when a user is watching a secret agent movie, they might show particular interest in a suit, watch, car, location, or action. Insights into these kinds of details will allow for improved personalization systems.

[0115] Adjusting content enjoyment with richer attention information can be done in real-time and / or over a long period by aggregating attention metrics. These detailed attention metrics can also lead to more informed recommendations and provide content creators with real-time feedback on how users are paying attention to content. When making any personalized decisions using attention information, attention tracking systems can be leveraged again to determine the impact on the attention of one or more users. Continuous use of attention tracking systems can form a closed loop in which many types of personalization can be optimized, or at least improved. In this section, we will detail some of the possibilities and use cases for personalization systems that utilize sensors to employ richer attention metrics in real time.

[0116] As used in this article, the phrase “content enjoyment experience” can refer to any factor that may affect a user’s experience when enjoying content, such as changing playback (e.g., volume, TV backlight level), content (e.g., selected content, changing the storyline, adding elements), and control systems (e.g., pause, replay).

[0117] As used in this article, "personalization" can refer to any content enjoyment experience tailored to a user. "Enhancement of experience" refers to real-time changes to the content enjoyment experience (including but not limited to real-time changes). "Recommendation" refers to any content suggested to a user. Personalization, enhancement of experience, and recommendation can all be based on attention information. These terms will be described in more detail in their respective sections.

[0118] “Linear content” refers to any content format designed to be played from beginning to end without diverging paths or controlling the flow of content (e.g., an episode of a Netflix series, an audiobook, a song, a podcast, or a TV show).

[0119] Detectable attention list and metadata This section provides a more detailed look at the short list of potential detectable attentions briefly described in the introduction. The scope of detectable attentions can be specified in the list. Examples of attention types that can appear in this list can be one of the following forms: - A specific response in which the user does exactly what they want. For example, the user says "yes," "no," or something that doesn't match, or it can be detected that the user raises their left hand, right hand, or neither hand is raised.

[0120] - A response type where the reaction matches the response type to some extent. For example, you might ask a user to "start moving," and the target attention type is the user's movement. In this case, detecting the user twisting their body would be a strong match. Another example might involve content saying, "Are you ready?", and the content is looking for an affirmative response. There are many valid user responses to indicate affirmation, such as "yes," "of course," "let's begin," or a nod.

[0121] - Emotional responses, where a user's emotion or a subset of their emotions is detected. For example, a content provider wants to understand a consumer's emotion towards their latest post. They decide to add emotions to a short list of detectable attention. Users consuming content begin a conversation about the content, and only their emotions are inferred as attributions for their emotional responses. Another example involves a user who only wants to share their emotional level on the arousal dimension. When a user feels disgusted by the content they are watching, their disgust is not detected. However, a low level of arousal is reflected in the attention detection.

[0122] - The topics discussed, where ATS determines what topics the response content generated. For example, content producers want to know what issues their films raise with viewers. After listing the topics to be discussed as options in the attention list, they found that people typically talk about how funny the film was or about global warming.

[0123] There can be even more types of attention. The list of attention can be provided from a range of different providers, such as device manufacturers, users, content producers, etc. If multiple detectable attention lists are available, any method of combining these lists can be used to determine the resulting list, such as using only the user list, the union of all lists, the intersection of the user list and the content provider list, etc.

[0124] The detectable attention list can also provide users with a degree of privacy, allowing them to provide their own lists of detectable content and their own rules for how their lists are combined with external parties. For example, users can be given phased options on what content they should be detected (e.g., through a graphical user interface (GUI) and / or one or more audio cues), choosing to detect only emotions and specific responses. This gives users peace of mind when using their (multiple) ATS-enabled devices.

[0125] A list of detectable attention indicators can reach a user's device in several ways. Two examples include: - A list of detectable attention indicators is provided to the device in the content's metadata stream.

[0126] - A list of detectable attention indicators is pre-installed on (multiple) ATS-enabled devices for users. This list can be applied to a wide range of content and user attention indicators. In some examples, users are able to select from these detectable attention indicators, such as those described above.

[0127] The list of detectable attention indicators associated with content segments can be learned from users whose ATS-enabled devices detect a larger set of attention indicators. In this way, content providers can discover how users are paying attention to their content and then add those attention indicator types to the list they want to detect for users with a more limited set of detectable attention indicators. In some examples, there might be an upstream connection along the content stream, allowing this learned metadata to be sent to the cloud for aggregation. This is in... Figure 2 This was briefly discussed in Content Analysis and User Affinity Block 203.

[0128] The option to have a list of detectable attention indicators applies to all use cases listed in the 'Use Case List' section.

[0129] Example use cases Long-term personalization We use "personalization" to refer to a way of tailoring experiences to users based on long-term analysis. The effectiveness of personalization adjustments can be evaluated by testing how they change the user experience (using (multiple) ATS-enabled devices). This forms a closed loop, allowing personalization adjustments to be continuously improved and user preferences tracked. In some alternative examples, personalization adjustments can be applied in an open-loop manner, where the effects of the adjustments are not measured. Changes applied to the experience don't always have to follow user preferences to provide natural variation and avoid creating isolated information cocoons for users (such as political segregation).

[0130] Determine user preferences Suppose one or more users consume a series of content on playback devices equipped with ATS (Active Content Service). In some examples, each user's content-related preferences can be determined over time by aggregating results from the ATS. The range of preferences that can be tracked can be broad and may include content type, actors, themes, effects, topics, locations, etc. User preferences can be determined in the cloud or on the user's(s) ATS-enabled devices. In some instances, the terms "user preference," "interest," and "affinity" are used interchangeably.

[0131] Before long-term user preference aggregation is available, short-term estimates of what users are interested in can be built. These short-term estimates can be made using recent attention information and hypothesis testing using attention feedback loops.

[0132] Personalize content based on user preferences Content can be personalized based on user preferences. In some examples, when personalization references the preferences of multiple users, the adjustments can be optimized in a way that takes all users into account. Some example methods for personalization based on the preferences of multiple users include calculating one or more average attention-related values, determining the maximum attention-related value for the group, determining learned combinations of preferences (e.g., via neural networks), etc. Attention-related values ​​can correspond to user preferences, including but not limited to predetermined / previously known user preferences.

[0133] In some examples, personalization tweaks (e.g., alternative scenarios) can be delivered in the content stream as different playback options (e.g., as user-selectable playback options). Alternatively, only the personalized version of the content is streamed to the user. A different option is to generate personalization tweaks, such as via a neural network. The generation of tweaks can occur on the user's (multiple) ATS-enabled devices or in the cloud, depending on the specific implementation.

[0134] Different forms of personalization will be detailed in the following sections.

[0135] The storyline of linear content is changed based on user preferences. An example of extending "personalizing linear content based on user preferences" is adjusting the content's storyline to suit user preferences. The storyline can be adjusted by replacing, adding, or removing segments within the content. The content's storyline can also be tailored to the preferences of multiple users. This can be useful, for example, when many people are watching content together in the same viewing session. In some such examples, the storyline can be optimized by using a blend of their preferences to improve the overall enjoyment for the group.

[0136] Specific examples: Bob dislikes gore and is easily disturbed. He watched a war movie that included a scene of soldiers having limbs amputated. The gory scene was replaced for Bob with a shorter one that only showed the amputee's face.

[0137] Alice was identified as a Stephen Curry fan because multiple ATS-enabled devices equipped with cameras detected that her gaze primarily followed Curry while watching a game. In the next video she watched, the content was personalized using generative technology, incorporating a cameo appearance by Stephen Curry.

[0138] John has a short attention span and therefore prefers shorter content. This was detected by his (multiple) ATS-enabled devices, and in the videos he watched next, lengthy scenes were cut into shorter clips.

[0139] A parent and child are watching a movie together. The child is easily startled, so a sudden frightening scene was removed. The parent prefers movies without happy endings, so the film's originally happy ending was left unresolved.

[0140] Generative enhancement of content based on user preferences Another example of extending "personalizing content based on user preferences" is using generative elements to enhance content. Generative elements refer to machine-generated media, which can include video, audio, lighting, or a combination thereof. We use "enhance" to mean changing aspects of the content (e.g., objects in a scene, musical instruments being played, lighting in a scene, a narrator's voice). Some examples include overlaying generated images to replace objects in a scene (e.g., turning an apple into an orange) or changing a sound from one type to another (e.g., changing strings into horns).

[0141] Specific examples: John was watching a new movie, *Top Gun: Maverick*, starring Tom Cruise as Captain Pete Mitchell. John is a huge Brad Pitt fan and would have preferred Pitt to play Mitchell. Using generative machine learning, Brad Pitt's face was superimposed onto Tom Cruise's face throughout the film. Tom Cruise's voice was also replaced using generative technology to make it sound like Brad Pitt. From John's perspective, Mitchell is now played by Pitt, not Cruise.

[0142] Personalized options for selecting linear content based on user preferences Another extension of "personalizing linear content based on user preferences" is that the adjustment involves selecting personalized options for the content stream. Personalization options can include things like choosing a favorite commentator or creating a stream tailored to one team in a two-team sports match.

[0143] Specific examples: John is watching the Super Bowl on one of his (multiple) ATS-enabled devices, which uses a camera to detect that he is wearing a red hat. The system (e.g., AAE 201) infers that he must be supporting the Chiefs (the red team) and selects a streaming version tailored for Chiefs fans. In some examples, John may be shown more replays of Chiefs scores and fewer replays of opposing teams.

[0144] In the same scenario as above, John wasn't wearing a red hat, so ATS initially didn't know which team he supported. As the game progressed, John began interacting by cheering "Go Chiefs!", booing the opposing team, and cheering for the Chiefs' points. This information was detected by ATS, which then selected the appropriate stream for John.

[0145] In the same scenario as above, John neither wore the red hat nor commented on the game. Based on John's past user preferences (he had previously supported the Chiefs), the system selected the Chiefs version of the stream for him.

[0146] - In the same scenario as above, the selected stream version may also include the selection of a commentator for John's support team.

[0147] John and Alice watched the Super Bowl together on (multiple) devices that supported ATS, which were equipped with lights and cameras. It (for example, AAE 201) inferred that John was a Chiefs fan because he was wearing a red hat, but inferred that Alice was a Lions fan because she was wearing blue and cheering for Lions scores. A balanced version of the content presentation was provided for both John and Alice. In some examples, the lights could turn red on one side of the room where John was sitting and blue on the other side where Alice was sitting.

[0148] Accents appearing in personalized content based on user preferences Another extension of "personalizing content based on user preferences" is changing the accent in the content to suit user preferences.

[0149] Specific examples: Bob is listening to an audiobook playing on his ATS device, which detects that he speaks with a British accent. The audiobook is played in an American accent by default; however, the UK version of the audiobook was selected for Bob to make it sound more familiar and natural to him.

[0150] Jane is watching a movie starring an Australian. She finds the accent difficult to understand, and the ATS detects that Jane has a French English accent. The Australian's voice is then generatively replaced with a French English accent.

[0151] - In the same scenario as above, Jane's preference for a French English accent was determined, so the voice assistant on her ATS-enabled smart device was also changed to a French English accent.

[0152] Personalized experience enhancement based on user preferences Another extension of "personalizing content based on user preferences" is that personalization is pre-loaded experience enhancements. Examples of experience enhancements will be described in detail under the next heading. In some examples, pre-loaded experience enhancements can be determined based on previous sessions where users with similar preferences enjoyed the same experience enhancements. Such implementations can improve the user experience and can also help cover situations where the ATS fails to detect attention responses.

[0153] Specific examples: Alice is watching a comedy on (multiple) of her ATS-enabled devices. There's a joke in the comedy that Alice doesn't laugh at. However, users with similar preferences to Alice usually laugh at the joke. The system plays virtual laughter for this joke because it receives positive responses from users with similar preferences to Alice.

[0154] Replay and experience adjustments based on user habits and characteristics Here, "user habits" refers to the regular ways users interact with content (positively or negatively), and these habits can include useful insights into how to optimize playback and the user experience. User habits can be determined through detection, such as which seat a user usually sits in when interacting with content using an ATS, or what time of day a user usually starts making annoying sighs due to noisy scenes in the content.

[0155] We use the term "user characteristics" to refer to user attributes that may influence how users experience content, such as a user's hearing, vision, and attention span. User characteristics can be estimated by the ATS based on how users interact with content, for example: - Always missing jokes when they're being told quietly suggests poor hearing; and / or - When an actor describes an object solely by its color, it becomes difficult to understand which object the actor is referring to, suggesting the actor may be colorblind.

[0156] User habits and characteristics can also be considered a form of user preference.

[0157] The phrase "playback adjustments" refers to changes in the presentation of content. Playback adjustments can include changes to spatial rendering of content, volume adjustments, and changes to the color balance on the screen.

[0158] Based on some examples, adjustments to the user experience can be designed to ensure that the user's environment is ideal for enjoying the content. Some such examples may include features such as automatically pausing content playback when interrupted, warning the user that certain things need to be evaluated before starting content, etc. These adjustments to playback and the experience can be managed on the user's (multiple) playback devices.

[0159] In some examples, adjustments can be made jointly to simultaneously improve the experience for multiple users of ATS-enabled systems. The following sections detail specific examples of user habits, user characteristics, and the resulting playback and experience adjustments.

[0160] Personalize content playback accessibility based on user characteristics. Another extension of "Playback and Experience Adjustments Based on User Habits and Characteristics" is to adjust playback based on user characteristics to improve content accessibility. User characteristics that may need optimization include myopia, color blindness, hearing loss, dizziness, epilepsy, etc. Playback adjustments that can be applied to these characteristics to improve content accessibility can include increasing the size of subtitle fonts, enhancing the color or appearance of objects in the scene, applying compression to streaming audio, reducing camera shake, and skipping scenes with flickering lights, etc.

[0161] Specific examples: Jake watches his favorite comedian's performances on his (multiple) ATS-enabled devices and laughs at almost every joke the comedian tells. The jokes that don't make Jake laugh are most often those told quietly. In some examples, dialogue enhancement could be enabled to improve the audibility of these jokes.

[0162] Personalized car notifications based on user habits Another extension of "playback and experience tuning based on user habits and characteristics" is improving the car experience based on seat occupancy. Seat occupancy can be determined by sensors indicating the user's position in the vehicle as they interact with content. In some examples, the ATS system can be configured to detect when seat occupancy is different from usual. The improved car experience can appear in the form of user warnings or notifications. These warnings or notifications can inform the user that they may have forgotten someone who would normally be on the trip. This improves the driving experience because the user can feel reassured knowing that a common participant in the trip hasn't been forgotten.

[0163] Specific examples: A user drives his parents-in-law to dinner every weekend. He starts backing out of his in-laws' driveway, but his mother-in-law isn't in the car. The car notifies the user that his mother-in-law is not there, thus averting a crisis.

[0164] Short-term experience enhancement Add elements to content based on reactive dynamics. Suppose one or more users are consuming content on a playback device with an ATS (Active Time System). The ATS can be configured to detect specific categories of responses (e.g., laughter, cheering, shouting "yes"), based on metadata in the content stream. These added elements can enhance the experience for one or more users in real time.

[0165] Based on some examples, the metadata might specify a fixed set of user sound events (e.g., cough, laughter, applause) or keywords (e.g., "Slam", "Ooof", "Higher", "Lower"), along with details of the appropriate response element. In some such examples, there might also be a set of emotions associated with the appropriate action (e.g., excitement, fear). Based on some examples, these elements may also be passed along with the content in the metadata stream to achieve this functionality.

[0166] In some examples, a locally stored element library may also exist (e.g., within a TV or other playback device), which can be applied to many content streams. In such examples, the scope of the element library can be broader than elements specific to a particular content presentation. Alternatively or additionally, generative AI techniques can be applied to automatically generate elements, such as text descriptions based on content presentation.

[0167] According to some examples, response intensity (e.g., the volume of a sound being played, the intensity of an audio effect, the size of an icon, the saturation of an icon's color, the opacity of an icon, the intensity of a visual effect) can be proportional to the intensity of a user's reaction (e.g., the volume of a shout, the length of a laugh, the pitch of a song).

[0168] The following sections highlight examples of the element types that can be added.

[0169] Adding auditory elements to content based on reaction dynamics Another extension of "dynamically adding elements to content based on response" is that the added elements are played audibly. According to some examples, a corresponding auditory element is added to the content whenever one of the users makes a specific response. Auditory elements can include playing sounds or adding audio effects. Examples of sounds that can be played include crowd noise, comedic sounds, impact sounds, etc. Audio effects can be such as reverb, distortion, spatial movement, etc. In some examples, auditory elements can be combinations of these elements or similar elements. Sounds and audio effects can include using Dolby Atmos™ or other types of object-based audio systems. An example of using Atmos™ sounds is playing a virtual laugh soundtrack by placing different laughs in different spatial locations. Another example of spatial audio effects can include hearing the baseball move around the user after the batter hits the ball.

[0170] Specific examples: Multiple users are watching a cat video stream. When a user's laughter is detected, a virtual laughter soundtrack is played through the speakers. The virtual laughter soundtrack can be created by placing different laughs in different spatial locations.

[0171] - During a baseball game, when a user's supported team hits the ball, a spatial "whoosh" sound and audio effect are added to the content. The spatial movement of the sound makes it sound like the ball is flying past the user. The ball's impact sound can also be encoded in the metadata provided along with the content stream.

[0172] Based on the dynamic addition of other users' auditory elements to the content, the system dynamically adds these elements to the content. Another extension of "dynamically adding auditory elements to content based on responses" is that the added auditory elements originate from other remote users on one or more other ATS systems. Such examples could allow users in a local environment to hear other people's responses to content while other users in the local environment are consuming it. For instance, when user A responds to content, that response can be recorded and sent to user A's friends (including user B). Later, when user B watches the same content, user A's response can be sent to user B's playback device via the content's metadata stream. User A's response can be played back for user B at the same time they are watching the content. In some examples, the responses played back for user B could include responses from multiple users simultaneously. The responses played back for user B don't necessarily originate from user B's friends.

[0173] Specific examples: Jane is watching a football match between Liverpool and Manchester. She cheers for all of Liverpool's goals and boos for some of Manchester's. Jane's friend Bob later goes to watch the match. Bob gets to experience the game with Jane cheering and boos.

[0174] TV programs are typically watched by millions of people. As users watch a program, their reactions can be recorded and sent to the cloud, where they are mixed to form a collective reaction to the TV program. When another user goes to watch the program, the collective reaction audio can be played along with the content.

[0175] Add visual elements to content dynamically based on reactions. Another extension of "dynamically adding elements to content based on responses" is that the added elements appear visually. Whenever one of the users produces a response of a selected type, a visual element can be overlaid or composited onto the displayed content. The visual element can be an icon, animation, emoji, etc. Alternatively or additionally, the visual element can be a visual effect added to the content (e.g., dithering, color changing).

[0176] Specific examples: Jane was watching the Oscars broadcast on her ATS-supported TV. When the Best Actor category was introduced, the host asked families to shout out the name of their favorite actor. Jane, a Tom Cruise fan and a fan of *Top Gun: Maverick* this year, shouted "Tom Cruise!" In response, an overlay image appeared on Jane's TV: a virtual Oscar statuette held aloft by a cartoon image of Tom Cruise, while her neighbor Ben, a Brad Pitt fan, saw an overlay image: a virtual Oscar statuette held aloft by Brad Pitt.

[0177] Injuries are common in ice hockey games. When watching an ice hockey game, an ATS-enabled TV can be configured to display an ambulance animation whenever the user says "Ambulance, come quick!"

[0178] In the 1960 Batman movie, whenever a character was hit during a fight scene, words like "Smash!", "Bam!", and "Pow!" would be overlaid on the screen in cartoon font. In 2026, Batman watched on an ATS-enabled TV could display similar words based on the user's dialogue during fight scenes. For example, if the Penguin is hit by Robin, the user would shout "Pow!". Based on this, an ATS-enabled TV could be configured to overlay the text "Pow!" instead of the other available options "Smash!" and "Bam!".

[0179] Lighting elements are added to the content dynamically based on the response. Another extension of "dynamically adding elements to content based on responses" is adding elements that are changes to lighting. In some examples, lighting changes can be emitted whenever a user interacts with the content in a predetermined way. These lighting changes can include a general change in lighting or can trigger animations that utilize the lights. A general change can include altering the color and intensity of all lights. In some such examples, each light can be changed to its independent color and intensity. In some examples, animations can be emitted that play on the lights in response to user reactions. Some example animations are waves, strobes, and flickering like flames.

[0180] Specific examples: - A user is watching a horror movie and verbally states that they are too scared while watching it. In response, the light brightness is increased, thereby reducing the fear induced by the content.

[0181] A family is watching a rugby game, with the two teams colored red and blue. ATS detects that the family is responding positively to the red team, so the lights in the room turn red.

[0182] Enhanced content based on responsiveness to improve comprehensibility. Suppose one or more users are watching content on a playback device associated with an ATS. In some examples, whenever one or more users interact with the content aloud (e.g., laugh, cheer), the content can be adjusted to improve the comprehensibility of parts that might be missed by one or more users. Comprehensibility can be improved by delaying the next scene until the user's reaction subsides, increasing the dialogue volume, or enabling subtitles so that the next scene is still understandable. In some examples, the content stream may include additional media to optionally extend the scene to achieve a time delay. Without said additional media, a time delay can be achieved by pausing the content until the intensity of the user's response decreases. Alternatively, generative AI can be used to generate additional media to extend the scene as needed, preventing content from progressing until the user is ready to continue watching. In some examples, the increase in dialogue volume may be proportional to the intensity of the user's response.

[0183] Specific examples: Bob is watching a comedy with Jane. Bob finds a particular joke in a scene hilarious and laughs for a long time. The next scene of the comedy can be delayed until Bob can control his laughter to a certain level.

[0184] A family is watching the Super Bowl, cheering loudly for a long time on a touchdown score for their favorite team. During periods of user interaction, the commentators' captions automatically turn on. This allows users to understand what the commentators are saying.

[0185] - A user is watching a stand-up comedy show and is laughing loudly at the jokes told by the comedian for an extended period of time. The system increases the comedian's volume as the user continues to laugh.

[0186] - A user is having difficulty understanding dialogue while watching a movie. Conversation enhancement features can be enabled to improve the audibility of speech within the room.

[0187] - In the same scenario as above, instead of enabling dialogue enhancement, enable hidden captions.

[0188] Personalization through feedback from content creators This section shares similarities with long-term personalization and short-term experience enhancement, but the difference lies in the fact that personalization is performed by content producers based on user attention metrics detected by one or more ATS systems. Attention feedback can come from explicit user responses to content provider requests or from implicit user engagement with the content.

[0189] Examples of explicit feedback in requests and responses include: The host said, "Make some noise!" Then the viewers either made a sound or didn't.

[0190] The host says, "What do you think, viewers watching at home?" Viewers may respond with emotions such as booing or cheering.

[0191] The host asks, "Should we choose the red or the blue object?" Viewers respond with "red," "blue," or a mismatch.

[0192] Examples of implicit user attention to content can include all types of attention mentioned in this disclosure that are not explicitly requested in response content.

[0193] Personalization decisions can be made by the content producer or by an automated system implemented by the content producer. In some examples, personalization may involve tailoring content to a general viewer. In some examples, personalization may involve generating multiple versions of the personalized content for multiple audience categories (such as users in a specific country or region, users known to have similar interests or preferences, etc.). In some examples, user attention responses may be sent to the service providing the content (e.g., the same server or data center) or different services.

[0194] Producers personalize content for the next iteration based on feedback. Content provided by content producers can be viewed by consumers at any time. Content producers can obtain information about how users engage with their content in an ATS-enabled environment. In some examples, this information can be used by content producers to personalize the next iteration of their content (e.g., episodes, albums, stories).

[0195] Specific examples: An influencer received attention metrics for their previous short video and found that viewers were highly engaged and wanted to see more similar products they had showcased. The influencer decided to feature similar products in their next short video for their audience.

[0196] A video blogger who creates a series of videos about video games asks viewers at the end of his videos, "Which character should I play next?" Viewers indicate (e.g., through sound, pointing) that they would like to see the blogger play that character next. The blogger then takes these attentional results into account when deciding what to play next.

[0197] The TV program ended with an open-ended cliffhanger. Based on ATS user responses, there were four main types of reactions to the ending. The TV program producers used this information to decide on four versions of the next episode. Based on how users reacted to the previous episode, they were shown episodes designed for their preferred audience.

[0198] Reaction-based interactive live media This section describes examples of what can be achieved when user attention information is sent back to content producers in real time. Content producers can then decide to adjust their content in real time based on these analyses.

[0199] In some examples, ATS-enabled systems can provide other users with the option to see how the audience responds. For instance, content producers can provide attention information within their content. Alternatively, real-time attention information can be sent back to the user in the content's metadata stream. Seeing other users respond in a particular way can increase the likelihood that a particular user will respond in the same way. Similarly, seeing that others do not respond might make users realize that they can have a greater impact on content presentation by responding in a particular way (e.g., cheering for a character).

[0200] Reaction-based interactive live games Other extensions of "reaction-based interactive live media" could be that the content is or includes game content, whether that game content involves individual user games, game competitions, etc.

[0201] Specific examples: Jane is livestreaming herself playing a video game. A choice appears in the game: she can choose to go left or right. She asks viewers to shout out which path she should take, "left" or "right." The viewers' attention response sends back to Jane, saying that 70% of them want her to go "left." Jane decides to go right to joke with her viewers.

[0202] Bob is hosting a live stream of a trivia video game. He asks his viewers to identify the correct answer in a multiple-choice format. The answer most frequently chosen by his viewers turns out to be wrong. Bob boosts the stream for a random segment of users by adding camera shake and says, "We got it wrong, so we have to take the camera shake together!" Reaction-based interactive live music content Another extension of "response-based interactive live media" is that the media is music-related content, such as bands, DJs, live music events, or one or more musicians talking and playing different instruments. Some attentional responses that can be used for live music streaming include clapping, singing along, dancing, and moving to the rhythm of the music. Content producers may explicitly ask users to make attentional responses, such as asking users to sing the most memorable lyrics of a song.

[0203] Specific examples: A DJ is performing a show that includes participants joining remotely via their ATS-enabled devices. The DJ gradually shifts to more energetic songs. Attention metrics show the DJ that the audience is dancing more enthusiastically at this point in his show than usual. Based on this information, the DJ decides to switch to more energetic songs earlier than originally planned.

[0204] A musician was livestreaming from his studio. While playing a guitar solo, the musician noticed reports indicating a low level of focus on the performance. Additionally, some viewers responded with "bongo." The musician then decided to start playing the bongo drums.

[0205] Reaction-based interactive live user-generated content Another extension of "reaction-based interactive live media" is that the media is user-generated content.

[0206] Specific examples: A video blogger is livestreaming from abroad. The blogger enters a store to buy snacks and showcases the selection on the shelves. The blogger asks viewers what they think he should buy. The blogger uses attentional information to decide what to purchase, helping to personalize the experience for viewers.

[0207] Reaction-based interactive live streaming Another extension of “reaction-based interactive live media” is that the media is programs from broadcasters (e.g., reality TV programs, drama TV programs, award shows, live movie broadcasts).

[0208] Specific examples: During the People's Choice Awards ceremony, users were allowed to vote for Best Picture using their ATS-enabled devices. The host asked, "Which movie do you think should win Best Picture?" Users then shouted out the names of the movies they wanted to win. The host then announced the winner based on votes cast using attention metrics.

[0209] - A new interactive Jummy's Goofy Show™ movie is being filmed and streamed live. They allow users of ATS-enabled devices to vote on who should wear a cape and jump down a ramp on a quad bike.

[0210] - In the Eurovision Song Contest™ finals, the winner is determined based on real-time attention indicators from ATS-enabled devices.

[0211] - During the broadcast of Saturday Night Live™, the producers were interested in understanding how people were engaging with the content in real time. They discovered that people found a joke the host had just made particularly funny to viewers. Based on this, the producers encouraged the live audience to laugh louder (for example, through signs visible to the live audience but not to TV viewers), and the host was instructed in their earpiece to continue improvising on the joke with the guest for a longer period.

[0212] Reaction-based interactive live media viewing sessions Another extension of "reaction-based interactive live media" is that content presentation is a viewing session of a program, which may be a replay. Multiple participants join the viewing session via multiple devices (or more) that support the ATS, sometimes simultaneously. This viewing session can be hosted by a live producer or by an automated system. Some examples of suitable viewing session programs include "Choose Your Own Adventure" shows, quiz shows, trivia games, and children's TV programs (such as those involving call-and-response elements).

[0213] Specific examples: - In one segment of the show, Mr. Hug Man will hug a celebrity. During a viewing session, several participants voted to predict how long Mr. Hug Man would hug the celebrity, until the celebrity taps him to signal the end of the hug.

[0214] - A game hosted by an automated system where viewers must find a list of hidden items within a scene.

[0215] recommend Some publicly available examples involve recommending content to users based on user preferences and / or current user state, as determined through analytics such as using an ATS (Advanced Feature Controller). Recommendations can be computed in the cloud or on the user's (multiple) ATS-enabled devices. Recommendations can be based on short-term or long-term information, or a combination of both. Recommendation systems can also be optimized in the cloud to better understand what a broader population (rather than specific environments, such as people in a particular household) likes in specific situations (e.g., user mood, viewing session length, news events).

[0216] Content can be recommended before, during, or after playback. In one example of recommending content during playback, the user did not respond positively to the currently playing content, so a different piece of content that matches their preferences was recommended.

[0217] Emotion-based content recommendation Users can receive recommended content based on emotions detected by their (multiple) ATS-enabled devices. ATS can detect a user's laughter, sighs, valence and / or arousal level in speech, and lack of response to content that would normally excite them.

[0218] In some examples, ATS systems can also leverage closed loops that can be created as described above. For instance, an ATS system can recommend content based on whether the user's last consumption of a particular type of content had a positive impact when the user was in their current mood. Some examples include exploring whether uplifting or sad music is more helpful in helping the user shake off a negative mood, or whether watching sports is more enjoyable than watching comedy when the user is in a pleasant mood. Recommendations can also take the form of short-term sentiment estimates that do not consider user preferences to avoid forming closed loops for recommendation improvement.

[0219] Long-term analytics can also be used to determine sentiment-based content recommendations. Suppose the ATS (Advanced Transactional System) is unaware of the user's mood (e.g., the user has just woken up). In some examples, the ATS can refer to long-term sentiment information to infer possible moods. This inference can take the form of discovering patterns (e.g., the user is usually in a good mood in the morning), detecting trends (e.g., the user's mood has been improving as summer approaches), or using filtered sentiment information (e.g., the user has been in a good mood all week, so they are probably in a good mood now).

[0220] Specific examples: Bob was in a low mood, which ATS detected because he wasn't laughing at jokes in light comedies that matched his usual preferences. A dark comedy was recommended to Bob as the next option. ATS detected an improvement in Bob's mood through a few soft laughs. This information was stored (possibly on his device or in the cloud) to help guide the recommendation system when Bob found himself in a similar mood again.

[0221] A child was watching a TV program and started to feel scared. The child was then advised to switch to a calmer program.

[0222] Content recommendations based on user interest in specific views The term "view" can refer to any visual element that represents a particular type of content. Some examples of views include thumbnails when browsing content options, different screens playing content simultaneously (e.g., a tablet playing a game while a TV program is playing), multiple windows on a screen, and so on.

[0223] Attention to a specific view can be determined by one or more ATS-enabled devices, through gaze tracking, via audio direction-of-arrival information, or otherwise. In some examples, using this information, along with knowledge of the playback device's location, ATS can estimate what content a user is most likely to be interested in and can recommend more similar content. For example, knowing how long a user gazes at each thumbnail is a form of measurable attention.

[0224] Specific examples: - As a user scrolls through the list of content options, (multiple) ATS-enabled devices can report which content thumbnails the user is most interested in. As the user continues scrolling to see more options, recommendations will point to content that the user has implicitly expressed interest in and paid attention to.

[0225] Music-attention-based content recommendations Music attention can be determined by (multiple) ATS-enabled devices by detecting sing-alongs, tapping in the air or on an object (steering wheel), or calling out the name of a song or artist when it begins playing. These music attention indicators can then be used to determine a user's music preferences.

[0226] Content recommendations can take the form of recommendations based on the user's music attention indicators or automatically arranged playlists. Recommended content can be similar in terms of genre, era, artist, and song structure. Other media formats can also be recommended based on music preferences. For example, a movie can be recommended because its soundtrack matches the user's musical taste.

[0227] Specific examples: Alice is on a road trip and has a habit of clapping along to her favorite songs on the steering wheel. Using this information, her musical preferences were determined, and a curated road trip playlist was recommended for the remainder of her journey.

[0228] Exercise examples Exercise use cases can involve (in some instances simultaneously) long-term personalization, short-term experience enhancement, personalization through feedback from content creators, or a combination thereof. Exercise content can include yoga classes, fitness classes, music that users enjoy listening to while exercising at home, etc. Specific sounds that can be listened to to detect exercise-related attention cues on (multiple) ATS-enabled devices may include: -Gasping sounds; - A deep breath; - Verbally count the number of repetitions; -Increased respiratory rate; - Body shaking; -Visual detection training movements; - The technical quality of visual inspection users; -Emotion recognition, etc.

[0229] In some examples, the ATS can provide attention-related information to a human or virtual coach in real time. This attention-related information can correspond to the attention of one or more users during the exercise. Alternatively or additionally, the ATS can provide the coach with user tracking records, performance trends, attention information from previous sessions, and so on.

[0230] Specific examples: Jane uses her ATS-enabled devices (multiple devices) to exercise every Monday. Her progress is tracked over time, and the difficulty of the workouts is adjusted to match her progress. Jane also gets a summary of her progress on her ATS-enabled devices (such as a dashboard on her phone display).

[0231] Bob wants to take a workout class that matches his level and preferences. Bob watches a pre-recorded workout class that is personalized for him, tailored to his needs by selecting 'tracks' of videos for him. 'Tracks' can refer to different sections of pre-recorded video, thus allowing users to choose workouts in the order they most desire.

[0232] - Multiple users (e.g., two) decide to participate in a workout class together. They have different preferences for exercise movements and different ability levels in areas such as flexibility, strength, and aerobic exercise. They provide input to ATS, indicating that they want to complete core training, and ATS automatically generates a workout class suitable for their abilities and preferences.

[0233] The workout session was led by a virtual coach. ATS detected that the user was having difficulty completing the final repetitive movements due to asthma. ATS then instructed the virtual coach to discourage the user from completing these repetitive movements.

[0234] - In a workout session led by a virtual coach, one participant did not make an effort to complete a set of exercises. Multiple ATS-enabled devices detected this, and the virtual coach encouraged the participant to try harder and persevere.

[0235] A coach is leading a live workout session, with many participants using multiple ATS-enabled devices. The coach is informed that John is not keeping up with the current exercise. However, the coach also gains information about John's attentional performance in previous sessions. The coach offers encouragement, for example, by shouting, "Keep it up, John! You did it last week, I know you can do it!" - In another live workout session led by a coach, ATS detected that one of the participants was performing the current exercise with particularly poor posture. Based on the ATS feedback, the coach decided to speak directly with the participant through the user's content stream to offer suggestions on how to improve the exercise.

[0236] Replay optimization and control use cases This section provides examples of long-term personalization and short-term experience enhancements. "Playback optimization" in this section refers to optimizing playback across coordinated devices, which in some examples is based on an objective function. This objective function may involve, for example, maximizing attention, maximizing intelligibility, maximizing spatial quality for the listening location, etc. Updates required to implement playback optimization can be computed on the user's ATS-enabled device, or computed in the cloud and sent to the user's device (possibly via metadata streams provided with the content). "Playback controls" refers to features such as play, pause, rewind, and next episode.

[0237] Specific examples: Six users are in a car equipped with ATS (Automatic Sound System) and playing music. Two users in the back seat are talking. They are not paying attention to the music but are focused on their discussion. The music playback in the back seat is muted, but the volume remains high for the other users enjoying the song.

[0238] Alice likes to watch TV in the living room, which is adjacent to the kitchen, while Bob prepares dinner. Bob likes to listen to music while preparing dinner. Playback has been jointly optimized for both Alice and Bob so that they can both hear spatial audio centered on their respective room areas. Furthermore, the audio from Alice's content is much quieter in Bob's area, and vice versa.

[0239] - Many users are watching a movie in a room, but then all of them leave to greet a visitor. ATS detects that the room is empty (e.g., based on camera data) and that attention levels to the content have significantly decreased. ATS pauses the movie.

[0240] Bob received a phone call while watching a movie. Multiple ATS-enabled devices detected that Bob wasn't paying attention to the content because his posture showed he was holding the phone to his ear. ATS paused the movie for Bob while he answered the call.

[0241] In a similar scenario, Bob had previously shown a preference for resuming content during calls by automatically pausing and then restarting. The system noticed Bob resumed the call during the movie and, based on his learned preference, did not pause the content for him.

[0242] Advertising use cases using attention feedback Previously deployed advertising systems have limited ways of determining user attention. Current methods include having users click on ads, allowing them to choose whether to skip ads, and having test audiences complete questionnaires about the ads. With such methods, advertisers may not know whether the user was present when the ad was displayed.

[0243] Smarter advertising can be achieved using ATS (Adaptive Testing System). Smarter advertising allows for improvements in current technologies, such as ad performance, ad optimization, audience sentiment analysis, user interest tracking, and intelligent ad placement. Furthermore, smarter advertising creates new advertising opportunities, such as: -Interactive advertising; -Personalized advertising; - Attention-driven shopping; and - Stimulate attention to enhance advertising effectiveness; This section will explore these concepts in detail through a list of use cases.

[0244] Example use cases User attention-based advertising analysis When a user consumes content on their ATS-enabled device, the ads playing on the system can also utilize ATS. Rich advertising analytics can be obtained from ATS. Some examples of analytics that ATS can provide include: - How is the overall level of user interaction? -What aspects of the advertisement did the user interact with? - Is the user's interaction positive or negative? - Who interacted with the ad (e.g., based on demographic characteristics, user preferences, and interests); - When and in what specific way users interact with ads (e.g., morning / evening, Monday / Tuesday, etc.); - Whether a particular actor in the advertisement attracts more or less attention; and - Which parts of the content elicit higher or lower attention (the ad is placed at the beginning, middle, or end of the content; or after a suspenseful, humorous, or romantic scene, etc.).

[0245] Specific examples: The marketing team behind a new 3D-printed basketball wanted to understand how their latest ads were performing to determine if their spending was worthwhile. The team ran the ads on ATS-enabled devices and has now received analytics on ad performance.

[0246] Following the 2023 Super Bowl, and to satisfy public curiosity, a ranking of the best commercials from the event was desired. The Super Bowl was broadcast on devices supporting ATS, allowing for the analysis and comparison of each commercial to determine the rankings.

[0247] Optimize ads based on user attention "Attention-based ad analytics" can be used to optimize ads to achieve the desired attention response. This optimization can be achieved by iteratively improving the ad based on analytics after releasing a new version. Alternatively or additionally, this optimization can be achieved by simultaneously releasing multiple versions of the ad and observing how user responses differ.

[0248] Specific examples: A company wanted to optimize its ads to maximize positive user attention before fully launching its marketing campaign. The company decided to soft-release the ads to 1,000 ATS users. Using the analytics from this soft-release, the company created another version of the ads and soft-released it again, iterating until they were satisfied with the ads. Finally, they launched the campaign with full confidence in their ads.

[0249] - In the same scenario as above, instead of a soft release of content, it involves iterative focus group testing. The focus group views the content using ATS-enabled devices.

[0250] A company had many excellent ideas for promoting its new product. They decided to create multiple versions of their advertisement and release them simultaneously. Using analytics from devices that support ATS (Advanced Data Settlement), they discovered that one version of the advertisement significantly outperformed the others. They decided to continue advertising only with the best-performing ad.

[0251] Product sentiment analysis based on user attention Users' ATS-enabled devices can be used during content and ad replays to determine their sentiment toward the product. This sentiment can also be assessed on a user or group basis. For example, a group could be users of a specific age group, users with shared interests (e.g., cars), users with certain attentional characteristics (e.g., users whose laughter frequency falls within the typical range for adult women), group segmentation obtained through learning (e.g., users who sing with a specific timbre within a certain frequency range), or the total number of all ATS users, etc.

[0252] Demographic characteristics and user interests are subsets of factors that can be used to select a group for evaluating sentiment analysis. Demographic characteristics and user interests can be specified by the user, estimated by the ATS, or both. For example, user interests can be estimated in the manner specified in the 'Determining User Preferences' section. Demographic characteristics can be estimated based on attentional indicators, such as the characteristics of their responses (e.g., their laughter frequency range, jump height, their visual appearance indicating age), the types of attentional responses they use (e.g., the types of words they use and dance movements), etc.

[0253] Emotions can be determined by a user's implicit or explicit reactions to a product. Implicit reactions can include yawning, looking at the product when it appears on the screen, etc. Explicit reactions can include positive reactions, such as "I love <product name>", or negative reactions, such as "It's <product name> again", and then leaving the room, etc.

[0254] Specific examples: The latest James Bond film featured James wearing an Acme watch. Acme wanted to understand the watch's brand awareness during this product placement and obtained sentiment analysis on the watch from ATS at a group size. Acme decided to sign a similar product placement agreement for the next James Bond film.

[0255] - Continuing with the previous example, Acme had already launched an advertising campaign on TV before the movie's release. After the James Bond movie's release, they used ATS-enabled devices to capture emotions associated with the watches featured in the TV commercials. As the film's viewership increased, Acme found that the watches in its TV commercials garnered greater attention and more positive emotions.

[0256] A company is trying to expand the age reach of its product. They want to understand how people aged 40 to 50 perceive the product. They determine the product's sentiment among this demographic by combining two types of attention analysis: one from users who identified as belonging to this age group, and the other from users who did not provide age information but were estimated to belong to this age group by their Attention Analyzer (ATS). Based on this data, the company gains a wealth of sentiment analysis information from the ads it runs on devices that support ATS.

[0257] A company is trying to target its product to introverts. The company hypothesizes that introverts respond more calmly than other groups. They decide to filter their product's sentiment analysis to those who don't interact loudly and whose reactions are relatively subtle.

[0258] Interactive advertising using ATS Ads can be interacted with using ATS-enabled devices. Interactive components of an ad can be sent via the ad's metadata stream, including information such as the type of attention to respond to and the corresponding action the ad should take. Types of actions an ad can take include changing the ad's content, implementing control mechanisms (such as skipping the ad), and storing user responses (e.g., on the user's device or in the cloud) for use the next time the same product ad is played. These actions can help determine whether a user has interacted and allow for gamification of the ad. For example, a user could be rewarded for skipping an ad because they have demonstrated some level of interaction with the device.

[0259] Specific examples: ABCD launched a marketing campaign for their new car. They decided to place ads on ATS-enabled devices to provide an interactive experience. Their ads included a racing game. The ads allowed users to lean left or right, or say "left" or "right" to control the vehicle. Compared to traditional advertising methods, users responded more positively and showed greater awareness of the new car.

[0260] An insurance company wanted to demonstrate that minor accidents, which are commonplace in life, aren't so bad if you have insurance. To achieve this, the company ran a "Choose Your Own Adventure" ad on multiple ATS-enabled devices. Each time the ad appeared, users could choose a new path for the story, and the result was always "You should get insured."

[0261] A company is releasing its new game with an accompanying ad designed to engage people and generate excitement for the release. This ad is displayed on (multiple) ATS-enabled devices and occupies multiple ad slots to complete this series of ads mimicking game features. Each time the ad from this campaign is played, users gain access to mechanisms for personalizing the ad, such as customizing the character's appearance. The next time an ad from this campaign is played, the character will have the customized appearance, and the story will continue based on any other interactions the user has previously made.

[0262] - A user is watching content with embedded ads. The user is rewarded for not skipping the ads and actively engaging with them (e.g., making strong eye contact or discussing the ads). Example rewards could include not receiving any more ad interruptions during the subsequent 30-minute replay of the content.

[0263] Personalized advertising based on user preferences Ads can be personalized in a manner similar to the personalized content adjustments detailed in the "Long-Term Personalization" section. As described in the "Long-Term Personalization" section, these adjustments can be transmitted via metadata streaming, sent directly as a selected version, or created through a generation process. In some examples, ads can be selected that simply match user interests.

[0264] Specific examples: Two users watching TV together were detected to enjoy comedy through an ATS-enabled device, so they were shown humorous versions of the ads. Conversely, another household preferred straightforward content, so they were shown more serious versions.

[0265] - A user watching a basketball game on (multiple) ATS-enabled devices was identified as a fan of Stephen Curry due to his positive response to the game. Advertisements for tickets to the next game were personalized by selecting the version starring Stephen Curry.

[0266] In the same scenario as above, the user never talked about Stephen Curry. However, ATS knew that Stephen Curry was the user's favorite player because the user was always watching him on the court.

[0267] Jane and Bob are watching a show about traveling around the world. They're discussing how much they want to travel to Greece, a topic detected by (multiple) ATS-enabled devices playing the show. During the next commercial break, an advertisement for a resort in Greece is shown to Jane and Bob.

[0268] - When Bob watches the latest James Bond movie, he says, "I really like his suit," thus establishing his interest in James Bond-related suits. During the next ad break, Bob is shown an ad for a similar-looking suit. As mentioned above, this example can be enhanced by listening for pre-filled words relevant to the content ("suit" or "ABCD").

[0269] In a similar scenario, the user didn't mention liking the suit. Instead, a camera connected to the ATS allowed the system to determine that the user was particularly interested in the James Bond watch. Therefore, during the next ad break, an advertisement for the watch was shown to that user.

[0270] Attention-driven shopping By using (multiple) ATS-enabled devices, numerous shopping-related opportunities can be realized. This can encompass a variety of content types. Some examples include shopping-related channel content and virtual assistants. New shopping-related opportunities may include: - Similar to the product sentiment analysis detailed in "User Attention-Based Advertising Analytics", but specifically for shopping-related content (e.g., online stores, shopping TV channels). - Interact directly with shopping materials via ATS-enabled devices; -Use user preferences and interests to optimize their shopping experience and the products, deals, etc. displayed; There may also be more attention-driven shopping opportunities.

[0271] Specific examples: - When a user browses an online store, eye tracking or verbal comments like "Oh, I like that" indicate interest. The item will then be automatically added to their shopping cart.

[0272] - The directly interactive shopping experience supported by ATS explicitly tells the audience, "If you want this product today, clap!" Part of the audience who participate in the product interaction and support ATS receive the product for free. The remaining users who do not receive the product add it to their cart and receive a price update for the product.

[0273] Shopping TV channels use attention analysis to determine which products users are most interested in. Users' devices can also sense their owners' emotions towards the products.

[0274] Shopping TV channels use attention analytics to determine which aspects of advertised products (e.g., features, price) help users make a purchase or cause them to lose interest. For example, if a user is engaged during a product presentation but loses attention during a price display, then the pricing strategy is flawed.

[0275] - Users can interact with the virtual assistant using ATS. The virtual assistant can access and update the user's shopping interactions. For example, when a user points to a product they are interested in on the screen, they can say, "Listen, Dolby, add that product to my cart," and the virtual assistant can automatically add the product to the user's cart. The products a user is interested in can come from non-advertising related content, such as movies, podcasts, etc.

[0276] When a user consumes shopping-related content, the model of each product on the screen is a version selected based on their preference. For example, John insists that all his personal items are yellow. If a yellow option is available, he will see a yellow version of each product.

[0277] -ATS detected that Alice responds better to shopping content when items are sorted by price in ascending order. The next time Alice enters the online store, this shopping preference will be automatically selected for her.

[0278] - The user expressed interest in a car appearing on the shopping channel. The playback system responded, "Would you like me to schedule a test drive for you?" The user interacted by saying "yes," and the test drive was automatically scheduled via an ATS-enabled device.

[0279] A user expressed interest in a vacuum cleaner on a shopping TV channel. The user said, "Wow, this looks amazing, I want it," and (multiple) ATS-enabled devices added in the playback, "Would you like to order now?" The user's response to this question was, "No, I'll think about it," and the order was rejected due to the user's reaction.

[0280] Smart ad placement based on user attention Attention information, such as that detected by ATS, can be used to guide ad placement. We use "ad placement" to refer to when and what type of ad is placed. Ad placement decisions can be made using long-term trends and optimization, or in real-time using information about the user's current interaction. Furthermore, a combination of both can be used, where real-time ad placement decisions can be optimized over the long term. Examples of decisions that can be made using this information include: - Place ads in locations where users interact with the content the least to minimize ad annoyance.

[0281] - Place ads where users interact with the content most to maximize ad attention.

[0282] - When a topic appears in the content, place an ad related to that topic.

[0283] - When a user interested in a particular product type interacts with the content, an advertisement for that product is displayed.

[0284] - Utilize the closed loop supported by ATS to optimize ad placement based on ad or content performance.

[0285] -A combination of any of the above.

[0286] The decision-making process for ad placement based on learning can be conducted at different levels, such as: - By user (e.g., preferring ads to be at the beginning of content); -By group type (e.g., Gen Z viewers, or French viewers, jazz fans); - By scenario (e.g., some users are more likely to interact with an ad when a particular user appears on the screen: after a scenario in which Ryan Gosling appears on the screen, a Tag Heuer ad featuring Ryan Gosling is played for users who responded to that scenario). - By episode (e.g., learn the part of the episode where receiving ads is least likely to be offensive). - By series (e.g., attention is usually highest in the last five minutes before the end of a series); etc.

[0287] Specific examples: An original equipment manufacturer (OEM) that produces components for devices that support ATS sells attention information to broadcasters. The broadcasters then sell ad slots based on the attention levels they receive.

[0288] - A mobile video game generates revenue by advertising other games during replays. The game studio that developed this video game used ATS-enabled phones to determine that user engagement was lowest after a battle in the game ended. The studio wanted to minimize ad intrusion, so they decided to place ads after battles.

[0289] A TV program production company values ​​the viewing experience of its programs. Therefore, they want to optimize ad placement to maximize content performance. The production company uses attention levels after ad insertions to determine the impact of ad placement on the program. Some types of attention they can focus on include: user excitement about the program's return, all users leaving the room and being absent, and users now being more focused on their phones.

[0290] - During TV programming, ad placement will be delayed until the user reaches at least a certain level of attention. Alternatively, ads may be reduced when the user is focused on the content.

[0291] Predictive ad placement based on user attention Drawing on the aforementioned section on "Smart Ad Placement Based on User Attention," models can be developed in some examples to predict attention types, attention levels, the optimal timing for placing ads within content, or combinations thereof. When training the model, no additional attention information is required to determine ad placement. For example, the model can learn by referencing one or more content presentations (e.g., audio, video, text). If the model is tasked with predicting attention levels, content providers or users can use this information to decide where to place ads. Models trained to predict optimal ad placement can be trained to predict different ad placements based on user attributes such as age, user interests, user preferences, location, etc.

[0292] Specific examples: - Created a creative tool that attempts to predict the optimal timing for ad placement. A TV broadcaster used the tool to automatically determine where to place ads in its 24 / 7 programming.

[0293] Attention stimulation Incentives can be offered to users of (multiple) ATS-enabled devices to encourage them to interact with content or advertisements. These incentives can be provided by content providers, content producers, TV manufacturers, etc. Examples of possible types of attention-incentivizing rewards include expressing feelings about the content (e.g., through sound, movement). Example reward types could include things with monetary value, additional content, etc.

[0294] Specific examples: Jane listened to a lot of her favorite artists' music, and ATS determined that Jane often sang along to the artists' music. Based on ATS data, the artist's record company provided Jane with free tickets to the artist's performances because she was one of their most loyal fans.

[0295] - One user watched every Manchester United football match and cheered loudly for every goal. They can enjoy a discount on tickets to watch Manchester United live matches.

[0296] A group of children watched an animated film. Based on the group's positive emotional response and their apparent enjoyment of the film, additional content was unlocked. At the end of the film, they were given fictional bonus features and extra "behind-the-scenes" segments.

[0297] The user received an advertisement for a sound system and responded positively. The user received a offer to receive a free TV with the purchase of the sound system.

[0298] Use cases for content performance evaluation using attention feedback Current content performance evaluation methods typically involve having a test audience preview the content. Obtaining metrics through test audiences has several drawbacks, such as requiring manual labor (e.g., reviewing questionnaires), failing to represent the final audience, and potentially being costly. In this section, we will detail how to overcome these problems by using an attention tracking system (ATS).

[0299] Having an Attention Strategies (ATS) allows for precise determination of how users respond during content playback. ATS can be used on end-user devices, enabling all content consumers to become test viewers, reducing content evaluation costs and eliminating the problem of unrepresentative test viewers. Furthermore, the analytics generated by ATS require no human labor. Because the analytics are collected automatically in real time, content can be automatically improved by machines. However, the option of manually optimizing content remains available. Moreover, using ATS in the content improvement process can create a closed loop, where decisions made using attention information can be tested for effectiveness by reusing ATS. This section details examples of how to leverage ATS for content performance evaluation and content improvement.

[0300] In this section of this disclosure, we refer to a type of metadata that specifies where a user is expected to make a particular attentional response. For example, laughter can be expected to occur at a specific timestamp or within a certain time interval. In some examples, a certain emotional atmosphere can be expected for the entire scene.

[0301] In some implementations, the performance analysis system may receive the expected level of response to the content (specified by the content creator and / or based on response statistics detected by the ATS) and then output a score that can serve as an indicator of content performance.

[0302] Event analyzers can receive attention information (such as events, signals, embeddings, etc.) to identify key events within content that elicit responses from (multiple) users. For example, an event analyzer can perform clustering on response embeddings to identify regions or events within content where users react in similar ways. In some examples, probe embeddings can be used to locate times when similar attentional indications occur.

[0303] Example use cases Adding value to content creators The section on 'Using Attention Feedback for Content Performance Evaluation' highlights the value that Attention Feedback (ATS) adds for content creators. Several aspects can benefit content creators and providers by implementing ATS. These aspects include: -Content performance evaluation; -Content improvements; and - Real-time content improvements Content performance evaluation Content performance is evaluated based on user attention. Attention metrics from users' Attention Scale (ATS) can be used to determine content performance. For example, a user leaning forward while watching a screen providing content suggests interest. Conversely, a user discussing topics unrelated to the content may indicate disinterest. This information about user attention can be aggregated to gain insights into overall user response. These aggregated insights can then be compared to results from other content or portions thereof to compare performance. Examples of content or portions of content include TV shows, programs, games, levels, etc. Differences in attention levels can reveal useful insights into content performance. Furthermore, attention information can indicate what users are focusing on (e.g., topics, objects, effects, etc.). Note that any content performance assessment obtained using an ATS can be combined with traditional assessment methods such as surveys.

[0304] Use the written metadata to evaluate content performance. Another extension of “evaluating content performance based on user attention” is listing potential user responses in the metadata. Suppose one or more users are watching content (e.g., an episode of a Netflix series) on a playback device with an associated ATS. The ATS can be configured via metadata in the content stream to detect specific categories of responses (e.g., laughter, shouting “yes,” “oh my god”).

[0305] In some examples, content creators or editors can specify what the expected audience response should be. Content creators or editors can also specify when the expected response should occur, such as at a specific timestamp (e.g., at the end of a witty remark, during a funny visual event, such as a cat smoking), during a specific time interval, for a type of event (e.g., a specific type of joke), or for the entire piece of content.

[0306] Based on some examples, the anticipated response can be delivered along with the content in the metadata stream. There can also be a user response type library that can be applied to many content streams (e.g., stored on the user's device, another local device, or in the cloud), which can be more broadly applied. The metadata regarding which attention indicators are expected to be acquired can be those attention indicators that are only listened to and have the user's permission, aiming to give users greater privacy protection while providing content producers and providers with the attention analytics data they need.

[0307] User reactions to the content can then be collected (in some examples, these reactions correspond to metadata). Statistics based on these reactions and metadata can be used to evaluate the content's performance. Example evaluations for specific types of content include: joke Content creators can add metadata specifying where they expect the audience to respond with laughter. Laughter can even be categorized into different types, such as 'heart-pounding laughter,' 'chuckle,' 'gasping laughter,' and 'machine-gun laughter.' Additionally, content creators can choose to detect other verbal responses, such as someone repeating a joke or attempting to predict a witty remark.

[0308] During content streaming, in some examples, metadata can instruct the ATS (Action, Substance, and System) to detect whether a specific type of laughter response occurs. Statistics on the responses can then be collected from different audiences. Performance analytics systems can then use these statistics to evaluate the content's performance, which can serve as useful feedback for content creators. For example, if statistics show that a particular segment of a joke or skit didn't garner much laughter from the audience, it means that the segment's performance needs improvement.

[0309] fright In horror movies, audience responses can be expected to include things like "Oh my god," a noticeable jump, or a gasp. Analysis of ATS information collected based on this compiled metadata can show that specific segments of a horror scene elicited almost no startle response. This suggests that this part of the horror scene needs improvement.

[0310] Controversial topics Some channels stream debates about different groups, events, and policies, which can garner numerous comments and discussions. During content streaming, metadata instructs the Accessibility Strategies (ATS) to detect whether supportive or debating responses have emerged. Statistics on user responses can then be collected from different audiences. This ATS data can help content creators analyze the receptiveness of topics.

[0311] Incendiary scenes Content creators can add metadata specifying where they expect a strong negative reaction from viewers (such as "Oh, disgusting!" or turning away). This can be used for horror movies, user-generated "disgusting" content, etc. During content streaming, such as a video of someone eating a spider, in some examples, metadata can instruct the ATS (Active Data System) to detect whether an inflammatory response has occurred. Aggregated data might show that users did not exhibit a disgusted response during the scene. Content creators can then decide that additional work is needed to make the scene more inflammatory.

[0312] Exciting scenes Content creators can add metadata to specify where they expect a strong positive response from viewers (such as "Wow," "So beautiful," etc.) within their content presentation. This technology can be used in film, sports broadcasting, user-generated content, and more. For example, in snowboard broadcasting, slow-motion sequences of exciting moments are expected to receive a strong positive reaction. Receiving user reaction information from ATS (Automatic Data Setter) and performing data aggregation-based analysis can provide content creators with insights. Content creators can determine whether their audience likes the content and then adjust the content accordingly.

[0313] Evaluate content performance using learned metadata. Another extension of "evaluating content performance using written metadata" is learning metadata from user attention information. During content playback, the Event Analyzer (ATS) can collect responses from the audience. The statistics of these responses can then be fed into an event analyzer to help create meaningful metadata for the content. Metadata can be generated, for example, based on one or more specific dimensions (e.g., fun, tension). In some instances, the event analyzer can use techniques such as peak detection to determine what metadata to add to pinpoint where events might occur. While the content may already have written metadata, it's still possible to generate additional learned metadata using such publicly available methods.

[0314] Specific examples: - A live stand-up comedy show tagged all the jokes in the metadata using information from ATS-enabled devices.

[0315] A program that already had metadata learned additional metadata using an ATS-enabled device. The additional metadata showed that viewers laughed at a moment that was unintentionally funny. The content producer decided to highlight that moment.

[0316] Create highlights based on audience feedback Another extension of “using learned metadata to evaluate content performance” is creating highlight reels based on the most interactive segments. For example, during a match broadcast, viewers might react excitedly, perhaps shouting “Go! Go!” or saying “Oh no!” when their team loses a fight. These reactions can be detected by the ATS and can be predefined reaction types. Overall, the statistics from the ATS data can show the relatively most prominent moments in the broadcast. These prominent moments can be used to automatically create highlight reels, or require some effort from editors. Furthermore, highlight reels can be created for each user simply based on their attention indicators from watching the match on an ATS-enabled device. Similarly, highlight reels can be created for subsets of ATS users (such as a group of friends, “red” team supporters, etc.).

[0317] Users can enhance their viewing experience by hearing other users' reactions to highlights in the compilation. This can be achieved in the same or similar way as detailed in the "Adding Auditory Elements from Other Users to Content Based on Reaction Dynamics" section of "Personalization and Enhancement Using Attention Feedback".

[0318] Popular events based on user attention Another extension of "creating highlights based on audience response" is that the highlights originate from a specific set of events (e.g., jokes, steals, catches). These events may appear in different content, such as TV shows, games, books, songs, etc.

[0319] Specific examples: - Based on responses detected by ATS-enabled devices, the top ten best football goal highlights are determined. Reactions that ATS looks for can include supportive responses such as clapping, cheering, shouting "Great goal!", etc.

[0320] - A baseball game automatically generates highlights based on the sequence of most exciting moments for the user.

[0321] - A stand-up comedian's ten funniest jokes are automatically filtered based on user attention and used in his promotional videos. Reactions that ATS can look for include laughter, applause, cheers, etc.

[0322] Real-time content improvement Integrating voting into podcasts In some examples, podcasts can collect live audience votes by gathering predefined words (such as "yes" or "no"). This process makes viewers feel more interactive, as if they are part of the live stream rather than passively receiving content.

[0323] An example: A musician is broadcasting live music to their viewers. Viewers are asked to choose which song to play from four options. By saying 'A,' 'B,' 'C,' or 'D,' or pointing to an option, listeners can select their choice, and their attention is detected by their Attention Scale (ATS). These responses are then aggregated immediately. The broadcaster is informed of the listeners' results. The results show that 67.5% of viewers prefer A, 20.1% want to hear C, and the other options have less than 10%. The broadcaster decides to play song A for the viewers.

[0324] Cross-media location management of user attention Obtaining real-time user attention metrics opens up numerous possibilities. In this section, we'll outline several use cases illustrating what can be achieved when this attention information is combined with contextual information and / or shared across devices, media types, etc. We'll primarily focus on location as one form of contextual information. However, other types of contextual information are also applicable, such as weather, local holidays, and events. Many of the examples covered in this section revolve around a core idea: continuously tailoring experiences across devices, media types, and locations using attention information detected by ATS.

[0325] This section covers user attention use cases across media and location. We use "content" to mean any form of consumable media, such as TV programming, radio, bus stop advertising, billboards, satellite navigation, etc. In addition to traditional media forms, we also consider driving or riding in a vehicle as a form of content that users can pay attention to. We discuss what can be achieved when sharing attention information across different media types. When we discuss using attention information across media types, we are referring to the idea of ​​combining user reactions with what we see in other media. This combination can help to more robustly determine user preferences and interests. Furthermore, such an approach can allow ATS-based systems to use previously determined user preferences and interests to guide decision-making across different media. Geographic information is a valuable resource, providing an additional dimension to geographic actions such as recommendations for shops, restaurants, and events. In some examples, information collected using ATS can be combined with location information to enrich decision-making and / or create new applications that would otherwise be impossible.

[0326] This system has at least three elements that can help us implement the use cases outlined below. They are: - Information obtained This can include content sets, geographic information, and detected attention information. For example, a billboard is a form of content, the user's location is a form of geographic information, and the user's attention is captured by the ATS.

[0327] - Equipment Platform Example device platforms include automobiles, TVs, mobile devices, etc.

[0328] - action Recommendations (such as those discussed in the 'Recommendations' section), highlighting, purchasing, etc.

[0329] Based on these three elements, the following are different scenarios and use cases.

[0330] Example use cases Tracking user preferences across devices and media As detailed in the 'Determining User Preferences' section, individual user interests and preferences can be tracked using ATS-enabled devices. This can be extended to determine user interests across multiple ATS-enabled devices and across media types. Furthermore, when tracking user preferences across devices and modes, the learned information about their interests can be used to guide decisions on new devices and / or media. User preferences can be determined on the user's device or in the cloud. In some examples, their preferences can be synchronized using content metadata streaming or other methods. Such preferences can also be stored in user profiles (e.g., Google profiles, Facebook profiles, Dolby ID, etc.) that are being used by the ATS runtime.

[0331] User preferences determined by ATS can also be transferred to devices that do not support ATS (e.g., using cloud-based profiles) to allow for a personalized experience when ATS is unavailable.

[0332] Specific examples: Jane watched a movie that included musical clips. She sang along to the songs, which was detected as positive attention to music on her (multiple) ATS-enabled devices. Her preferences were updated accordingly. The next day, while she was in her car, an automatically generated playlist considered including songs similar to those she had responded to positively the previous day.

[0333] In a virtual reality (VR) game, Bob decided to give his character a medieval appearance. The VR system, which supports ATS (Advanced Time and Space Technology), detected that Bob enjoyed the game more when he was in the medieval appearance. This suggests that Bob may have an interest in medieval themes. When Bob looked for his next audiobook to listen to, several audiobooks with medieval themes were added to his recommendation list.

[0334] When Steve was listening to a professional basketball game in his car, he seemed very unhappy every time the Lakers scored. When he got home and turned on the TV, it automatically switched to the channel that was playing the game and automatically selected a commentary stream designed to attract Warriors fans, because Steve didn't seem to like the Lakers.

[0335] Highlighting local interests based on user attention User interests and preferences determined by the ATS can be used to highlight nearby interests (e.g., landmarks, shops, places) to users (e.g., notifications, markers on maps). User interests and preferences can be determined using the methods described in the 'Determining User Preferences' or 'Tracking User Preferences Across Devices and Media' sections. Combining these user preferences with location information can enable recommendation-like systems. Some examples of such recommendations include highlighting a location on satellite navigation, notifying users that they are passing through a location that matches their interests, etc.

[0336] Specific examples: - A user has been detected as a fan of a particular artist. The user will receive a notification when driving past locations where the artist has previously performed or is about to perform.

[0337] - When users view a four-wheel drive vehicle exhibition, they show a particularly strong reaction to vehicles from a certain brand. When users drive past a dealership of that car, the dealership is highlighted on the map. The car's voice assistant will also proactively suggest arranging a test drive appointment.

[0338] - A user watched with great interest a series of historical programs about Pompeii. When the user drove past an exhibition about Pompeii, the vehicle notified the user, and the satellite navigation highlighted a nearby museum.

[0339] - A user showed interest in Omega watches while watching several movies and TV shows featuring Omega watches. When the user looked up GPS directions on his phone, he was shown the nearest Omega store.

[0340] Location-based ads that match user attention By combining location information with user attention information, new geo-related advertising strategies can be created and existing strategies improved. Current examples of location-based advertising include billboards and ads on public transportation. This strategy can help advertisers use ATS (Actions Based on Location) to determine the success of their real-world targeted advertising. Additionally, this strategy offers new advertising opportunities where ads can be virtually placed in specific locations.

[0341] Real-world location-based advertising can track attention using devices that support Attention Tracking System (ATS). For example, many current advertising strategies, such as billboards, are associated with location (e.g., where they are placed). Compared to generalized data like traffic data, ATS can further refine which geolocation-based locations will garner more attention. It can detect when a user is looking at the billboard and capture other user responses associated with it. Interactions with specific physical ads can be determined based on GPS data, camera data, microphone data, etc. Furthermore, physical ads can encourage attention responses through interactive mechanisms such as Q&A. Even if the answer is incorrect, an action can be taken after answering a Q&A question.

[0342] Virtual advertising can be implemented in a similar way to the ideas detailed in "Highlighting Local Interests Based on User Attention," except that local interests are driven by advertising. New advertising strategies are possible, such as: every car on the road could act as an advertisement for its own model, and if a user is detected paying attention to another car, it can be known that the user may be interested in that car. Alternatively, locations with geographically relevant ads that match user preferences could be marked on the user's map.

[0343] Users can connect with advertisers in the following situations, for example: -When a user interacts with one of the geo-linked ads; and - When a user is near a geo-related ad that indicates their interest based on previous attention indicators.

[0344] In some examples, successfully establishing customer contacts through ads that support ATS can enable automotive OEMs to receive referral bonuses from advertising sponsors.

[0345] Specific examples: A geotagged billboard ad poses a trivia question: "What was the first movie to use Dolby audio?" The answer is "A Clockwork Orange." A correct answer earns a free movie ticket; an incorrect answer displays a "Learn More" link.

[0346] A user walks past a bus with a humorous, location-linked ad on the side. The user is detected laughing using their ATS-enabled mobile device. The next ad that appears on their phone is the same product advertised on the bus.

[0347] - When users begin to discuss the topic presented by the billboard advertisement, they are detected as interacting with the ad. Statistics on billboard performance are sent back to advertisers or billboard managers to differentiate pricing based on audience attention.

[0348] Alice drove past an advertisement for a live music event. She responded positively to the advertisement and received an offer to buy tickets to the event.

[0349] - In the same example as above, Alice was not offered to buy tickets, but instead was offered to play some of the band's new music for herself in the car.

[0350] A user was driving when they saw an advertisement for a concert on a bus and said, "I didn't expect [name] to be having a concert at [location], I'll definitely be there." The ATS-enabled car captured the user's reaction to the advertisement using the car's current location. This was then used to suggest the user buy tickets.

[0351] - When a user drives past an Acme watch advertisement at a bus stop, the user says, "Acme watches are nice." Afterward, the next time the user approaches an Acme store, that store will be highlighted on the map.

[0352] - A musical paid for customized virtual georeferencing ads for its upcoming shows. If a user's interests align with the themes of the musical or an upcoming show, the location of the venue is pinned to their GPS navigation.

[0353] - Car brands might decide to feature their vehicles as interactive, location-linked ads. A user sees someone else driving their favorite car and says, "Those cars look good, don't they?" This is detected by their (multiple) ATS-enabled devices. The user then receives personalized ads based on their preference for that particular car model.

[0354] - Users glance at the billboards each time they drive by, which helps determine their potential interest in the billboard's advertising content.

[0355] Continuous attention and customization across devices and locations This section provides examples of various possibilities that can be achieved using combinations of the following parts: - Track user preferences across devices and media; - Highlighting local interests based on user attention; and - Location-based ads that align with user attention.

[0356] Dynamic continuous attention tracking optimization can be possible, for example: Device A detects the user's active attention to something of interest and transmits this information to other devices; Device B creates highlight reels of things that interest you; and Device C advertises things that the user is interested in.

[0357] Specific examples: Users see geotagged billboard ads for the watch while in their cars. Later, during their smart TV's advertising time slots, they see ads for the same watch brand again. Finally, when users search for a restaurant for lunch on their map app, the watch brand's store is highlighted on their map.

[0358] User attention-based navigation route planning The routes determined by the navigation system can use attention analysis to help guide decision-making. This can be short-term (e.g., taking users on scenic routes because they are in a low mood and not in a hurry) or long-term (e.g., optimizing city traffic based on conditions that generate the highest overall level of active driving attention). Routes can also be personalized to lead users along paths that better match their interests or preferences. For example, user interests and preferences can be determined using the methods described in the 'Determining User Preferences' or 'Tracking User Preferences Across Devices and Media' sections. Recommended route planning can also take into account the user's emotional state determined by the ATS to suggest departure times.

[0359] Specific examples: - Suburban traffic during peak hours is particularly problematic. Route planning is generated for each ATS-enabled user in the suburbs based on their preferences during peak hours. In some examples, the ATS-enabled navigation system can suggest that some people depart earlier in the morning, as they are known to typically wake up early, and vice versa. The navigation system can also suggest different routes based on the user's preference for certain roads and their proximity to their destination. In these ways, suburban traffic route planning can be optimized holistically.

[0360] - It could be suggested that the user depart a little earlier today, as traffic and weather conditions are favorable for driving. Furthermore, in this example, the user is in a good mood and may therefore be more inclined to depart earlier. Based on this reasoning, navigation route planning using ATS information would advise the user to take the earlier route.

[0361] - A user wants to drive to a destination, and there are two different routes that take similar amounts of time. Suggest the user take the route that has shops and / or advertisements that match their interests.

[0362] Mr. Zheng was planning a road trip. He wanted to take a "scenic" route. The navigation system, based on ATS information, suggested a route that would allow for the most coastal driving, as Mr. Zheng frequently watched ocean documentaries and surfing videos.

[0363] Yvonne is planning a similar road trip. Her "scenic" route will pass by top-rated restaurants and wineries, as she frequently watches cooking shows and talks about wine in the car.

[0364] Using ATS knowledge-based Q&A Many knowledge-based question-and-answer experiences are possible through the use of ATS. Note that knowledge-based question-and-answer is a content format. Context-sensitive knowledge-based questions can be optionally obtained by leveraging location information. Knowledge-based question-and-answer questions can also be provided by advertisers.

[0365] Specific examples: As John drives his ATS-enabled vehicle through town, a game similar to a spot-check game automatically generates trivia questions for him. Nearby landmarks and towns can serve as the basis for trivia questions, such as "When was this town founded?" - In a similar scenario, a subtle form of advertising was mixed into the knowledge-based quiz questions John received. For example, "What is the only watch brand for...?" The answer could be a watch brand that John prefers, such as Acme.

[0366] - A quiz game entirely centered around the advertisement could be offered to the user. This game could serve as entertainment and / or potentially offer rewards. An example reward after completing a quiz ad segment might be that the user no longer receives any ads on any of their devices for the rest of the day.

[0367] Geographic sentiment mapping based on user attention Combining attentional and location information allows for mapping regions to emotions on a per-user basis. This type of mapping can prove to provide insights for geographic decision-making. In some examples, computer programs can post-process this mapping to help users digest the information. For instance, the program can create summaries, tag locations, etc.

[0368] Specific examples: A user looking to buy a new home is trying to identify an area that might be a good fit. Geographic sentiment mapping data can prove useful because it can provide insights into areas where the user previously responded positively (e.g., laughing, excitedly talking), rather than areas where they responded negatively (e.g., swearing, complaining). This can be used to provide the user with a list of recommended areas where they might want to buy a home.

[0369] A couple is considering buying a house. Each time they view a property, they discuss it in their car as they drive away. The mood at each location is tracked by their vehicle's ATS (Automatic Vehicle System). When the couple is ready to make a final decision, they can see the ATS information displayed to them in the form of a dashboard, heatmaps, and location markers for the places they visited during their house viewings. Using this information summary, the couple can be able to make a more informed decision based on how they feel about a potential home purchase, without having to meticulously record their previous impressions.

[0370] Enhance safety and well-being through attentional feedback. In recent years, products targeting well-being have become increasingly popular. Apps designed to monitor and enhance user well-being will greatly benefit from information that strongly reflects a user's current state and infers their level of happiness. This type of information can be collected from attention tracking systems (ATS).

[0371] Attention analytics can be used to enhance user safety. Security systems can glean information from various types of attention data, such as mood, focus, and positive / negative emotions. In this section, we will provide examples of how Attention Strategies (ATS) technology can improve user safety and well-being.

[0372] Example use cases Using ATS to detect user speech as a proxy for attention An in-vehicle ATS (Automatic Driver Response System) can be used to determine whether the driver is focused on driving and is alert. This assessment can be performed in real time and / or improve over time as the ATS-enabled system learns about user behavior. Understanding user behavior can help determine what is normal for a particular user and detect changes from that normality. For example: Singing along to music playing in the car or laughing at jokes on podcasts can indicate that the driver is alert and conscious.

[0373] • Swearing can indicate that a driver is frustrated or suffers from "road rage" and may make unwise decisions.

[0374] • It is known that a user frequently uses profanity. When this user's profanity is detected, it is less likely to be attributed to "road rage".

[0375] Snoring can indicate that the driver is asleep! • Blinking or nodding can indicate that the driver is about to fall asleep.

[0376] Adaptive acoustic driver mindfulness application Several smartphone applications targeting mindfulness, meditation, sleep, calmness, and psychological well-being exist in the prior art. Some publicly available examples provide mindfulness apps configured to enhance focus on tasks such as driving a car. In the proposed technology and system, feedback from the ATS is used to guide the playback of various acoustic responses (e.g., from the in-vehicle audio system).

[0377] Here are some examples of voice responses that drivers can hear from this system: • “I noticed you were swearing at other drivers much more than usual this morning. You sound a bit frustrated and not focused on driving. Here are some soothing whale songs to help you concentrate.” Gamification enhances driver focus. In this example, drivers could be required to participate in simulated knowledge quizzes or other games designed to reward focused driving. Insurance companies could offer reduced car insurance rates to drivers who consistently score well in such games.

[0378] For example, drivers can be periodically asked to answer questions about road conditions, such as: • What color was the car that just merged from the right lane? • What is the current speed limit? • "How many meters after which will the left lane end?" These examples can be extended to other use cases, such as operating heavy machinery and air traffic control.

[0379] Automatic rest suggestion Systems supporting ATS can be configured to detect when a driver becomes increasingly distracted, increasingly angry, increasingly fatigued, or less attentive to the road over time. In some examples, such systems can be configured to detect driver fatigue, anger, or less attentiveness to the road over time using methods specified in 'Detecting User Verbal Awareness as a Proximity Agent via ATS'. Based on this, the system can suggest that the driver take a break to ensure their own safety and the safety of passengers (if any).

[0380] This system can also personalize recommendations, for example, as described in the sections on 'Personalization and Enhancement Using Attention Feedback', 'Long-Term Personalization', and 'Recommendations,' thus showing users more preferred rest options. This might include things like favorite food, coffee, service stations, etc. Preferences can be determined using a user's attention to driving after resting at a specific location. Insurance companies can offer reduced car insurance rates to drivers who consistently follow such recommendations and do not drive in a damaged state.

[0381] Keep the children in the back seat quiet. Children shouting or playing in the back seat during a car trip can distract the driver. In some examples, the Automatic Attention System (ATS) can be configured to detect the sounds of shouting, playing, or bored children and automatically switch the content played by the in-car audio system to something more suitable for the child's entertainment. In some examples, the ATS can be configured to ensure that a specific target attention level is reached (e.g., indicated by laughing at a joke, singing along, or responding to a call-and-response segment of a children's podcast). If the child's attention level falls below a certain point, in some examples the system can automatically try different content until the children quiet down, thus helping the driver to refocus.

[0382] Using discourse as a proxy to measure mental health An individual's words and other responses, analyzed over hours or days, can indicate their mental health. For example, an ATS (Attention Scale) configured to listen for laughter can calculate an average weekly number of laughs. This can be converted into a score indicating a person's general mental health or overall well-being. Deviations from a user's previous attention patterns or averages can be detected, suggesting a change in the user's well-being.

[0383] Other examples include: Frequent swearing can indicate poor mental health.

[0384] • A monotonous way of speaking can indicate poor mental health, especially when compared to a baseline of an individual's usual way of speaking.

[0385] The frequency with which someone smiles can indicate their level of happiness.

[0386] • The likelihood that a person is singing along to an uplifting song can indicate good mental health.

[0387] Figure 12 This is a flowchart outlining an example of the disclosed method. Method 1200 can be, for example, by... Figure 1A The control system 160, by Figure 2 The control system 160 or one of the other control system examples disclosed herein may execute the method. As with other methods described herein, the blocks of method 1200 need not be executed in the indicated order. According to some examples, one or more blocks may be executed in parallel. Furthermore, some similar methods may include more or fewer blocks than those shown and / or described.

[0388] Method 1200 can be comprised of, including Figure 1A Control system 160 Figure 2 The method 1200 is performed by an apparatus or system of control system 160 or one of other examples of control systems disclosed herein. In some examples, blocks of method 1200 may be performed by one or more devices within an audio environment, such as an audio system controller (e.g., a device that may be referred to herein as a smart home hub) or by another component of the audio system (e.g., a television, television control module, laptop computer, game console or system, mobile device (e.g., a cellular phone), etc.). However, in some embodiments, at least some blocks of method 1200 may be performed by one or more devices (e.g., one or more servers) configured to implement cloud-based services.

[0389] In this example, box 1205 relates to the control system acquiring sensor data from a sensor system during content presentation. Content presentation can be, for example, a television program, a movie, an advertisement, music, a podcast, a game session, a video conferencing session, an online learning course, etc. In some examples, the control system may acquire sensor data from one or more sensors of sensor system 180, as disclosed herein in box 1205. Sensor data may include sensor data from one or more microphones, one or more cameras, one or more eye trackers configured to collect gaze and pupil size information, one or more ambient light sensors, one or more thermal sensors, one or more sensors configured to measure skin conductance responses, etc.

[0390] According to this example, box 1210 relates to estimating a user response event by the control system based on sensor data. In some examples, box 1210 may be performed by one or more Device Analysis Engines (DAEs). The response event estimated in box 1210 may include, for example, detected phonemes, emotion type estimates, heart rate estimates, body posture estimates, one or more latent spatial representations of sensor signals, etc.

[0391] In this example, box 1215 relates to generating user attention analysis by a control system based at least in part on estimated user response events corresponding to estimated user attention to content intervals presented with the content. In some examples, box 1215 may be performed by an attention analysis engine. A “content interval” may correspond to a time interval, which may or may not have a uniform size. In some examples, a content interval may correspond to a uniform time interval of approximately 1 second, 2 seconds, 3 seconds, etc. Alternatively or additionally, a content interval may correspond to an event in the content presentation, such as an interval corresponding to the presence of a particular actor, the presence of a particular object, the presentation of a theme (e.g., a musical theme), a joke, a dramatic event, a romantic event, a scary event, a beautiful event, etc.

[0392] According to this example, box 1220 relates to modifying content presentation by a control system, at least in part, based on user attention analysis. In some examples, box 1220 may be performed by a content presentation module based on input from an attention analysis engine. Modifying content presentation may, for example, involve adding or changing a laughter track, adding or changing audio corresponding to the response of one or more other persons to the content presentation (or similar content presentation), extending or shortening the duration of at least a portion of the presented content, adding, removing, or changing at least a portion of the visual content presentation and / or audio content presentation, etc. In some examples, modifying one or more aspects of the audio content may involve adaptively controlling an audio enhancement process, such as a dialogue enhancement process. According to some examples, modifying one or more aspects of the audio content may involve modifying one or more spatialization properties of the audio content. In some examples, modifying one or more spatialization properties of the audio content may involve rendering the at least one audio object at a location different from where it would originally be rendered.

[0393] In some instances, content presentation can be modified, at least in part, based on one or more user preferences or other user characteristics. Based on some examples, modifying content presentation can involve personalizing or enhancing the content presentation, for example, based at least in part on one or more user preferences or other user characteristics. In some examples, personalizing or enhancing content presentation can involve changing audio playback volume, one or more other audio characteristics, or a combination thereof. According to some examples, personalizing or enhancing content presentation can involve changing one or more display characteristics, such as brightness, contrast, etc. In some examples, personalizing or enhancing content presentation can involve changing the storyline, adding characters or other story elements, changing the time period involving characters, changing the time period dedicated to another aspect of the content presentation, or a combination thereof.

[0394] Personalizing or enhancing content presentation, as illustrated by some examples, can involve providing personalized advertising content. In some such examples, providing personalized advertising content can involve offering advertising content that corresponds to an estimated user attention level for one or more content segments involving one or more products or services.

[0395] In this example, box 1225 relates to the control system causing the provision of modified content presentation. In some examples, box 1225 may relate to providing modified content presentation on a television screen, one or more speakers, etc.

[0396] Some aspects of this disclosure include a system or apparatus configured (e.g., programmed) to perform one or more examples of the disclosed methods, and a tangible computer-readable medium (e.g., a disk) storing code for implementing one or more examples of the disclosed methods or steps thereof. For example, some disclosed systems may be or include a programmable general-purpose processor, digital signal processor, or microprocessor programmed with software or firmware and / or otherwise configured to perform any of a variety of operations on data, including embodiments of the disclosed methods or steps thereof. Such a general-purpose processor may be or include a computer system including input devices, memory, and a processing subsystem programmed (and / or otherwise configured) to perform one or more examples of the disclosed methods (or steps thereof) in response to data asserted thereto.

[0397] Some embodiments may be implemented as a configurable (e.g., programmable) digital signal processor (DSP) configured (e.g., programmed and otherwise configured) to perform necessary processing on (multiple) audio signals, including execution of one or more examples of the disclosed methods. Alternatively, embodiments of the disclosed system (or elements thereof) may be implemented as a general-purpose processor (e.g., a personal computer (PC) or other computer system or microprocessor that may include input devices and memory) programmed with software or firmware and / or otherwise configured to perform any of a variety of operations, including one or more examples of the disclosed methods. Alternatively, elements of some embodiments of the system of the invention are implemented as general-purpose processors or DSPs configured (e.g., programmed) to perform one or more examples of the disclosed methods, and the system also includes other elements (e.g., one or more speakers and / or one or more microphones). A general-purpose processor configured to perform one or more examples of the disclosed methods may be coupled to input devices (e.g., a mouse and / or keyboard), memory, and a display device.

[0398] Another aspect of this disclosure is a computer-readable medium (e.g., a disk or other tangible storage medium) that stores code (e.g., an encoder executable to perform one or more examples of the disclosed methods or steps) for performing the disclosed methods or steps.

[0399] While specific embodiments and applications of this disclosure have been described herein, it will be apparent to those skilled in the art that many changes can be made to the described embodiments and applications without departing from the scope of this disclosure as described and claimed herein. It should be understood that although certain forms of this disclosure have been shown and described, this disclosure is not limited to the specific embodiments described and shown or the specific methods described.

[0400] Various aspects of this disclosure can be understood from the following enumerated example embodiments (EEE): EEE 1. A system comprising: Main unit; loudspeaker system; Sensor systems; and The control system includes: One or more device analytics engines, configured to estimate user response events based on sensor data received from the sensor system; and A user attention analysis engine is configured to generate user attention analysis based at least in part on estimated user response events received from the one or more device analysis engines, the user attention analysis corresponding to estimated user attention to content intervals presented via the host unit and the speaker system. The control system is configured as follows: The content presentation is altered, at least in part, based on the user attention analysis; and The modified content is presented by the host unit, the speaker system, or both the host unit and the speaker system.

[0401] EEE 2. The system as described in EEE 1 further includes an interface system configured to provide communication between the control system and one or more other devices via a network, wherein the modified content presentation is or includes modified content received from the one or more other devices via the network.

[0402] EEE 3. The system as described in EEE 2, wherein causing the content presentation to be modified involves sending user attention analysis from the user attention analysis engine via the interface system and receiving the modified content in response to the user attention analysis.

[0403] EEE 4. The system of any one of EEE 1 to 3, wherein changing the content presentation involves personalizing or enhancing the content presentation.

[0404] EEE 5. A system as described in EEE 4, wherein personalizing or enhancing the presentation of the content involves changing one or more of the following: audio playback volume, audio rendering position, or one or more other audio characteristics, or a combination thereof.

[0405] EEE 6. A system as described in EEE 4 or EEE 5, wherein the host unit is a television or includes a television, and wherein personalizing or enhancing the presentation of the content involves changing one or more television display characteristics.

[0406] EEE 7. The system of any one of EEE 3 to 6, wherein personalizing or enhancing the presentation of the content involves changing the storyline, adding characters or other story elements, changing the time interval involving the characters, changing the time interval dedicated to another aspect of the presentation of the content, or a combination thereof.

[0407] EEE 8. The system of any one of EEE 3 to 7, wherein personalizing or enhancing the presentation of the content involves providing personalized advertising content.

[0408] EEE 9. A system as described in EEE 8, wherein providing personalized advertising content involves providing advertising content corresponding to an estimated user attention to one or more content segments relating to one or more products or services.

[0409] EEE 10. The system as described in any one of EEE 3 to 9, wherein personalizing or enhancing the presentation of the content involves providing or altering a laughter soundtrack.

[0410] EEE 11. The system as described in any one of EEE 1 to 10, wherein the host unit, the speaker system, and the sensor system are in a first environment; and The control system is further configured to modify the content presentation based at least in part on sensor data corresponding to one or more other environments, estimated user response events, user attention analysis, or a combination thereof.

[0411] EEE 12. The system of any one of EEE 1 to 11, wherein the control system is further configured to cause the content to be paused or replayed.

[0412] EEE 13. The system as described in any one of EEE 1 to 12, wherein the sensor system includes one or more cameras, and wherein the sensor data includes camera data.

[0413] EEE 14. The system of any one of EEE 1 to 13, wherein the sensor system includes one or more microphones, and wherein the sensor data includes microphone data.

[0414] EEE 15. The system as described in EEE 14, wherein the control system is further configured to implement an echo management system to mitigate the effects of audio played back by the speaker system and detected by the one or more microphones.

[0415] EEE 16. The system as described in any one of EEE 1 to 15, wherein one or more first portions of the control system are deployed in a first environment, and a second portion of the control system is deployed in a second environment.

[0416] EEE 17. The system as described in EEE 16, wherein one or more first portions of the control system are configured to implement the one or more device analysis engines, and wherein a second portion of the control system is configured to implement the user attention analysis engine.

[0417] EEE 18. A system as described in EEE 4 or EEE 5, wherein the host unit is a digital media adapter or includes a digital media adapter.

Claims

1. A system comprising: Main unit; loudspeaker system; Sensor systems; as well as The control system includes: One or more device analytics engines, configured to estimate user response events based on sensor data received from the sensor system; and A user attention analysis engine is configured to generate user attention analysis based at least in part on estimated user response events received from the one or more device analysis engines, the user attention analysis corresponding to estimated user attention to content intervals presented via the host unit and the speaker system. The control system is configured as follows: The content presentation is altered, at least in part, based on the user attention analysis; and The modified content is presented by the host unit, the speaker system, or both the host unit and the speaker system.

2. The system of claim 1, further comprising: An interface system configured to provide communication between the control system and one or more other devices via a network, wherein the modified content presentation is or includes modified content received from the one or more other devices via the network.

3. The system as described in claim 2, wherein, The alteration of the content presentation involves sending user attention analysis from the user attention analysis engine via the interface system and receiving the altered content in response to the user attention analysis.

4. The system as described in any one of claims 1 to 3, wherein, Changing the presentation of the content involves personalizing or enhancing the presentation of the content.

5. The system as described in claim 4, wherein, Personalizing or enhancing the presentation of the content involves changing the audio playback volume, audio rendering position, one or more of other audio features, or a combination thereof.

6. The system as claimed in claim 4 or claim 5, wherein, The host unit is a television or includes a television, and wherein personalizing or enhancing the presentation of the content involves changing one or more television display characteristics.

7. The system as claimed in any one of claims 3 to 6, wherein, Personalizing or enhancing the presentation of the content involves changing the storyline, adding characters or other story elements, changing the time period involving the characters, changing the time period dedicated to another aspect of the presentation of the content, or a combination thereof.

8. The system as claimed in any one of claims 3 to 7, wherein, Personalizing or enhancing the presentation of the content involves providing personalized advertising content.

9. The system of claim 8, wherein, Providing personalized advertising content involves providing advertising content that corresponds to an estimated user attention level for one or more content segments involving one or more products or services.

10. The system as claimed in any one of claims 3 to 9, wherein, Personalizing or enhancing the presentation of the content involves providing or changing the laughter soundtrack.

11. The system as claimed in any one of claims 1 to 10, wherein, The host unit, the speaker system, and the sensor system are located in the first environment; and The control system is further configured to modify the content presentation based at least in part on sensor data corresponding to one or more other environments, estimated user response events, user attention analysis, or a combination thereof.

12. The system as claimed in any one of claims 1 to 11, wherein, The control system is further configured to pause or replay the content.

13. The system as claimed in any one of claims 1 to 12, wherein, The sensor system includes one or more cameras, and the sensor data includes camera data.

14. The system as claimed in any one of claims 1 to 13, wherein, The sensor system includes one or more microphones, and the sensor data includes microphone data.

15. The system of claim 14, wherein, The control system is further configured to implement an echo management system to mitigate the effects of audio played back by the speaker system and detected by the one or more microphones.

16. The system as claimed in any one of claims 1 to 15, wherein, One or more first parts of the control system are deployed in a first environment, and a second part of the control system is deployed in a second environment.

17. The system of claim 16, wherein, The first portion of the control system is configured to implement the one or more device analysis engines, and the second portion of the control system is configured to implement the user attention analysis engine.

18. The system as claimed in claim 4 or claim 5, wherein, The host unit is a digital media adapter or includes a digital media adapter.