System and method for assessing viewer engagement
Patent Information
- Application Number
- JP2023090934
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2021-12-09
- Filing Date
- 2023-06-01
- Publication Date
- 2025-07-10
- Estimated Expiration
- 2042-11-18
AI Technical Summary
Conventional methods of television audience measurement fail to accurately determine viewer engagement, as they only count the number of people in the room and do not account for actual viewing time, response to programs or advertisements, and lack demographic specificity.
A system that uses cameras and microphones to capture images and audio data from a viewing area, processing this data locally to determine the number of viewers and their engagement levels, including facial recognition and sentiment analysis, and transmitting processed data to a remote server for statistical analysis across households.
Provides accurate, granular data on viewer engagement, allowing content providers to optimize programming and advertising based on demographic-specific metrics, improving ROI and enhancing viewer data privacy through local processing and minimal data transmission.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
Background Art
[0001] [Cross - Reference to Related Applications] This application is a partial continuation of U.S. Patent Application No. 15 / 702,229, a bypass continuation of PCT Application No. PCT / US2,017 / 012,531, filed on January 6, 2017, titled "SYSTEMS AND METHODS FOR ASSESSING VIEWER ENGAGEMENT", which is hereby incorporated by reference in its entirety, and claims priority to U.S. Application No. 62 / 275,699, filed on January 6, 2016, titled "SYSTEMS AND METHODS FOR ASSESSING VIEWER ENGAGEMENT", which is also hereby incorporated by reference in its entirety.
[0002] Conventional methods for measuring television viewers include collecting data from viewers using devices such as people meters and diaries. These methods generally attempt to identify the people (potential viewing members) in the room where the television is located. This method can also include capturing a series of images (e.g., television programs or commercial advertisements) being played on the television. And for each image, the number of people in the room at the time a particular image is displayed can be estimated.
[0003] These methods have several drawbacks. First, the data collected by these methods usually only includes the number of people in the room where the television is located. This data usually does not indicate how often the viewer is actually watching the television (the measurement is taken when the television is on). Second, the collected data can show how often a person tunes to a particular channel. However, since it does not measure the reaction to a program or advertisement, it does not indicate the effectiveness of the program or advertisement. Third, the television ratings are not given for a particular demographic group within a household or community.
Brief Description of the Drawings
[0004] Those skilled in the art will understand that the drawings are primarily for illustrative purposes and are not intended to limit the scope of the subject matter of the inventions described herein. The drawings are not necessarily to scale, and in some instances, different aspects of the subject matter of the inventions described herein may be exaggerated or enlarged in multiple drawings to facilitate the understanding of multiple different features. In the drawings, similar reference numerals generally refer to similar features (e.g., functionally similar and / or structurally similar elements).
[0005] [Figure 1] This figure shows a schematic diagram of a system for evaluating the degree of viewer engagement among television viewers.
[0006] [Figure 2A] This figure shows a method for quantifying user involvement using the system shown in Figure 1.
[0007] [Figure 2B] This diagram shows how to train a computer vision model to quantify user involvement.
[0008] [Figure 3A] This diagram shows methods for assessing viewer engagement, including face and eye tracking, facial recognition, and sentimental analysis.
[0009] [Figure 3B] This diagram illustrates the concepts of visibility index and attention-grabbing index.
[0010] [Figure 4A] This diagram shows the process for evaluating viewer engagement, including the estimation of the visibility index.
[0011] [Figure 4B] This diagram shows the process for evaluating audience engagement, including the estimation of an attention index.
[0012] [Figure 5] A diagram showing a process for evaluating viewer engagement, including determining the orientation of each person's face within the viewing area.
[0013] [Figure 6] A diagram showing a process for detecting skeleton, face, identification information, emotion, and engagement.
[0014] [Figure 7] A diagram showing a schematic of a data acquisition architecture in an exemplary method for evaluating viewer engagement.
[0015] [Figure 8A] A diagram showing a commercial message (CM) curve obtained using the architecture shown in FIG. 7. [Figure 8B] A diagram showing a commercial message (CM) curve obtained using the architecture shown in FIG. 7. [Figure 8C] A diagram showing a commercial message (CM) curve obtained using the architecture shown in FIG. 7. [Figure 8D] A diagram showing a commercial message (CM) curve obtained using the architecture shown in FIG. 7. [Figure 8E] A diagram showing a commercial message (CM) curve obtained using the architecture shown in FIG. 7. [Figure 8F] A diagram showing a commercial message (CM) curve obtained using the architecture shown in FIG. 7. [Figure 8G] A diagram showing a commercial message (CM) curve obtained using the architecture shown in FIG. 7.
[0016] [Figure 9] A diagram showing the ratio of each CM curve among the sampled TV stations.
[0017] [Figure 10]A diagram showing a classification model using a decision tree, and the determination results of the decision tree are shown in Table 5.
[0018] [Figure 11] A diagram showing the visibility rate with respect to the length of the CM.
[0019] [Figure 12] A diagram showing the correlation between the elapsed time from the start of the program and the visibility rate.
[0020] [Figure 13] A diagram showing the communication of viewer engagement data obtained using the technologies shown in FIGS. 1 to 12.
[0021] [Figure 14] A diagram showing the spread and use of viewer engagement data obtained using the technologies shown in FIGS. 1 to 12.
[0022] [Figure 15] A diagram showing the big data analysis and visualization of viewer engagement data obtained using the technologies shown in FIGS. 1 to 12.
[0023] [Figure 16] A diagram showing a model for obtaining additional data for complementing viewer engagement data obtained using the technologies shown in FIGS. 1 to 12. [Figure 17] A system diagram including a packet inspection module. [Figure 18] A system diagram further showing the functions of the packet inspection module.
Modes for Carrying Out the Invention
[0024] The systems and methods disclosed herein acquire image data of a viewing area in front of a display (e.g., a television, computer, or tablet) playing video (e.g., a television show, a movie, a web show, an advertisement, or other content). An exemplary system uses the image data to determine how many people are in the viewing area and which of them are actually watching the video. The system also samples the soundtrack of the video with a microphone and uses the soundtrack samples to identify the video. The system stores (and / or retains) information about the video, the number of people in the viewing area, and the number of people watching the video in local memory and transmits the information to a remote server via the Internet or other network connection.
[0025] Embodiments of the present invention include apparatus, systems, and methods for evaluating the degree of viewer engagement of television viewers. In one example, a system for quantifying viewer engagement with an image being played on a display includes at least one camera positioned to image a viewing area in front of the display and to acquire image data of the viewing area. A microphone is positioned close to the display and to acquire audio data emitted by a speaker coupled to the display. The system further includes a memory operably coupled to the camera and the microphone for storing processor-executable instructions, and a processor operably coupled to the camera, microphone, and memory. When a processor-executable instruction is executed, the processor receives image data from the camera and audio data from the microphone and determines identification information of the image displayed on the display based on at least a portion of the audio data. The processor also estimates, at least partially, the number of people present in the viewing area and the number of people engaged with the image in the viewing area based on at least a portion of the image data. The processor further quantifies the degree of viewer engagement with the image based at least partially on the first and second numbers.
[0026] In another example, a method for quantifying viewer engagement with an image displayed on a display includes the step of acquiring an image of the viewing area in front of the display with at least one camera while the image is displayed on the display. The method also includes the step of acquiring audio data representing the soundtrack of the image emitted by a speaker coupled to the display with a microphone. The method further includes the steps of receiving image data from the camera and audio data from the microphone, determining identification information of the image displayed on the display with a processor operably coupled to the camera and microphone based on at least a portion of the audio data, and estimating with the processor a first number of people present in the viewing area and a second number of people engaged with the image in the viewing area while the image is displayed on the display, based on at least a portion of the image data. The method also includes the step of transmitting the image identification information, the first number of people and the second number of people to a remote server with the processor.
[0027] In yet another example, a system for evaluating viewer engagement with an image being played on a display is disclosed. The display is coupled to a speaker that emits a soundtrack of the image. The system includes a visible camera that acquires a visible image of the viewing area in front of the display at a first sampling rate while the display is playing an image. An infrared camera is included in the system to acquire an infrared image of the viewing area in front of the display at a first sampling rate while the display is playing an image. A microphone is positioned close to the display to acquire samples of the soundtrack emitted by the speaker at a second sampling rate lower than the first sampling rate while the display is playing an image. The system also includes a processor operably coupled to the visible camera, infrared camera and microphone that (i) identifies the image based on the soundtrack samples, (ii) estimates the number of people in the viewing area and the number of people engaged with the image based on the visible and infrared images while the display is playing an image, and (iii) overwrites, erases and / or discards samples of the soundtrack, visible image and infrared image. The system also includes memory operably coupled to the processor, which stores representations of video identification information, the number of people in the viewing area while the video is being played on the display, and the number of people involved in the video. The system further includes a network interface operably coupled to the processor for sending representations to a server.
[0028] In yet another example, a method for quantifying viewer engagement with a unique image in multiple videos includes the steps of acquiring image data of the viewing area in front of the display in each of several households, and determining whether the display is showing an image in the multiple videos. The method also includes the steps of estimating (i) viewership and (ii) attendance rates for each unique image in the multiple videos, based on the image data and demographic information about each of the several households. Viewership represents the ratio of the total number of people in the viewing area to the total number of displays showing the video, and attendance rate represents the ratio of the total number of people in households with displays showing the video to the total number of people in several households. The method also includes the step of determining a visibility index for each unique image in the multiple videos, based on the viewership and attendance rates.
[0029] It should be understood that all combinations of the aforementioned concepts and additional concepts discussed in more detail below (provided that such concepts are not mutually contradictory) are intended to be part of the subject matter of the inventions disclosed herein. In particular, all combinations of claimed subject matter appearing at the end of this disclosure are considered to be part of the subject matter of the inventions described herein. It should also be understood that terms used expressly herein and incorporated by reference into any disclosure should be given meanings that best correspond to several specific concepts disclosed herein.
[0030] Unlike conventional systems that measure viewer engagement with video by identifying video based on a digital watermark embedded in the video itself, an example of the system of the present invention identifies video based on the video's soundtrack. As a result, the system of the present invention does not need to be connected to the viewer's display, set-top box, or cable connection. This makes installation and removal easier (and therefore more readily adopted). It also reduces the likelihood of recording the impression of malfunction or "false detection" caused by leaving the set-top box on while the display is powered off.
[0031] Furthermore, the system of the present invention processes image data locally, i.e., within the viewer's home, to determine the number of people in the viewing area and the number of people involved in the video. It can also identify when people are present in the viewing area by processing audio data locally. This data is stored locally, i.e., in the memory of a local device in the viewer's home, or in memory coupled to it. Since the processed image and audio data consumes far less memory than the raw image and audio data, this local memory can store information over a longer period. In other words, the device of the present invention uses memory more efficiently because it stores processed data instead of raw data.
[0032] The local device processes raw image data, which may include both visual and depth information, acquired from the viewing area to assess viewer engagement. The local device can analyze viewer gestures, movements, and facial orientation using artificial intelligence (AI) and machine learning techniques. It can also recognize individual faces of video viewers and determine each viewer's emotions from the image data. In this process, personal images are not transmitted outside the individual's home. Recognition can be performed on a local device within the home. Each individual in a household can receive a unique identifier during the household's onboarding process. If a match is found in the recognition process, this identifier can be assigned to the match and then sent to a remote server. Furthermore, the processing is performed over streaming video or audio data (including images). In other words, the video or audio data is not stored in local memory.
[0033] The local device processes raw audio data by matching or comparing it to samples in an audio database to identify the specific video being viewed (e.g., a television channel, program, or advertisement). Alternatively or additionally, the local device may submit queries based on the audio data to a third-party application programming interface (API) that identifies and returns identifying information about the content to which the audio belongs. In some cases, the database or API may return multiple possible matches, and the remote server may select the best match using information about the television schedule, subsequent audio samples, or data collected from other sources, including but not limited to the set-top box, cable / Internet connection, or the content provider itself.
[0034] In some implementations, the local device does not store raw image or audio data for later retrieval. Instead, the local device writes the raw image and audio data to one or more buffers for processing, and then overwrites or clears the buffers after the raw image and audio data has been processed. In other words, the local device simply temporarily holds the raw image and audio data while processing it. As used herein, "holding" raw image and audio data on the local device means temporarily storing this data for a short period of time (e.g., less than 100 milliseconds, less than 80 milliseconds, less than 60 milliseconds, less than 50 milliseconds, or less than 40 milliseconds, including any value and subranges in between). Overwriting or clearing raw image and audio data offers various advantages, such as reducing the amount of memory required by the local device. It can also facilitate compliance with personal data protection laws by eliminating image or audio data that could be used to identify individuals, including children, within the viewing area or microphone range.
[0035] Processing and storing image and audio data locally offers another technical advantage: it reduces the bandwidth required to transfer viewing habit information from the local device to a remote server. Processed image and audio data consumes less memory than raw image and audio data, thus requiring less bandwidth for transmission. Furthermore, processed image and audio data fills memory more slowly than raw image and audio data, allowing for less frequent transmission to the remote server. Local devices can leverage this flexibility by scheduling burst transmissions during periods of relatively low network bandwidth usage, such as late at night or early in the morning. Additionally, transmitting processed image and audio data that does not necessarily contain personally identifiable information, including children within the viewing area or microphone range, can ensure or facilitate compliance with personal data protection laws.
[0036] A remote server collects image and audio data processed from local devices in different households. By processing this data and statistically analyzing viewer engagement information collected from different households within the community, viewer engagement across the community is evaluated. For example, the server can quantify the ratio of viewer engagement to the total length of the detected program from the granular data collected from each household. In one embodiment, an audio fingerprint is created on the local device but is then matched against a database that is not permanently located on the local device. The audio fingerprint is generated on the local device from a 6-second audio segment. This fingerprint is then sent to the remote database for matching. The remote database may return 1 or 100 matches. (For example, an episode of The Simpsons may be played on one or more linear TV channels, multiple streaming services such as hulu.com, or a panelist may play it from a DVR device). All returned matches are stored on the local device. In one embodiment, a new audio fingerprint is created every 8 seconds and sent to the remote server for matching, and the process of receiving and storing matches is repeated. In one embodiment, the stored matches are uploaded to a remote data processing infrastructure approximately every hour. Here, a predictive algorithm is applied to the time series of matches uploaded by the local device. This algorithm includes the following: 1. For fingerprint matching, use contextual matching to predict the most likely match (recall from above that audio fingerprints can match multiple episodes (same theme music), multiple channels, and multiple streaming services). The goal is to determine as precisely as possible which channel or service presented which content to the audience. 2. Although data can be reported in seconds, voice fingerprints are taken every 8 seconds, and each fingerprint spans a period of 6 seconds, so this algorithm also determines the most likely match in the interleaved seconds and transmits over those durations.
[0037] Statistical analysis can further consider demographic information of the people and / or households watching the video (e.g., age, gender, household income, ethnicity, etc.). Based on all this information, the server can calculate various indices, such as the visibility index and the attention index (both defined below), to quantify viewer engagement. These viewer engagement indices can be based on any and all information provided by the local device, including viewer gestures, movements, viewer face orientation, and video information. These quantitative indicators can, among other things, show (i) who is actually watching the display, (ii) how often viewer members watch the display, and (iii) viewer reactions to programs and advertisements on the display.
[0038] Subsequently, quantitative metrics can be transferred by a remote server to central storage (e.g., a cloud-based database) where third parties, including but not limited to television advertising agencies and television networks, can access the metrics and, in some cases, other data. Alternatively, raw data collected by sensors may be analyzed in the manner described herein and transferred to central storage on the cloud, making it available to interested third parties. Third parties may optionally access the raw data through the system. In this example, raw data includes data collected after processing of video and audio streams (instead of the video and audio streams themselves). Generally speaking, raw data may include viewer unique identifiers, viewer attention, and the programs being watched by viewers, on a sub-second basis (e.g., every half-second or less). Using this raw data, more quantitative metrics (see below for further details) can be calculated on the remote server.
[0039] This acquired and analyzed data allows collecting entities, such as content providers or advertising agencies, to accurately assess the impact of their content, including unprecedented measures of individual demographics that could be valuable to advertisers. For example, advertising agencies can use this data to determine the best commercial slots for their target audience. With demographic information, the data can be matched to audience types and effectively drive purchasing decisions, thereby increasing the return on investment (ROI) of programs. Television networks can also benefit from the data, gaining a more accurate understanding of their programs, audience types, reactions, and advertising slot predictions. This allows them to determine the most valuable advertising slots for specific target demographic segments, further improve programs to better suit audience types, and eliminate unpopular shows.
[0040] Furthermore, the acquired and analyzed data enables a variety of business models. For example, the collecting entity can provide performance-based television viewership data and raw analytical data collected from motion-sensing devices installed in selected user homes representing national and / or regional demographics to television networks, advertising agencies, and other interested third parties, and indirectly to advertisers who acquire data from advertising agencies.
[0041] A system for evaluating viewer engagement
[0042] Figure 1 shows a schematic diagram of a system 100 for evaluating audience engagement in a home, sports bar, or other space with a display. The system 100 includes local devices 105 placed in each home to collect audience engagement data, and a remote server 170, such as a cloud storage and computing device, which includes memory for storing the data and a processor (also called a remote processor) for analyzing the data. The local devices 105 are communicably coupled to the remote server 170 via a network connection 172, such as an internet connection. For example, the local device 105 may include a network interface 165, such as a WiFi antenna or an Ethernet® port, for connecting to a home local area network (LAN). This LAN is further connected to a wide area network (WAN) via a cable or fiber optic connection provided by an Internet Service Provider (ISP).
[0043] The local device 105 in Figure 1 includes an IR radiator 110 to illuminate a viewing area 101 in front of a display 11, such as a television (TV), computer screen, tablet, or other device, with infrared (IR) light. This IR light can be structured or modulated to produce an illumination pattern that scatters or reflects off objects (including human viewers) within the viewing area 101. The local device 105 also includes an IR sensor 120 to detect the IR light reflected or scattered by these objects. A processor 150 (also called the local processor 150) coupled to the IR radiator 110 and the IR sensor 120 uses information about the illumination pattern and the detected IR light to generate one or more IR depth images or IR depth maps of the viewing area 101. More specifically, the processor 150 converts the information derived from the reflected beam into depth information that measures the distance between the viewer and the sensor 120. The processor 150 uses these IR depth images to determine how many people are in the viewing area and which of those people are looking at the display. Furthermore, the processor 150 may derive information from IR depth images relating to the identification information of people looking at the display, depending on the case, by recognizing their faces or postures, or by determining their demographic information (e.g., age, gender, etc.).
[0044] The local device 105 further includes an RGB sensor 130 (also called a visible camera) that captures a color image of the viewing area 101. The processor 150 is coupled to the RGB sensor and can use the color image alone or in combination with an IR depth image to estimate the number of people in the viewing area, the number of people involved in the display, and information about people in the viewing area. The color image may also be used for face recognition. In some cases, the processor 150 uses both the color image and the IR depth image to improve the fidelity of the estimation of the number of people in the viewing area and the number of people involved in the image.
[0045] The local device 105 also includes one or more microphones 140 positioned to detect sounds emitted by a speaker 13 coupled to the display 11. During operation, the speaker 13 plays a soundtrack of the image shown on the display 11. The microphones 140 also capture audio samples of the soundtrack played by the speaker 13. A processor 150 coupled to the microphones 140 uses these audio samples to create an audio fingerprint (soundtrack) of the image and identifies the image shown on the display 11 by comparing it with other audio fingerprints in a dedicated or third-party database. In one embodiment, the local device stores and runs a packet inspection module 1702, which is described in further detail below.
[0046] System 100 may further include a Bluetooth® receiver 180 matched to a Bluetooth® transmitter 185. In some cases, the Bluetooth® transmitter 185 may be included in a wristband or watch worn by the viewer. During operation, the Bluetooth® transmitter 185 transmits a low-power Bluetooth® beacon that is received by the Bluetooth® receiver 180. The processor 150 may then measure the viewer's distance from the display 11 based on the received Bluetooth® beacon. Each Bluetooth® transmitter 185 may also have a unique ID that can be recognized by the processor 150. The transmitter ID may further be associated with a unique viewer (for example, each viewer in a household may have their own transmitter). Thus, viewer identification information can also be determined.
[0047] In some cases, system 100 may include more than one Bluetooth® receiver. These receivers may be positioned at different locations so that each receiver can receive a different Bluetooth® signal strength from transmitter 185. This configuration allows processor 150 to estimate not only the distance from the viewer to display 11, but also the viewer's relative position (e.g., to the left or right of display 11).
[0048] System 100 may include other motion sensing devices, such as a three-axis accelerometer, for detecting position and motion. The motion sensing device may be connected to a data analysis and processing device, such as a desktop machine, for example, via a USB cable.
[0049] Figure 1 shows the data acquisition components, here, the IR radiator 110, IR sensor 120, RGB sensor 130, and microphone 140, as part of the local device 105 (e.g., within the same housing). In other embodiments, one or more of these components may be implemented as separate devices coupled to the processor 150 by one or more wired connections, such as USB, RS232, Ethernet®, fiber optic, or one or more wireless connections, such as WiFi, Bluetooth®, or other RF or infrared connections. For example, the IR radiator 110 and IR sensor 120 may be (or contained within) a commercially available device such as a Microsoft Kinect, which is connected to the processor 150. Similarly, the microphone 140 may be implemented as an array of microphones positioned around the listening area or near the speaker 13. The microphone array may more preferably be able to extract the audio input from ambient noise. The local device 105 may also include or be coupled to other sensors.
[0050] In system 100, the processor 150 is employed to process raw data acquired by sensors, including an IR radiator 110, an IR sensor 120, an RGB sensor 130, and a microphone 140. Processing can occur when processor-executable instructions stored in memory 160 coupled to the processor 150 are executed. In one example, the user can manually store instructions in memory 160 by downloading them from a remote server 170. In another example, the local device 105 may be configured to (periodically) check whether there are updated instructions available for download from the remote server 170. If so, the local device 105 can automatically download the updates via network connection 172 and network interface 165. In yet another example, the remote server 170 may be configured to send a notification to the local device 105 when it is ready to download an update or a new set of instructions. Upon receiving the notification, the user can decide whether to download and / or install the update. In yet another example, the remote server 170 may be configured to send the update notification to another user device, such as a smartphone. Upon receiving the notification, the user can decide whether or not to download and / or install the update.
[0051] Furthermore, the memory 160 in the local device 105 stores the processed data (e.g., estimation of the number of people in the viewing area, estimation of the number of people involved in the display, identification of images, and demographic information or indices derived from raw image and audio data). Once the memory 160 has accumulated sufficient processed data, the processor 150 sends the processed data to the remote server 170 via the network interface 165 and network connection 172 for aggregation, further processing, and reporting. The local memory 160 also temporarily holds image and audio data during local processing. In some cases, this processing is completed in less than a quarter of a second.
[0052] Collection and processing of image and audio data on local devices
[0053] Figure 2A shows a process 200 for collecting and processing image and audio data acquired by a system similar to the system 100 shown in Figure 1. As described above, the system may include a visible sensor, an IR sensor, or both to image the viewing area in front of the display (202). In one example, the RGB sensor 130 and the IR sensor 120 operate independently of each other, and the sensors acquire images asynchronously. In another example, image acquisition by the RGB sensor 130 and the IR sensor 120 is substantially synchronized. Whenever the RGB sensor 130 acquires a visible image, the IR sensor 120 acquires an IR image, for example, simultaneously or in an interleaved manner.
[0054] A local processor (e.g., processor 150) detects the number of people in the image of the viewing area (204) and determines which of those people are engaged with the display (206). For example, the local processor may use techniques known in the field of computer vision / image processing, including skeletal detection techniques, facial recognition techniques, and eye-tracking techniques, as described below. In some cases, the local processor 150 can determine (208) further indices that may be derived from audio data, such as described below, relating to the duration each viewer is in the viewing area, the duration each viewer is engaged with the display, and identification information of the displayed image (222).
[0055] The local processor can further identify each person detected as being in the viewing area 101 at a demographic level (e.g., males aged 25-30, girls aged 12-15) (210). If the local processor 150 has access to information about the household where the local device 105 is located, via local memory 160 or a remote server 170, it can use this demographic information to provide a more reliable estimate of demographic information for each person detected in the viewing area 101. Furthermore, the local processor can identify specific individuals within the viewing area within the household.
[0056] Furthermore, the local processor 150 can estimate the mood or emotion of each person detected in the viewing area 101 (212). The emotions that can be determined by the processor 150 may include, for example, happy, sad, or neutral. When viewing video on the display 11, classifying the viewer's emotions can be used to measure the viewer's reaction to the video, thus facilitating targeted advertising.
[0057] To estimate each person's mood or emotion, the local processor 150 can capture visual information in real time from both RGB and IR channels (e.g., from images of the viewing area 101). The visual information can be further processed to extract patterns and features that may be signatures of different mood or emotion states. Features extracted from both channels can be merged as a unified feature. A classifier can be trained to handle such features as input. The emotion / mood estimation can then be performed based on the classifier / response to a specific pattern at each time step.
[0058] In some cases, mood or emotion estimation can be achieved by the following method. The method includes a step of collecting training images of people exhibiting various emotions, such as smiling and frowning. Features representing each emotion are extracted from these training images (e.g., by a processor). The features and images are then used to train a classifier to correlate each feature with the corresponding emotion. In this way, the classifier can assign these features to various emotions. The method also includes a step of deploying the classifier to a local device to recognize the viewer's emotions in real time.
[0059] When the system collects visible and IR images synchronously, the visible and IR cameras can collect images to train a computer vision model used by the processor to detect people (204), count engaged viewers (206), demographically identify viewers (210), and estimate their mood (212). This training can be employed to establish "ground truth." By collecting image data from both IR and RGB sensors almost in parallel, a human can annotate the person detected in each image. This manual data can be fed into a training algorithm, resulting in two separate models, one trained on the visible RGB spectrum and the other on the IR spectrum. The detection rates of each model against "ground truth" are then compared to select the better-performing model. Further details of this training are described below with reference to Figure 2B.
[0060] Furthermore, the synchronization of the two cameras (e.g., sensors 120 and 130 in Figure 1) allows the local processor to double-check the image processing. For example, processor 150 can compare the number of people identified in each image, or remove errors where an individual is visible in one image but barely visible or invisible in the other. If the results match, processor 150 can record the results. Otherwise, processor 150 can detect a possible error in at least one of the images. Alternatively, processor 150 can generate a warning for human intervention. Processor 150 can also generate a flag indicating that the data estimated from these two images may be less reliable. In subsequent analysis, if images taken immediately before or after the problematic pair of images can provide reliable person recognition, this data may not be used at all.
[0061] In one example, the local device 105 always uses the visible sensor and IR sensors 120 and 130 to acquire image data. In another example, the local device 105 may acquire image data using only one of sensors 120 or 130. In yet another example, the local device 105 may use one sensor as the default sensor and the other as a backup sensor. For example, the local device 105 may use the RGB sensor 130 for imaging in most cases. However, if the processor 150 is unable to satisfactorily analyze the visible image (e.g., if the analysis is not as reliable as desired), the processor 150 may turn on the IR sensor 120 as a backup (or vice versa). This may occur, for example, when the ambient light level in the viewing area is low.
[0062] The local processor may also adjust the image acquisition rate of the visible sensor, the IR sensor, or both based on the number of people in the viewing area, their position within the viewing area, and identification information of the image on the display (214). Generally, the image acquisition rate of either or both sensors can be substantially equal to or greater than about 15 frames per second (fps) (e.g., about 15 fps, about 20 fps, about 30 fps, about 50 fps or greater, including any value and sub-range in between). At this image acquisition rate, the sensors can detect eye movements well enough for the local processor to assess viewer engagement (206).
[0063] The local processor may increase or decrease the image acquisition rate based on the number of people in the viewing area 101. For example, if the processor determines that there are no people in the viewing area 101, it may reduce the image acquisition rate to reduce power and memory consumption. Similarly, if the processor determines that the viewer is not engaged with the video (for example, because the viewer appears to be sleeping), it may reduce the image acquisition rate to save power, memory, or both. Conversely, the processor may increase the image acquisition rate (for example, to more than 15fps) if the viewer appears to be shifting attention rapidly, watching fast-paced video (for example, a soccer match or an action movie), rapidly changing channels (for example, channel surfing), or if the content is changing relatively rapidly (for example, between a series of advertisements).
[0064] If the system includes both an IR sensor and a visible image sensor, the local processor can also vary image acquisition based on lighting conditions or relative image quality. For example, in low-light conditions, the local processor may acquire IR images faster than visible images. Similarly, if the local processor obtains better results processing visible images than IR images, it may acquire visible images faster than IR images (or vice versa, if the same is true).
[0065] The system also records samples of the video's soundtrack using microphone 140 (220). Generally, the audio data acquisition rate or audio sampling rate is lower than the image acquisition rate. For example, the microphone acquires an audio sample at a rate of once every 30 seconds. With each acquisition, microphone 140 records an audio sample with a finite duration to enable identification of the video associated with the audio sample. The duration of the audio sample may be substantially equal to or longer than 5 seconds (e.g., about 5 seconds, about 6 seconds, about 8 seconds, about 10 seconds, about 20 seconds, or about 30 seconds, including any value and subranges in between).
[0066] The local processor uses audio samples recorded by microphone 140 to identify the video being played on the display (222). For example, processor 150 can create a fingerprint of the audio data and use the fingerprint to query a third-party application programming interface (API), which responds to the query with the identification information of the video associated with the audio data. In another example, processor 150 can determine the identification information of the video by comparing the fingerprint against a local table or memory.
[0067] As mentioned above, identifying video using a sample of the video soundtrack offers several advantages compared to digital watermarks used to identify video by conventional television survey devices. It eliminates the need to insert digital watermarks into the video and the need to collaborate with content creators and providers. This simplifies content production and distribution, enabling a wider range of video content identification and evaluation, including creators and distributors who cannot or do not wish to provide digital watermarks. It also eliminates the need to connect local devices via cables or set-top boxes.
[0068] Furthermore, using audio data instead of digital watermarks reduces the risk of "false positives," or instances where the system detects a person within the viewing area and identifies footage that is not actually being viewed, even when the TV is off. This can occur with traditional systems connected to set-top boxes, where household members may leave the set-top boxes powered on even when the TVs are off.
[0069] In some cases, the local processor adjusts the audio sampling rate based on, for example, video identification information, the number of people in the viewing area, the number of people involved in the video (224). For example, if the local processor cannot identify the video from a single fingerprint (for example, because the video soundtrack contains a popular song that appears in many different video soundtracks), the local processor and microphone may be improved to take samples faster or for longer periods to resolve any ambiguity in the video. The processor may also reduce the audio sampling rate to save power, memory, or both if there are no people in the viewing area 101 or the viewer is not involved in the video (for example, because the viewer appears to be sleeping). Conversely, the processor may increase the audio sampling rate if the viewer is rapidly changing channels (for example, channel surfing) or if the content is changing relatively rapidly (for example, between a series of advertisements).
[0070] Depending on the implementation, the microphone can record audio samples at regular intervals (i.e., periodically) or at irregular intervals (e.g., aperiodic or time-varying intervals). For example, the microphone can acquire audio data at a constant rate (e.g., about 2 samples per minute) throughout the day. In other cases, the microphone can operate at a certain sampling rate when the television is on or likely to be on (e.g., in the evening), and at a different, lower sampling rate when the television is off or likely to be off (e.g., early morning, midday). If the local processor detects from the audio samples that the television has been turned on (off), it can increase (decrease) the sampling rate accordingly. The local processor may also trigger the image sensor to start (stop) imaging the viewing area in response to detecting that the television has been turned on (off) from the audio samples.
[0071] While raw image and audio data is being processed or after processing, the local processor overwrites the raw image and audio data or erases it from memory (230). In other words, each image is held in memory 150, while processor 150 detects and identifies humans and measures their involvement and facial expressions. Detection, identification, and involvement data are collected frame by frame, and this information is persisted and eventually uploaded to backend server 170. Similarly, audio data is held in memory 160, while a third-party API processes audio fingerprints and returns identification information for the associated video. The identification information is stored and / or uploaded to backend server 170, as described below.
[0072] By overwriting or erasing (or otherwise discarding) raw image and audio data, the local processor reduces memory demands and diminishes or eliminates the ability to identify individuals within the viewing area. This reduces the exposure of information to potential targets attempting to infiltrate the system, thus preserving individual privacy. It also eliminates the possibility of an individual's images being transmitted to a third party. This is particularly useful for protecting children's privacy within the viewing area as defined by the Children's Online Privacy Protection Act.
[0073] In some cases, the local processor actively erases raw image and audio data from memory. In other cases, the local processor stores the raw image and audio data in one or more buffers in memory of a size that does not store more than a predetermined amount (e.g., one image or one audio sample). The local processor analyzes the raw image and audio data over the period between samples so that the next image or audio sample overwrites the buffer.
[0074] The local processor 150 also stores the processed data in memory 160. The processed data may be stored in a relatively compact format, such as comma-separated variable (CSV) format, to reduce the amount of memory required. The data contained in the CSV or other file may, for example, indicate whether there are people in each image, the number of people in the viewing area 101 of each image, the number of people actually looking at the display 11 in the viewing area 101, the emotional classification of each viewer, and the identification information of each viewer. The processed data may also include indicators of the operating status of the local device, such as the IR image acquisition rate, visible image acquisition rate, audio sampling rate, and current software / firmware updates.
[0075] The local processor sends the processed data to a remote server (e.g., via a network interface) for storage or further processing (236). Because the processed data is in a relatively compact format, upload bandwidth is significantly reduced compared to uploading raw image and audio data. Also, since the transmitted data does not contain images of the viewing area or audio samples that may contain the viewer's voice, the risk of violating viewer privacy is lower. Furthermore, because the audio and image portions of the processed data are processed locally, rather than being sent to a remote server for processing, there is a higher probability that they will be synchronized and maintain their state.
[0076] In some cases, the local processor may transmit the processed data remotely while it is being processed. In other cases, the local processor may identify a transmission window based on, for example, the available upstream bandwidth, the amount of data, etc. (234). These transmission windows may be predetermined (e.g., 2 a.m. Eastern Standard Time), set by a household member when the local device is installed, set by a remote server (e.g., via a software or firmware update), or determined by the local processor based on bandwidth measurements.
[0077] Figure 2B shows how to train a computer vision model to quantify viewer engagement. In 241, both RGB and IR sensors acquire video data that undergoes two types of processing. In 242a, the video data is manually annotated to identify faces in each frame. In 242b, faces in each frame are automatically detected using the current model (e.g., the default model or a previously used model). In 243b, the processor is used to calculate the accuracy of the automatic detection in 242b for the annotated video acquired in 242a. In 244, if the accuracy is acceptable, method 240 proceeds to 245, where the current model is set as a fabricated model for face recognition (e.g., used in method 200). If the accuracy is unacceptable, method 200 proceeds to 243a, where the video is split into a training set of video (246a) and a test set of video (246b). For example, you can select RGB video as training video 246a and IR video as test video 246b (or vice versa).
[0078] The training video 246a is sent in 247a to train the new model, while the test video (246b) is sent to stage 247b to test the new model. In 247b, both the training video 246a and the test video 246b are collected to calculate the accuracy of the new model in 247c. In 249, the processor recalculates the accuracy of the new model. If the accuracy is acceptable, the new model is set as the fabricated model (245). Otherwise, method 240 proceeds to 248, where the parameters of the new model are tuned. Alternatively, another new model may be built in 248. In any case, the parameters of the new model are sent back to 247a, and the training video 246a is used to train the new model. In this way, new models can be iteratively built to have acceptable accuracy.
[0079] Remote server operation
[0080] During operation, the remote server 170 collects data transmitted from different local devices 105 located in different homes. The remote server 170 can periodically read the input data. The remote server 170 can also analyze the received data and combine the video recognition data and the speech recognition data using the timestamps of when each was saved.
[0081] Furthermore, the remote server 170 can correct mislabeled data. For example, the remote server 170 can use data from preceding and succeeding timestamps to correct blips where viewers are not identified or are misidentified. If a person is identified in an image preceding the image in question, and also in an image following the image in question, the remote server 170 can determine that this person also appears in the image in question.
[0082] Furthermore, the remote server 170 can load data received from the local device 105 and / or data processed by the remote server 170 into a queryable database. In one example, the remote server 170 can also provide access to users who may then use the stored data for analysis. In another example, the stored data in the queryable database can also facilitate further analysis performed by the remote server 170. For example, the remote server 170 can use the database to calculate attention indices and audience indices.
[0083] Evaluation of viewer engagement
[0084] Figures 3A to 6 illustrate a method for quantifying viewer engagement with video using measurements such as the visibility index and attention index. The following definitions may help in understanding the method and apparatus of the present invention for quantifying viewer engagement with video.
[0085] Program duration is defined as the total duration of a unique program, e.g., in seconds, minutes, or hours. The actual units used (seconds, minutes, or hours) are irrelevant as long as the durations of different programs can be compared.
[0086] Commercial duration is defined as the total duration (in seconds or minutes) of a single commercial.
[0087] Viewing duration (seconds) is defined as the total duration (in seconds) that a unique program or commercial was viewed in each household. Alternatively, viewing seconds may be defined as the total duration (in seconds) of the program minus the total time (in seconds) that no household was watching the program.
[0088] Total viewing duration (seconds) is defined as the total viewing time (in seconds) of a unique program or commercial across all households.
[0089] The positive duration ratio is defined as the percentage of time a program or commercial was viewed. More specifically, the positive duration ratio for a program or advertisement can be calculated by multiplying the ratio of aggregated viewing duration to the total duration of the program or advertisement by the number of households.
[0090] Viewer count (VC) is defined as the total number of viewers within the viewing area, across all households that had a positive viewing time for a given program or commercial advertisement.
[0091] Viewership rate (WR) is defined as the ratio of the total number of people in all households with the television on to the total number of people in all households. For example, consider 100 households with a total population of 300 people. If 30 of these households with 100 people each have their television set on, the viewership rate is 33.3% (i.e., 100 / 300). However, if the same 30 households each have 150 people, the viewership rate is 50% (i.e., 150 / 300).
[0092] Viewership rating (VR) is defined as the ratio of the total number of people in the viewing area of all households to the total number of televisions that are turned on. For example, if there are 100 people in a viewing area defined by 40 different televisions (one television defines one viewing area), the VR is 2.5 (i.e., 100 / 40).
[0093] Attention Rate (AR) is defined as the ratio of the total number of people paying attention to the television in all households to the total number of people in the viewing area across all households. For example, if 100 people are in the viewing area across all individuals considered by the method, but only 60 are actually watching television (the remaining 40 may be doing other things with the TV on), the attention rate is 0.6 or 60%.
[0094] The Visibility Index (VI) is defined as the average of the viewership ratings (VR) for each program and commercial.
[0095] The attention index is defined as the average of the attention rate (AR) for each program and commercial.
[0096] Figure 3A shows Method 300 for evaluating viewer engagement (e.g., box 206 in Method 200 in Figure 2A), which includes face and eye tracking 310, face recognition 320, and sentimental analysis 330. A processor (e.g., local processor 150 shown in Figure 1) may be used to implement Method 300. Input data in Method 300 may be data acquired by a local device 105 shown in Figure 1, such as image data, audio data, or depth data of the viewing area. Face and eye tracking 310 is employed to identify and track feature data points as the face moves to determine whether the user is looking at the screen. Face recognition 320 is employed, for example, to determine viewer identification information using artificial intelligence. Sentimental analysis 330 is employed, for example, to determine viewer sentiment using artificial intelligence to analyze facial features, posture, and heart rate, among other things.
[0097] The information obtained, including whether viewers are actually watching the screen, viewer identification information, and viewer sentiment, is used to determine various video ratings 340. In one example, the obtained information is used to estimate individual video ratings for each household. In another example, the obtained information is used to estimate individual video ratings for each demographic region. In yet another example, the obtained information is used to estimate an overall video rating for a group of videos. In yet another example, the obtained information is used to estimate viewer reactions to specific videos (e.g., programs and advertisements). The obtained information can also be used to determine quantitative measures of viewer engagement, such as the visibility index and the attention index, as described below.
[0098] Steps 310, 320, and 330 in Method 300 can be achieved using pattern recognition techniques. These techniques can determine, for example, whether any viewer is in the viewing area by recognizing one or more human faces. If a recognized face is indeed present, these techniques can further determine who the viewer is by comparing the recognized face to a database containing facial data of the household in which the video is being played. Alternatively, if the viewer is not a person from that household, these techniques may use an expanded database to include facial data of more people (e.g., the entire community if possible). These techniques can also determine, for example, whether a viewer is watching the video by tracking facial movements and analyzing facial orientation.
[0099] Furthermore, artificial intelligence, machine learning, and trained neural network learning techniques can be used to analyze viewers' emotions. To this end, these techniques analyze, among other things, gestures (static posture for a set period of time), body movements (changes in posture), face orientation, face direction / movement / position, heart rate, and so on.
[0100] In another example, Method 300 can first recognize a face from image data acquired by, for example, an RGB sensor 140 and an IR sensor 120, as shown in Figure 1. Method 200 can also detect the position of the face, identify feature points on the face (e.g., the boundary points of the eyes and mouth, as shown in Figure 2A), and track them as the face moves. Method 300 can use eye-tracking techniques to determine whether the viewer is actually looking at the image (or, instead, simply sitting within the viewing area but doing something else). Method 300 can then match the viewer to a known person in the household by using trained neural network learning techniques to compare facial features from a database to similar areas. Once the viewer is identified, Method 300 can continuously track the viewer for notable facial features and determine the user's mood and / or emotions.
[0101] Method 300 also compares audio data (e.g., acquired by microphone 140 shown in Figure 1) with audio databases and other audio of video (e.g., a television show) to determine which video is being played at a particular time. In one example, video matching can determine which television channels are being watched by the viewer identified by Method 300. In another example, video matching can determine which television programs are being watched by the viewer. In yet another example, video matching can determine which commercials are being watched. Alternatively or additionally, the television channels, programs, or advertisements being watched can be determined from data collected from other sources, including but not limited to cable or satellite set-top boxes or other program provider hardware or broadcast signals.
[0102] Figure 3B shows the concepts of the visibility index and attention index, which can be estimated via the techniques described herein for quantifying viewer engagement. Generally, the visibility index quantifies the tendency of what is displayed on the screen to draw people into a room. The attention index quantifies the tendency of what is displayed on the screen to attract the viewer's attention. In other words, the visibility index can be thought of as the probability that an image (or other display content) initially attracts the viewer, and the attention index can be thought of as the probability that the image keeps the viewer in front of the display after the viewer is already in the viewing area. As shown in Figure 3B, the visibility index depends on the number of people in the viewing area, and the attention index depends on the number of people actually looking at the display.
[0103] Viewer engagement is evaluated using the visibility index and attention index.
[0104] Figure 4A shows a method 401 for quantifying viewer engagement using a visibility index. Method 401 can be implemented by a processor. Method 401 begins in step 411, where image data is acquired by a processor in each of several households participating in the method, for example, by installing or using a local device 105 in the system shown in Figure 1. The image data includes images of the viewing area in front of a display that can play video (e.g., television programs, advertisements, user-requested videos, or any other video). Furthermore, the processor also determines in step 411 whether the display is showing video. In step 421, the processor estimates the viewership and attendance rates for each video being played by the display. Viewership rate represents the ratio of the total number of people in the viewing area to the total number of displays showing video, as defined above. Similarly, attendance rate represents the ratio of the total number of people in a household with a display showing video to the total number of people in several households, as defined above.
[0105] The estimation of viewership and attendance rates is based on image data acquired in step 411 and demographic information about each household in multiple households. The demographic information may be stored in memory operably coupled to the processor so that the processor can easily retrieve the demographic information. In another example, the processor may retrieve the demographic information from another server. In step 330, the processor determines the visibility index based on the viewership and attendance rates for each unique video in multiple videos. The visibility index is defined above as the average of the viewership ratings for each video, such as programs and commercials.
[0106] Method 401 may further include a step of estimating the viewer count and positive duration ratio for each video being played on the display. The estimation is based on image data and demographic information about each household in multiple households. As defined above, the viewer count represents the total number of people involved in each unique video, and the positive duration ratio represents the ratio of the total time spent by people in multiple households watching a unique video to the duration of that unique video.
[0107] A balanced visibility index can be determined based on viewer count and location duration ratio. In one example, the balanced visibility index may be calculated as a weighted average of the visibility index (VI), taking into account the viewer count and positive duration ratio for each given program and commercial. In another example, the balanced visibility index may be calculated by normalizing the visibility index across unique images in multiple images.
[0108] Method 401 may further include a step of averaging the viewability indices across all programs and commercials over a finite period in order to generate an average viewability index. The viewability index for each program and commercial may be divided by the average viewability index (calculated, for example, on a daily, weekly, or monthly basis) in order to generate a final viewability index (a dimensionless quantity) for users such as advertising agencies, television stations, or other content providers. In one example, the finite period is approximately two weeks. In another example, the finite period is approximately one month. In yet another example, the finite period is approximately three months.
[0109] Image data can be acquired at various acquisition rates. In one example, image data may be acquired 50 times per second (50 Hz). In another example, image data may be acquired 30 times per second (30 Hz). In yet another example, image data may be acquired every second (1 Hz). In yet another example, image data may be acquired every 2 seconds (0.5 Hz). In yet another example, image data may be acquired every 5 seconds (0.2 Hz). Furthermore, Method 300 may acquire and classify image data for each viewer within a viewing area in order to derive viewer engagement information, taking into account household demographic information.
[0110] Figure 4B shows a method 402 for quantifying user engagement with video using an attention index. Method 402 includes a step 412 in which image data of the viewing area in front of the display is acquired for each household participating in the viewer engagement evaluation. In step 412, the processor determines whether any video is displayed on the display when the image data is acquired (for example, via audio data acquired by the microphone 140 in the local device 105 shown in Figure 1). In step 422, for each video being played by the display, the processor estimates an attention rate based on the image data and demographic information about the household. As defined above, the attention rate represents the ratio of the total number of people engaged with the video to the total number of people in the viewing area. Based on the attention rate of the video, in step 432, the attention index is determined to indicate the effect of the video.
[0111] Method 402 further includes a step of estimating the viewer count and positive duration ratio for the video being played on the display. Similar to Method 401, Method 402 can determine the viewer count and positive duration ratio based on image data and demographic information about each household. Using the viewer count and positive duration ratio, the processor can determine a balanced attention index. Method 402 may include a step of generating a normalized attention index by normalizing the attention index for a unique video in multiple videos over a given period (e.g., one week or one month).
[0112] Method 402 may further include a step of averaging the attention indices across all programs and commercials over a finite period in order to generate an average attention index. The attention index for each program and commercial may be divided by the average attention index to generate a final attention index (a dimensionless quantity) for clients such as advertising agencies, television stations, or other content providers.
[0113] Using facial recognition technology to evaluate viewer engagement
[0114] Figure 5 shows a method for evaluating viewer engagement with video using facial recognition technology and other artificial intelligence technologies. Method 500 begins in step 510, where an image of the viewing area in front of the display is captured (for example, using the system shown in Figure 1). For each acquired image, the number of people in the viewing area is estimated in step 520. In one example, the estimation may be performed using, for example, facial recognition technology. In another example, the estimation may be performed based on body skeleton detection.
[0115] In step 530, the orientation of each person's face relative to the display within the viewing area is determined. For example, the face may be facing the display, indicating that the viewer is actually looking at the image on the display. Alternatively, the face may be facing away from the display, indicating that the viewer is within the viewing area of the display but is not looking at the image. Therefore, based on the orientation of the viewer's face, the processor may evaluate in step 540 whether each person within the viewing area is actually engaged with the image. By distinguishing between those who are actually looking at the image and those who are not, the processor can make an accurate judgment about the effect of the image. The effect of the image can be quantified, for example, by how long the viewer's engagement with the image can be maintained.
[0116] Detects skeletal structure, face, identification information, emotions, and level of engagement.
[0117] Figure 6 is a flowchart illustrating method 600 for detecting skeletons, faces, identification information, emotions, and engagement, which can be further used for viewer engagement assessment as described above. Method 600 can be implemented by a processor (e.g., processor 150 or a processor in remote server 170). Method 600 begins in step 610, where image data of the viewing area in front of the display is provided (e.g., by memory or directly from an imaging device such as the RGB sensor 130 shown in Figure 1). In step 620, the processor acquires skeleton frames from the image data (i.e., image frames containing images of at least one potential viewer, see, for example, 230 in Figure 2A). In step 630, the processing loop is initiated, where the processor uses six individual skeleton data points / sets for each skeleton frame for further processing, including face recognition, emotion analysis, and engagement determination. Once the skeleton data has been processed, method 600 returns to acquiring skeleton frames in step 620 via a refresh step 625.
[0118] Step 635 in method 600 is a decision step, where the processor determines whether any given skeleton is detected in the selected skeleton data within the skeleton frame. If not, method 600 returns to step 630, where new skeleton data is picked up for processing. If at least one skeleton is detected, method 600 proceeds to step 640, where a bounding box is generated to identify the viewer's head region in the image data. The bounding box may be generated, for example, by identifying the head from the entire skeleton based on the skeleton information.
[0119] Step 645 is again a decision step, where the processor determines whether a bounding box has been generated (i.e., whether a head region has been detected). It is possible that the image contains the entire skeleton of the viewer, but the viewer's head is obscured and therefore not visible in the image. In this case, method 600 returns again to step 630, where the processor picks up new skeleton data. If a bounding box has been detected, method 600 proceeds to step 650, where the processor performs a second level of face recognition (also called face detection). In this step, the processor attempts to detect a human face within the bounding box generated in step 640. Face detection can be performed using, for example, a Haar feature-based cascade classifier in OpenCV. More information is described in U.S. Patent No. 8,447,139 B2, which is incorporated herein by reference in its entirety.
[0120] In stage 655, the processor determines whether a face was detected in stage 650. If not, the first level of face recognition is performed in stage 660. This first level of face recognition may be substantially the same as the second level of face recognition performed in stage 650. Performing face detection one more time can reduce the possibility of accidental failure of the face recognition technology. Stage 665 is a decision stage, similar to stage 655, where the processor determines whether a face was detected.
[0121] If a face is detected by either first-level or second-level face recognition, method 600 proceeds to step 670 to perform facial landmark detection, also known as facial feature detection or facial keypoint detection. Step 670 is employed to determine the location of different facial features (e.g., outer corners of the eyes, eyebrows, mouth, tip of the nose, etc.). More information on facial landmark detection can be found in U.S. Patent Publication 2014 / 0050358 A1 and U.S. Patent No. 7,751,599 B2, which are incorporated herein by reference in their entirety.
[0122] In step 672, the processor determines whether any facial features were detected in step 670. If not, method 600 returns to step 630 to select other skeletal data for further processing. If at least one facial feature is detected, the processor further determines in decision step 674 whether any face was detected in the second level of face recognition in step 650. If so, method 600 proceeds to step 690, where the detected face is identified (i.e., it is determined who the viewer is), and then the method proceeds to step 680, where the facial sentiment is predicted based on the facial features. In step 674, if the processor finds that no face was detected in step 650, method 600 proceeds directly to step 680 for the processor to estimate the viewer's sentiment. Sentiment analysis can be performed using, for example, a Support Vector Machine (SVM) in OpenCV. Further information can be found in U.S. Patent No. 8,488,023, which is incorporated herein by reference in its entirety.
[0123] In one example, the method shown in Figures 3A to 6 analyzes all available video (including television programs and advertisements) regardless of its duration or viewer count. In another example, the method shown in Figures 3A to 6 performs preliminary filtering to exclude videos that are too short or have too few viewer counts before performing a quantitative analysis of viewer engagement. In this way, the quantitative analysis can yield statistically more reliable results. For example, videos with a viewing time of less than a finite duration (e.g., less than 30 seconds, less than 20 seconds, or less than 10 seconds) may be excluded. Similarly, videos with fewer than a certain number of viewers over a finite period (e.g., less than 20 people, less than 15 people, or less than 10 people) may also be excluded.
[0124] In one example, the method shown in Figures 3A to 6 is performed on a live television broadcast. In another example, the method shown in Figures 3A to 6 is performed on a recorded television broadcast. If the timing of the program is recognized as being more than 10 minutes off from the original "timestamp" (e.g., the television station's database), the program is determined to be viewed as a recording. Otherwise, the program is determined to be viewed live.
[0125] Experimental evaluation of the effectiveness of commercial messages (CMs)
[0126] This section describes the collection and analysis of accurate viewing data to verify the effectiveness of commercial messaging (CM) management. The "visibility" index indicates whether a person is "in front of the television." The visibility index is created for this descriptive and research purpose to generate data. The research was conducted over two weeks, sampling 84 people from 30 households. A CM curve is defined as a pattern showing the time-series curve of visibility between two scenes. While individual viewing rates of a CM between scenes may be constant, visibility rates can change. As a result, seven patterns of CM curves were found. CM length and visibility variables can significantly contribute to the shape of the CM curve. Furthermore, a multinomial logit model may be useful in determining the CM curve.
[0127] This experiment investigated the relationship between commercial messages (CMs), programs, and human viewing behavior. The experiment also characterized the systems and methods described above. The correlation between program information, such as broadcast timing and television station, and viewing behavior using statistical methods was analyzed. Currently, in individual viewership surveys conducted in Japan, individuals are registered using colored buttons on television remote controls, and recordings are made when these buttons are pressed at the start and end of television viewing. Furthermore, a metric called the People Meter (PM) records what television viewers watched and who watched those programs (Video Research Ltd. (2014): "TV rating handbook," available in PDF format on the VIDEOR.COM website, incorporated herein by reference). However, even if viewership ratings are accurately captured, this type of viewership survey typically cannot distinguish between focused viewing and casual viewing.
[0128] Hiraki and Ito (Hiraki, A. & Ito, K. (2000): Cognitive attitudes to television commercials based on eye tracking analysis combined with scenario, Japanese Journal of Human Engineering, Vol.36, pp.239-253, incorporated herein by reference) proposed a method for analyzing the influence of commercials on image recognition using visual information based on eye movement analysis. They conducted a commercial viewing experiment using actual commercials in an environment that replicated viewing conditions. As a result, they found that auditory and visual information may hinder product comprehension.
[0129] In this experiment, in addition to individual viewership ratings, physical presence captured by the system was used as an indicator to measure viewing attitudes. For example, during commercials, people may stand up and turn their attention to each other even if they are not sitting in front of the television. Thus, during commercials, viewing attitudes were statistically analyzed using two indices: individual viewership ratings and physical presence. The latter index is referred to herein as "visibility."
[0130] From mid-November to the end of November 2014, a viewing attitude survey experiment was conducted with 84 people from 30 households. Data was collected 24 hours a day for 14 days.
[0131] Figure 7 shows a schematic diagram of a data acquisition system 700 for measuring viewer engagement within a viewing area 701 with respect to a program or advertisement displayed on a TV 702 or other display. The system 700 includes an image sensor 710 that captures images of the viewing area 701 while the TV 702 is on. The system 700 also includes a computing device 750 that stores and processes image data from the image sensor 710 and communicates the raw image data and / or processed image data to and from a server (not shown) via a communication network.
[0132] In some cases, the computing device 750 and / or server measure visibility in addition to individual viewership. Visibility indicates being "in front of the television," a term defined for viewers who are within a distance of approximately 0.5m to 4m from the television and whose faces are within a 70° angle to the left or right of the television's front. In one example, visibility is obtained in units of one second and is expressed as the number of samples per second divided by the total number of samples (84 in this case).
[0133] Figures 8A, 8B, 8C, 8D, 8E, 8F, and 8G show seven different shapes of the CM curve, illustrating the transition of the value obtained by dividing visibility by individual viewership. This value can represent the proportion of people who are actually watching television.
[0134] To explain the differences in the shape of the CM curves, data classification and modeling can be performed. The analytical methods employed in this experiment are discussed below. First, a multinomial logit model (e.g., Agresti, A. Categorical data analysis. John Wiley & Sons (2013) is incorporated herein by reference) can be employed for data modeling. Then, non-hierarchical clustering can be performed using the K-means method, given the large sample size (1,065). Next, a decision tree can be constructed. Explanatory variables are used, and all samples are classified using stepwise grouping. In general, a decision tree is a classification model that represents multiple classification rules in a tree structure. The Gini coefficient was used as the impurity function.
[0135] When using these methods to determine the shape of the commercial broadcast curve, the analysis should also consider methods or variables closely related to the determination of the commercial broadcast curve's shape. This may include variables that are observed substantially simultaneously with the commercial broadcast.
[0136] Data from the time slots with the highest viewership ratings each day will be used. In this experiment, the time slot with the highest viewership ratings is the 6 hours from 18:00 to 24:00. Viewer behavior towards commercials on five television stations will be analyzed. The ratio of the commercial curves for each television station is shown in Figure 9.
[0137] In the analysis, the shape of the commercial curve is the dependent variable and is classified into A to G, as shown in Figures 8A to 8G. The explanatory variables are the length of the commercial, the television station, the genre, the time elapsed since the start of the program, the average individual viewership rating of the commercial, the average viewability rating of the commercial, the average individual viewership rating of the previous scene, the average viewability rating of the previous scene, the viewability rating of the current scene divided by the individual viewership rating, the viewability rating of the previous scene divided by the individual viewership rating, the date, and the day of the week. The previous scene refers to the scene between commercials.
[0138] Table 1 shows the discrimination results based on the multinomial logit model. The discrimination rate with the multinomial logit model is 20% higher than the discrimination rate with random selection. The discrimination rate is particularly high when the shape of the CM curve is B or G.
[0139] This model uses seven explanatory variables: commercial length, television station, elapsed time since the start of the program, average individual viewer rating for the commercial, viewability rating, viewability rating of the commercial divided by the individual viewer rating, and viewability rating of the previous scene divided by the individual viewer rating. Of the seven variables, commercial length and television station contribute the most to the discrimination rate. [Table 1]
[0140] It is also possible to hierarchically represent the seven shapes of dependent variables. While several different types of hierarchies are possible, for efficient analysis, we compared the following two types of hierarchies.
[0141] Hierarchy 1: Monotonic shape types (C / D / E) and non-monotonic shape types (A / B / F / G). First, monotonic shape types without extrema and non-monotonic shape types with extrema were stratified. A multinomial logit model can be applied to each group to calculate the discrimination rate for each group. The discrimination results for Hierarchy 1 are shown in Table 2. The discrimination rate for monotonic shape types was 59.34%, but the discrimination rate for monotonic shape types was 51.72%, and the overall discrimination rate was 53.62%.
[0142] After hierarchically classifying monotonic and non-monotonic shape types, the overall discrimination rate was 15% higher compared to a multinomial logit model without hierarchy. Compared to a multinomial logit model without hierarchy, there were cases where the difference in discrimination rate due to the shape of the CM curve was correctly identified (D / E / G) and cases where it was not correctly identified (C).
[0143] The selected explanatory variables are as follows: For the monotonic shape type, six variables are selected: television station, elapsed time since the start of the program, average individual viewership rating for the commercial, viewability of the commercial, viewability of the previous scene, and viewability of the previous scene divided by the individual viewership rating. For the non-monotonic shape type, six variables are selected: commercial length, television station, elapsed time since the start of the program, average individual viewership rating for the commercial, viewability rating for the commercial, and viewability rating of the previous scene. Commercial length, which contributes to a non-hierarchical multinomial logit model, is not selected for the monotonic shape type. [Table 2]
[0144] Hierarchy 2: Simple shape types (A / B / C / D / E), complex shape types (F / G). Secondly, simple shape types, which have a maximum of one extremum, and complex shape types, which have more than one extremum, can be stratified. The classification results for Hierarchy 2 are shown in Table 3. The classification rate for simple shape types was 46.50%, while the classification rate for complex shape types was 77.55%, and the overall classification rate was 52.21%.
[0145] For the simple shape type, nine variables are selected: commercial length, TV station, elapsed time since program start, average individual viewership for commercials, commercial visibility, average individual viewership for the previous scene, visibility divided by the commercial's individual viewership, visibility of the previous scene divided by the average individual viewership, and date. Furthermore, for the complex shape type, only one variable, TV station, is selected. Since this model has only one variable, all samples are classified according to F. For the simple shape type, the selected variables are the same as those in a non-hierarchical multinomial logit model. [Table 3]
[0146] Cluster analysis using explanatory variables can be performed. The results of the cluster analysis are shown in Table 4. The discrimination rate was 15.77%, and there was no difference in the discrimination rate between cluster analysis and random selection. In other words, CM curves cannot be classified in non-hierarchical cluster analysis. [Table 4]
[0147] Figure 10 shows a classification model via a decision tree. The decision tree's results are shown in Table 5. The decision tree's discrimination rate is 40%. From Table 5, it can be seen that the discrimination rate for G is 0%, but the discrimination rate for D is 73%, which is higher than the other CM curves. The decision tree's discrimination rate is slightly higher than that of the multinomial logit model without hierarchy.
[0148] From Figure 10, the characteristics of each shape of the CM curve can be identified. Shape A occurs when visibility is high. Shape B occurs when visibility is low and the CM length is long. Shape C occurs when the visibility of the scene is not much different from the visibility of the previous scene. Shape D occurs when visibility is low and the CM length is short. Shape E occurs when the visibility of the previous scene is low and the CM length is short. Shape F occurs when the visibility of the scene is low but the visibility of the previous scene is high. [Table 5]
[0149] Comparison and Consideration The discrimination rates for each method are summarized in Table 6. Method 1 of stratification shows the highest rate among all methods. However, because the dependent variable is stratified, it is impossible to prove the entire connection. [Table 6]
[0150] The discriminant rate of a non-hierarchical multinomial logit model is almost the same as that of a decision tree. Decision trees are difficult to understand intuitively because their judgment is based on whether the visibility rate is higher than a fixed value, and the fixed value is not reproduced. Therefore, the most appropriate method for determining the CM curve is a non-hierarchical multinomial logit model.
[0151] In all methods, the length of commercials and viewability rates are the variables that contribute most to determining the commercial curve. Therefore, television viewing behavior does not depend on the genre of the program or the broadcast time, but does depend on the length of commercials and the viewability of the current and previous scenes.
[0152] In these five methods, the CM length variable and visibility rate significantly contribute to the interpretation of the CM curve. Therefore, we will consider two points: 1) the relationship between CM length and visibility rate, and 2) under what circumstances visibility rate is high.
[0153] Figure 11 shows the relationship between commercial length and viewability. Generally, the shorter the commercial, the higher the viewability. People stop watching television when they lose interest, so the longer the commercial, the lower the viewability.
[0154] Furthermore, the conditions that lead to high viewership rates were investigated. Viewership rates are high when a short amount of time (depending on the genre) has passed since the program started. As shown in Table 7, there are significant differences in the average viewership rates for each genre. New programs have low viewership rates, while movies and music programs have high viewership rates. Figure 12 shows the correlation between the time elapsed since the start of a program and the viewership rate. From Figure 12, it can be seen that the viewership rate is higher when the time elapsed since the start of a program is shorter. [Table 7]
[0155] This experimental study uses exemplary embodiments of the hardware and software components of the present invention to elucidate the relationship between commercials, programs, and human viewing behavior. The most appropriate method for determining the commercial curve is a multinomial logit model.
[0156] Variables observable during commercials are analyzed, and the relationship between the commercial curve and these variables is examined. In all methods employed, the variable of commercial length and visibility rate contribute most significantly to determining the commercial curve. Due to the high discrimination rate of monotonic shape types, it is easier to distinguish between changes and no changes. In other words, the shape of the commercial curve is not related to program characteristics such as genre and date. This indicates that the longer the commercial broadcast time, the more likely viewers are to tire of watching it. Also, if the preceding scene in the program is uninteresting to viewers, they will not watch the next commercial.
[0157] Application of viewer engagement data
[0158] Figure 13 shows a system for communicating data acquired using the method and system described herein. System 1300 stores and processes raw data 1310 acquired from a television viewer panel via a motion sensing device, which is then transferred to a computing device 1320, but not limited to a desktop machine. The method for evaluating viewer engagement can then be run, for example, on the desktop machine to analyze and process the data. The method converts the analyzed data into performance-based television audience rating data that can be used to determine (1) who is actually watching television (who are the viewers), (2) how often viewer members watch television, and (3) the viewers' response to television programs and advertisements. This processed and / or summarized data is then transferred on the cloud to a central storage location 1330, such as a server, where third parties, including but not limited to television advertising agencies 1340, television networks 1350, and any other potential clients 1360 that consider the data useful, can conveniently access the data at any time through the software, application programming interface, or web portal of the collecting entity, which is specifically developed for the clients of the collecting entity. Alternatively, raw data 1310 collected by the hardware component's sensors is transferred directly or indirectly via an internet connection to central storage 1330 in the cloud, where it is analyzed by software components and made available to interested third parties 1340-1360. These third parties can optionally access the raw data through the system.
[0159] Figure 14 shows the basic elements of an exemplary system 1400 that can utilize data acquired and analyzed by the systems and methods described herein. A collecting entity 1430 (e.g., TVision Insights) may compensate panel members 1410 (e.g., household members) who permit them to place the hardware component arrangement depicted in Figure 1 on their televisions in their homes for the purpose of collecting television audience data, in exchange for compensation or volunteer work. Panel members may be asked to provide additional information 1420, including but not limited to credit card transaction data, demographic and socioeconomic information, social media account logins, and data from tablets, smartphones, and other devices. This data is collected, and video and IR images are recorded by the system depicted in Figure 1, and the video may be analyzed by the methods described in Figures 2A to 6. Once analyzed, the data describing the video may then be transmitted to a collecting entity 1430 which may sell or otherwise provide the data to advertisers 1440, television stations 1460, television agencies 1450, and other interested third parties. The collecting entity 1430 may optionally provide access to the raw collected data for individual analysis. As part of the disclosed business model, the collecting entity 1430 may incentivize advertisers 1440 to encourage their television agencies 1450 to purchase this data.
[0160] Figure 15 shows a big data analysis and visualization based on data acquired in a method for evaluating viewer engagement. In these models 1500, the collection entity 1520 (e.g., TVision INSIGHTS) shown in Figure 15 can collect data from households 1510 that own televisions. In return, participating households 1510 can receive monetary compensation (or other benefits) from the collection entity 1520. The collection entity 1520 then analyzes the data collected from participating households using big data analysis 1530a and visualization techniques 1530b to derive information such as the effectiveness of a particular television program or advertisement. This data can then be provided to advertisers, advertising agencies, television stations or other content providers or promoters (collectively referred to as customers 1540) to instruct them to improve the effectiveness of their programs. In one example, customers 1540 can subscribe to this data service from the collection entity 1520 for a monthly fee. In another example, customer 1540 may purchase data from collecting entity 1520 regarding specific video content (e.g., campaign videos, special advertisements at sporting events, etc.).
[0161] Figure 16 shows an example of a collection of additional information 1600 from individuals and households (television viewers) participating in viewer engagement data collection. Television viewers may represent national and / or regional demographics useful to interested third parties. The collecting entity collects video data 1610 and demographic information, which is packaged with data collected by the system, analyzed in a manner consistent with television viewership ratings, and this information can be provided to customers for compensation. Examples of information that may be collected from television viewers include, among other things, any or all information that can be obtained through social media profiles 1620, such as but not limited to Twitter®, Instagram, and Facebook®. The information may further include multi-screen data 1630, including video and audio data 1640 (including both television audio and audio such as conversations from individuals within the household), smartphone and tablet search habits, internet search history, email account information, and credit card transaction data 1650. This list is not exhaustive and should not be interpreted restrictively.
[0162] The collected information and data enable advertisers to accurately assess the impact of television advertising, including unprecedented measurement of individual demographics. Advertisers can use the data to determine which ad slots are best suited to their target audiences. Furthermore, by delivering messages tailored to different audience types, they can effectively drive purchasing behavior and improve their return on investment (ROI).
[0163] Furthermore, television networks can benefit from the disclosed invention by obtaining more accurate data on the evaluation of their television programs, audience types, reactions, and advertising slot predictions. This enables them to determine the most valuable advertising slots for specific target demographic segments, as well as to improve programs to suit different audience types and eliminate unpopular programs. The data can also be used to compare programs across multiple channels at the same or different time slots for comparative evaluation of programs and advertisements. Similarly, television viewer data and behavior can be collected and compared with streaming content at any given program time slot. Television pilot programs can also be evaluated using this system before ordering episodes. In another embodiment with reference to Figures 17 and 18, another aspect of the invention includes the ability to identify a particular program or advertisement that was viewed, and the platform or service on which it was viewed. From this aspect, the streaming service playing the content (e.g., Netflix®, Hulu®, Paramount+, etc.) can be identified. In this aspect, the platform on which the service is running (e.g., Amazon Firestick, Samsung Smart TV, Apple TV, etc.) and the start, end, pause, or resume times of the streaming session are also identified. This is partially achieved by a software module 1702 running on the measuring device 105. Module 1702 collects and observes network packets outbanded from the streaming service while minimizing the impact on the quality of the video stream. The model is trained on the majority of the collected real data. This data consists of Ethernet® packets outbanded from the streaming application, with respondents logging their actions and various states of the streaming application. Respondent actions include: a. Turn on your streaming device. b. Start the streaming application on the device. c. Select some content within the application. d. Press the play button e. Press the pause button while the content is playing. f. Resume playback g. Navigate back to the home screen. The application status recorded by the respondents is as follows: h. Home screen on the display i. The application logo is displayed. j. The application has been started. k. Content Introduction l. Content is playing m. An ad is playing. Module 1702 further applies the above model to network packet data collected from panelists' homes to predict active streaming applications. This model uses packets captured while respondents were viewing content on streaming applications, along with logs of their recorded actions and application states. The timing of state transitions between playback, pause, and resume is also available. Module 1702 further applies this model to network packet data collected from panelists' homes to predict which streaming applications were active at any given time, which devices the streaming applications were running on, and which state the streaming applications were in at that time (stopped, playing, paused). This analysis yields a time series, sorted in ascending order over time, for each panelist's home that viewed any given streaming content. This time series data is then combined with content detected to be playing simultaneously on the panelists' televisions, as determined by Module 1702, along with viewer identification information and the level of attention paid to the content on the televisions, to obtain a second-by-second explanation defining which demographic groups were viewing which streaming applications and which streaming content on which streaming-enabled devices. Figure 17 shows the data collection on local device 105. Local device 105 first discovers that various streaming-enabled devices 1706 are active in the home. After this: 1. In one embodiment, the packet inspection module 1702 uses ARP poisoning to disguise itself as an internet gateway in the home. ARP poisoning is one example of a method to facilitate this, but many other methods are possible. At this point, the local device 105 is seen as an internet gateway for its location, but only to streaming devices that need to collect data from it. 2. As a result of ARP poisoning, any packets from the streaming device that were assumed to be sent to the gateway are instead sent to the local device 105. 3. The packet inspection module 1702 analyzes the content contained in these packets and records out-of-band information. These packets are usually encrypted, but the available information includes: a. IP header b. TCP header c. Lookup request d.TLS handshake packet The above outbound information is available on local device 105 and is recorded along with the time the packets were observed. 4. After recording the available information in the network packet, the local device 105 forwards that information to the "real" gateway 1708. 5. Gateway 1708 forwards the packet to its destination (not shown) via the Internet 1710. 6. The response is then received by gateway 1708 (as shown in "6"). 7. The response packet is not routed through the gateway to ensure that stream quality is not buffered due to delay. It proceeds directly to the streaming device 1706 (as shown in "7"). In most cases, this packet contains the streaming content that is to be displayed on the display device 1704. Packet acquisition and forwarding The first stage in the packet acquisition process is the discovery of the streaming-enabled device in the panelist's home. This process is performed when module 1702 listens for mDNS packets broadcast by the streaming-enabled device 1706. This discovery process also obtains the IP address of the streaming device. In one embodiment, once the IP address of the streaming device is obtained, the MAC address of that streaming device is obtained through a lookup in the ARP table. In one embodiment, the packet acquisition or redirection process is initiated using ARP spoofing. This leverages a feature of the Ethernet® protocol, which requires each host on the network to know the MAC address of other hosts in order to communicate with them. The only way these hosts can discover the MAC address of another host is to ask the other host for its MAC address and trust that what the other host replies is correct. Using the packet redirection process as described herein, the packet inspector 1706 convinces the target streaming that the MAC address of local device 105 is that of the internet gateway 1708. Here, the target streaming device 1706 is confident that the MAC address of the local device 105 is actually the MAC address of the internet gateway 1708, and therefore sends all packets destined for the internet gateway 1708 to the local device 105 instead. In this way, module 1702, running on the local device 105, can inspect all packets leaving the streaming device 1706. Once module 1702 has inspected the packets, it forwards them to the home internet gateway 1708. In particular, the latency of outband packets from the streaming device 1706 can increase by introducing extra hops to outgoing packets via the local device 105, which is constantly performing highly computationally intensive tasks such as running computer vision algorithms on multiple video frames per second, inspecting packet headers, and in some cases, also inspecting content.This delay can cause packet retransmissions, ultimately resulting in the streaming device 1706 not having enough data to continue playing the stream. Referring to Figure 18, kernel 1804 is illustrated. Kernel 1804 is part of device software 105. To prevent such the aforementioned delay scenario, the device 105 software, including the kernel software shown in 1804, utilizes (in one embodiment) Extended Data Path (XDP) technology. This allows the device 105 software, including module 1702, to parse incoming packets and correctly collect the necessary data within kernel 1804 before the packets traverse the TCP stack. Instead of traversing the TCP stack, packets are forwarded directly from kernel 1804 to the internet gateway. As a result, these packets do not need to be processed by code in user space 1802. By using XDP as a means of inspecting packets, the overhead of traversing the TCP stack or passing data into the user address space is avoided. This packet inspection technique is highly effective, allowing the software on device 105 to simultaneously monitor packets from multiple streaming devices 1706 without negatively impacting stream quality. A key improvement for high-speed observation of packet data is that aggregated data is held within a kernel data structure, allowing a user-space program to poll the kernel at a given frequency, typically every second, to collect the latest values for those data points. The user-space program 1802 can use the previous values collected for each data point, along with the new values read from the kernel, to determine the delta of the value for that data point since the last time the kernel was polled. This method avoids the need for the user-space program to observe data from every incoming packet. An example of aggregated data is a count of packets that are outbound from a streaming device to a specific IP and port. Referring again to Figure 18, the XDP hook on kernel 1804 reads the incoming packet data in step (1).By examining the packet header, the XDP hook determines whether the packet is for gateway 1708 or whether local device 105 is indeed the intended recipient. If local device 105 is the intended recipient, the XDP hook places the data on the normal TCP stack and ultimately delivers the packet to the intended user process. However, if the packet is instead destined for gateway 1708, the XDP program first updates the data in its kernel data structure, then changes the out-of-band MAC address to the MAC address of internet gateway 1708, and places the packet on the TX (transmit) queue. Data collection Since most packet content is encrypted, module 1702 focuses on the following data points and extracts them from each packet. If the packet is a TCP packet but not a TLS handshake packet, the information extracted is Destination IP Sourceport Destination port Includes, If the packet is a TCP packet and is a "client hello" TLS handshake packet, the software will: Destination IP Sourceport Destination port Server name from server extension Extract, If the packet is a UDP packet and the destination is a standard DNS port, the software will assume the packet contains a DNS name query. Server name being searched Let's assume we extract the following. The data described above is collected by local device 105 with participation by module 1702. This data is then uploaded by local device 105 to remote data processing server 170. In one embodiment, processing the uploaded data includes the step of mapping each IP to its own organization by performing a reverse DNS lookup. For example, the reverse DNS lookup may indicate that the IP belongs to a streaming service such as hulu.com or a CDN service such as Akamai. Once the IP is replaced with the name of the organization that owns the IP, The total number of seconds for all outbound packets to various services DNS lookup duration in seconds Server name with established TLS connection The data, including the above, is fed into a pre-trained predictive model. This model then determines the state of the streaming device, streaming service, and streaming app (as described above) every second that content is streamed using device 1706 in a particular home.
[0164] While various inventive embodiments are described and illustrated herein, those skilled in the art will readily conceive of various other means and / or structures for performing the functions described herein and / or obtaining one or more of the results and / or advantages, and each of such variations and / or modifications will be considered to fall within the scope of the inventive embodiments described herein. More generally, those skilled in the art will readily understand that all parameters, dimensions, materials, and configurations described herein are illustrative, and that actual parameters, dimensions, materials, and / or configurations will depend on the specific use or application in which the teachings of the present invention are used. Those skilled in the art will be able to recognize or elucidate many equivalents to several specific embodiments of the invention described herein using only a little ordinary experimentation. Therefore, it should be understood that the embodiments described herein are presented only as examples, and that inventive embodiments may be carried out separately from those specifically described and claimed within the scope of the appended claims and their equivalents. The multiple embodiments of the inventions of this disclosure cover the individual features, systems, articles, materials, kits, and / or methods described herein. Furthermore, any combination of two or more such features, systems, articles, materials, kits, and / or methods is included within the scope of the invention of this disclosure, provided that they are not mutually inconsistent.
[0165] The embodiments described above can be implemented in any of a number of ways. For example, the design and manufacturing embodiments of the technology disclosed herein can be implemented using hardware, software, or a combination thereof. If implemented in software, the software code can run on any suitable processor or set of processors, whether provided on a single computer or distributed across multiple computers.
[0166] Furthermore, it should be understood that computers can be embodied in any of many forms, such as rack-mount computers, desktop computers, laptop computers, or tablet computers. Additionally, computers may be incorporated into devices that are not generally considered computers but possess sufficient processing power, such as personal digital assistants (PDAs®), smartphones, or other suitable portable or fixed electronic devices.
[0167] Furthermore, a computer may have one or more input and output devices. These devices may, among other things, be used in the user interface of the present invention. Examples of output devices that can be used to provide a user interface include a printer or display screen for a visual representation of the output, and a speaker or other sound-generating device for an auditory representation of the output. Examples of input devices that can be used in the user interface include a keyboard and pointing devices such as a mouse, touchpad, or digitized tablet. In another example, a computer may receive input information by speech recognition or other audible forms.
[0168] Such computers may be interconnected by one or more networks of any suitable form, including local area networks or wide area networks such as enterprise networks, intelligent networks (IN), or the Internet. Such networks may be based on any suitable technology, operate according to any suitable protocol, and may include wireless networks, wired networks, or fiber optic networks.
[0169] The various methods or processes outlined herein can be coded as software executable on one or more processors employing any one of a variety of operating systems or platforms. Furthermore, such software may be written using one of a number of suitable programming languages and / or programming or scripting tools, and may be compiled as executable machine code or intermediate code that runs on a framework or virtual machine.
[0170] In this regard, various inventive concepts can also be embodied as computer-readable storage media (or a plurality of computer-readable storage media) (e.g., computer memory, one or more floppy disks, compact disks, optical disks, magnetic tapes, flash memory, circuit configurations in field-programmable gate arrays or other semiconductor devices, or other non-temporary or tangible computer storage media) encoding one or more programs that, when executed on one or more computers or other processors, perform methods for carrying out the various embodiments of the present invention described above. The computer-readable media or media can be movable so that the programs stored thereon or the programs can be loaded onto one or more different computers or other processors to carry out the various embodiments of the present invention described above.
[0171] The terms “program” or “software” are used herein in a general sense to refer to any type of computer code or set of computer-executable instructions that can be employed to program a computer or other processor to implement various aspects of the embodiments described above. Furthermore, it should be understood that, according to one aspect, one or more computer programs that perform the method of the present invention when executed do not need to reside on a single computer or processor, but may be modularly distributed among a number of different computers or processors to implement various aspects of the present invention.
[0172] Computer-executable instructions can take many forms, including program modules executed by one or more computers or other devices. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform a specific task or implement a specific abstract data type. Typically, the functions of program modules can be combined or distributed as desired in various embodiments.
[0173] Furthermore, data structures can be stored in computer-readable media in any suitable form. For the sake of simplicity in illustration, data structures may be shown as having fields that are related by their location within the data structure. Such relationships can also be achieved by assigning locations on the computer-readable media that indicate the relationships between fields to the storage for the fields. However, any suitable mechanism can be used to establish relationships between the information in the fields of a data structure, such as the use of pointers, tags, or other mechanisms for establishing relationships between data elements.
[0174] Furthermore, various inventive concepts can be embodied in one or more methods, and examples of such methods are provided. Multiple processes performed as part of a method can be ordered in any suitable manner. Thus, multiple embodiments may include multiple processes being performed in a different order than illustrated, and may also include performing several processes simultaneously, even if they are shown as multiple subsequent processes in multiple exemplary embodiments.
[0175] All definitions defined and used herein should be understood to govern dictionary definitions, definitions in documents incorporated by reference, and / or the ordinary meanings of the defined terms.
[0176] As used herein, the indefinite articles "a" and "an" as used herein and in the claims should be understood to mean "at least one" unless explicitly indicated otherwise.
[0177] As used herein, the phrase “and / or” as used herein and in the claims should be understood to mean “either or both” of the elements thus combined, that is, elements that exist in some cases incidentally and elements that exist in other cases incidentally. The elements enumerated by “and / or” should be interpreted in the same way, that is, “one or more” of the elements being combined. In addition to the elements specifically identified by the “and / or” expression, there may be other elements that are optionally present, whether related to or not related to the elements specifically identified by the “and / or” expression. Therefore, as a non-restrictive example, when the reference "A and / or B" is used with open-ended (non-restrictive) language such as "comprising," in one embodiment it may refer to A only (optionally including multiple elements other than B), in another embodiment it may refer to B only (optionally including multiple elements other than A), and in yet another embodiment it may refer to both A and B (optionally including other elements), and so on.
[0178] As used herein, and as used in the claims, “or” should be understood to have the same meaning as “and / or” as defined above. For example, when multiple articles are separated into a list, “or” or “and / or” should be interpreted as including, that is, including at least one, but also including, of the numerous or enumerated elements and, optionally, additional articles not enumerated. Only terms that explicitly indicate the opposite, such as “only one of” or “exactly one of,” or, as used in the claims, “consisting of,” refer to including exactly one element from the number or list of elements. In general, as used herein, the term “or” should be interpreted as indicating exclusive substitution (i.e., “one or the other but not both”) only when preceded by terms of exclusivity such as “either,” “one of,” “only one of,” or “exactly one of.” As used in the claims, “consisting essentially of” should have the usual meaning as used in the field of patent law.
[0179] As used herein and in the claims, the phrase “at least one” referring to a list of one or more elements should be understood to mean at least one element selected from any one or more elements in the list of elements, but not necessarily including at least one of each and all of the elements specifically listed in the list of elements, and not excluding any combination of elements in the list of elements. Furthermore, this definition allows for the optional presence of multiple elements, whether related to or unrelated to the specifically identified elements, in addition to the multiple elements specifically identified in the list of elements to which the expression “at least one” refers. Therefore, as a non-restrictive example, "at least one of A and B" (or, as in "at least one of A or B" or "at least one of A and / or B") may refer to, in one embodiment, at least one, optionally more than one, of A, but no B (and optionally more than one other element); in another embodiment, at least one, optionally more than one, of B, but no A (and optionally more than one other element); and in yet another embodiment, at least one, optionally more than one, of A, more than one, of B (and optionally more than one other element), and so on.
[0180] In the claims, as in the specification, all transitional expressions such as “comprising,” “including,” “carrying,” “having,” “containing,” “involving,” “holding,” and “composed of” are understood to be open-ended, meaning they include but are not limited to. However, as stated in the United States Patent Office Manual of Patent Examining Procedures, Section 2111.03, only the transitional phrases “consisting of” and “consisting essentially of” are closed or semi-closed transitional phrases, respectively. [Other possible items] [Item 1] A method for quantifying the degree of viewer engagement with images displayed on a screen, wherein the method is: The step of acquiring an image of the viewing area in front of the display with at least one camera while the image is displayed on the display in the respondent's home, wherein the respondent's home is the location of the measuring device in which one or more respondents have chosen to interact. The steps include: acquiring audio data representing the soundtrack of the video, emitted by a speaker connected to the display, using a microphone; A step of determining the identification information of the video based at least partially on the audio data using a processor operably coupled to the at least one camera and the microphone, The processor determines the identity of the in-home streaming service that is playing the streamed content. A method for providing this. [Item 2] The method according to item 1, further comprising the step of the processor determining which platform the streaming service is running on. [Item 3] The method according to item 1, further comprising the step of the processor determining when a streaming session should start, end, pause, and resume. [Item 4] The method according to item 1, wherein the step of the processor determining the identification of an in-home streaming service that is playing streaming content includes data packet redirection performed by a packet inspection module of the processor. [Item 5] Packet redirection is the method described in item 4, which includes a packet inspection module that disguises itself as an internet gateway within a home network. [Item 6] The packet inspection module intercepts packets and analyzes the contents within them, as described in item 5. [Item 7] The method according to item 5, wherein analyzing the contents within the packet includes recording outband packet information and recording the time the packet was intercepted. [Item 8] The method according to item 7, wherein after recording the outband packet information, the packet is forwarded to its original destination. [Item 9] The method of item 8, further comprising the step of receiving a response packet at the Internet gateway, wherein the response packet is not routed through the Internet gateway but rather routed directly to an appropriate streaming device. [Item 10] A system for quantifying the degree of viewer engagement with images displayed on a screen, wherein the system is A local device residing in the respondent's home, wherein the local device has a processor that executes multiple processes, including a packet inspection module, While the video is displayed on the display in the respondent's home, an image of the viewing area in front of the display is acquired by at least one camera, wherein the respondent's home is the location of a measuring device in which one or more respondents have chosen to interact, and the acquisition is performed. The process involves acquiring audio data representing the soundtrack of the video, emitted by a speaker connected to the display, using a microphone. The identification information of the video, based at least partially on the audio data, is determined by a processor operably coupled to at least one camera and the microphone. The processor determines the identification of the in-home streaming service that is playing the streamed content. A system with a local device that executes instructions that include at least the following. [Item 11] The system described in item 10, wherein the instructions to be executed further include the processor determining which platform the streaming service is running on. [Item 12] The system according to item 10, wherein the instructions to be executed further include the processor determining when a streaming session should start, end, pause, and resume. [Item 13] The system according to item 10, wherein the instruction to be executed further includes a step in which the processor determines the identification of an in-home streaming service that is playing streaming content, and the determination step includes data packet redirection performed by the processor's packet inspection module. [Item 14] Packet redirection is a system described in item 13, which includes a packet inspection module that disguises itself as an internet gateway within a home network. [Item 15] The packet inspection module is the system described in item 14, which intercepts packets and analyzes the contents within them. [Item 16] The system described in item 15, which includes analyzing the contents of the packet, recording outband packet information, and recording the time the packet was intercepted. [Item 17] The system described in item 16, which, after recording the outband packet information, forwards the packet to its original destination. [Item 18] The system according to item 17, further comprising the step of receiving a response packet at the Internet gateway, wherein the response packet is not routed through the Internet gateway but rather routed directly to an appropriate streaming device. [Item 19] A non-temporary computer medium in which instructions are stored, wherein the instructions are executed by a processor and perform a method, The step of acquiring an image of the viewing area in front of the display with at least one camera while the image is displayed on the display in the respondent's home, wherein the respondent's home is the location of the measuring device in which one or more respondents have chosen to interact. The steps include: acquiring audio data representing the soundtrack of the video, emitted by a speaker connected to the display, using a microphone; A step of determining the identification information of the video based at least partially on the audio data using a processor operably coupled to the at least one camera and the microphone, The process includes the processor determining the identification of the in-home streaming service that is playing the streamed content, which includes data packet redirection performed by the processor's packet inspection module. A non-temporary computer medium having [certain characteristics]. [Item 20] The medium described in item 19, further comprising the step of the processor determining which platform on which the streaming service is running. [Item 21] The medium described in item 19, further comprising a step of determining the time for a streaming session to start, end, pause, and resume using the processor. [Item 22] The step in the processor determining the identification of an in-home streaming service that is playing streaming content includes data packet redirection performed by the processor's packet inspection module, as described in item 21. [Item 23] Packet redirection is a medium described in item 22, including a packet inspection module that disguises itself as an internet gateway within a home network. [Item 24] The packet inspection module is a medium described in item 23 that intercepts packets and analyzes the contents within them. [Item 25] The medium described in item 23, which includes analyzing the contents within the packet, recording outbound packet information, and recording the time the packet was intercepted. [Item 26] The medium described in item 25, which, after recording the outband packet information, forwards the packet to its original destination. [Item 27] The medium according to item 26, further comprising the step of receiving a response packet at the internet gateway, wherein the response packet is not routed through the internet gateway but rather routed directly to a suitable streaming device. [Item 1] A method for quantifying the degree of viewer engagement with images displayed on a screen, wherein the method is: The steps include: acquiring an image of the viewing area in front of the display with at least one camera while the image is displayed on the display in the respondent's home, wherein the respondent's home is the location of a measuring device in which one or more respondents have chosen to interact; The steps include: acquiring audio data representing the soundtrack of the video, emitted by a speaker connected to the display, using a microphone; A step of determining the identification information of the video based at least partially on the audio data using a processor operably coupled to the at least one camera and the microphone, A step in which the processor determines the identification of a streaming service in the respondent's home that is playing streamed content, including data packet redirection performed by a packet inspection module of the processor, wherein the packet inspection module disguises itself as an internet gateway in the respondent's home, and the data packet redirection includes capturing packets and analyzing the content within the packets, wherein the analysis includes recording outband packet information, including the time it takes for the streaming session to transition between states. A method for providing this. [Item 2] The method according to item 1, further comprising the step of determining the streaming application provided by the streaming service using the processor. [Item 3] The method according to item 1, further comprising the step of the processor determining when a streaming session should start, end, pause, and resume. [Item 4] The method according to item 1, wherein after recording the outband packet information, the packet is forwarded to its original destination. [Item 5] The method according to any one of items 1 to 4, further comprising the step of receiving a response packet at the Internet gateway, wherein the response packet is not routed through the Internet gateway but rather routed directly to an appropriate streaming device. [Item 6] A system for quantifying the degree of viewer engagement with images displayed on a screen, wherein the system is A local device residing in the respondent's home, wherein the local device has a processor that executes multiple processes, including a packet inspection module, While the video is displayed on the display in the respondent's home, an image of the viewing area in front of the display is acquired by at least one camera, wherein the respondent's home is the location of a measuring device in which one or more respondents have chosen to interact, and the acquisition is performed. The process involves acquiring audio data representing the soundtrack of the video, emitted by a speaker connected to the display, using a microphone. The identification information of the video, based at least partially on the audio data, is determined by a processor operably coupled to at least one camera and the microphone. The processor determines the identification of an in-home streaming service playing streamed content, including data packet redirection performed by the packet inspection module of the processor, wherein the packet inspection module disguises itself as the in-home internet gateway, and the data packet redirection includes capturing packets and analyzing the content within the packets, wherein the analysis includes recording out-of-band packet information, including the time it takes for the streaming session to transition between states. A system with a local device that executes instructions that include at least the following. [Item 7] The system according to item 6, wherein the instructions to be executed further include the processor determining the streaming application provided by the streaming service. [Item 8] The system described in item 6, wherein the instructions to be executed further include the processor determining when a streaming session should start, end, pause, and resume. [Item 9] The system described in item 6, which, after recording the outband packet information, forwards the packet to its original destination. [Item 10] The system according to any one of items 6 to 9, further comprising a gateway that receives response packets, wherein the response packets are not routed through the internet gateway but rather routed directly to a suitable streaming device. [Item 11] A computer program comprising instructions, wherein the instructions, when executed by a processor, execute a method, The step of acquiring an image of the viewing area in front of the display with at least one camera while an image is displayed on the display in the respondent's home, wherein the respondent's home is the location of the measuring device in which one or more respondents have chosen to interact. The steps include: acquiring audio data representing the soundtrack of the video, emitted by a speaker connected to the display, using a microphone; A step of determining the identification information of the video based at least partially on the audio data using a processor operably coupled to the at least one camera and the microphone, A step in which the processor determines the identification of a streaming service in the respondent's home that is playing streamed content, including data packet redirection performed by a packet inspection module of the processor, wherein the packet inspection module disguises itself as an internet gateway in the respondent's home, and the data packet redirection includes capturing packets and analyzing the content within the packets, wherein the analysis includes recording outband packet information, including the time it takes for the streaming session to transition between states. A computer program that has [a certain characteristic]. [Item 12] The computer program according to item 11, further comprising the step of determining the streaming application provided by the streaming service using the processor. [Item 13] The computer program described in item 11, further comprising a step in which the processor determines when a streaming session should start, end, pause, and resume. [Item 14] A computer program as described in item 11, which, after recording the outband packet information, forwards the packet to its original destination. [Item 15] A computer program according to any one of items 11 to 14, further comprising the step of receiving a response packet at the Internet gateway, wherein the response packet is not routed through the Internet gateway but rather routed directly to a suitable streaming device.
Claims
1. A method for quantifying the viewer engagement with respect to video displayed on a display, the method comprising: obtaining, by at least one camera, image data of a viewing area in front of the display while the video is being displayed on the display; estimating, by at least one processor, viewer data including the number of people present in the viewing area and the number of people engaged with the video in the viewing area while the video is being displayed on the display, based at least in part on the image data; obtaining, by a microphone, audio data representing the sound track of the video emitted by a speaker coupled to the display; determining, by the at least one processor, identification information of the video based at least in part on the audio data; quantifying, based at least in part on the viewer data, the viewer engagement with respect to the video; quantifying, for each of a plurality of households, the viewer engagement with respect to the video based at least in part on the number of people present in the viewing area and the number of people engaged with the video in the viewing area while the video is being displayed on the display, wherein the step of quantifying the viewer engagement includes estimating a viewing rate of the video, the viewing rate representing a ratio of the number of people present in the viewing area while the video is being displayed on the display to the number of people engaged with the video in the viewing area, and determining a prominence index based on the viewing rate of the video among the plurality of videos for each unique video among the plurality of videos, and estimating a viewer count and a positive duration ratio based on the image data and demographic information for each of the plurality of households, the viewer count representing the number of people engaged with each unique video, and the positive duration ratio representing a ratio of the total time spent by people within the plurality of households watching the unique video to the duration of the unique video, A method comprising the above steps.
2. The step of quantifying the viewer engagement includes performing, by the at least one processor, one or more of face tracking, eye tracking, face recognition, and sentiment analysis. The method according to claim 1.
3. The method according to claim 2, wherein the step of performing sentiment analysis includes face feature detection and sentiment analysis.
4. The method according to claim 2, wherein the step of performing sentiment analysis includes obtaining skeletal frame data.
5. The method according to claim 1, further comprising the step of estimating demographic information about each person in the viewing area from the image data.
6. The method according to claim 1, wherein the step of estimating the demographic information includes estimating age, gender, ethnic group, and facial expression.
7. The method according to claim 1, wherein the step of determining that the video is a unique video is at least partially based on the audio data of the viewing area by signal fingerprinting technology.
8. The method according to claim 1, wherein the at least one processor further comprises generating a plurality of commercial message curves including a pattern showing a time series curve of the visibility rate between two scenes of one video.
9. The method according to claim 8, wherein the individual viewing rate of the commercial message between scenes may be constant, but the visibility rate may vary.
10. The method according to claim 9, wherein the length of the commercial message and the variable of the visibility rate can greatly contribute to the shape of the commercial message curve.
11. The method according to claim 10, wherein a multinomial logit model is used for determining the commercial message curve.
12. A method for quantifying viewer engagement with a video displayed on a display, the method comprising: Obtaining, by at least one camera, image data of a viewing area in front of the display while the video is being displayed on the display; Estimating, by at least one processor, viewer data including the number of people present in the viewing area and the number of people engaged with the video in the viewing area while the video is being displayed on the display, based at least in part on the image data; Obtaining, by a microphone, audio data representing the sound track of the video emitted by a speaker coupled to the display; Determining, by the at least one processor, identification information of the video based at least in part on the audio data. Quantifying the viewer engagement with respect to the video based at least in part on the viewer data Quantifying the viewer engagement with respect to the video based at least in part on the number of people present in the viewing area while the video is being displayed on the display in each of a plurality of households, and the number of people engaging with the video in the viewing area, wherein the step of quantifying the viewer engagement is a step of estimating the attention rate for the video, the attention rate representing a ratio between the number of people present in the viewing area while the video is being displayed on the display and the number of people engaging with the video in the viewing area, and for each unique video among a plurality of videos, determining an attention index based on the attention rate of the video among the plurality of videos Estimating a viewer count and a positive duration ratio based on the image data and demographic information for each of the plurality of households, the viewer count representing the number of people engaging with each unique video, and the positive duration ratio representing a ratio of the total time spent by people in the plurality of households watching the unique video to the duration of the unique video Determining the identification information of each person present in the viewing area based at least in part on the image data wherein the step of quantifying the viewer engagement with respect to the video includes quantifying the viewer engagement for each identified person A method comprising [
13. ] A method for quantifying viewer engagement with respect to a video being displayed on a display, the method comprising Acquiring, by at least one camera, image data of a viewing area in front of the display while the video is being displayed on the display Estimating, by at least one processor, viewer data including the number of people present in the viewing area while the video is being displayed on the display and the number of people engaging with the video in the viewing area based at least in part on the image data Acquiring, by a microphone, audio data representing the soundtrack of the video emitted by a speaker coupled to the display Determining, by the at least one processor, identification information of the video based at least in part on the audio data quantifying the viewer engagement of the video based at least in part on the viewer data; determining that the video is a unique video among a plurality of videos; estimating (i) a viewing rate and (ii) a viewing penetration rate based on the image data and demographic information about each of a plurality of households, in response to determining that the video is a unique video, wherein the viewing rate represents a ratio of the total number of people in the viewing area to the total number of displays on which the video is displayed, and the viewing penetration rate represents a ratio of the total number of people in households having a display on which the video is displayed to the total number of people in the plurality of households; determining a visibility index based on the viewing rate and the viewing penetration rate; estimating (iii) a viewer count and (iv) a positive duration ratio based on the image data and the demographic information about each of the plurality of households, wherein the viewer count represents the total number of people engaged with the unique video, and the positive duration ratio represents a ratio of the total time spent by people in the plurality of households watching the unique video to the duration of the unique video; and weighting the visibility index based on the viewer count and the positive duration ratio A method comprising the steps of: The method of claim 13, further comprising normalizing the visibility index across unique videos among the plurality of videos. A system for quantifying viewer engagement with a video displayed on a display, the system comprising: at least one camera arranged to image a viewing area in front of the display and acquire image data of the viewing area; a microphone arranged proximate to the display and acquiring audio data representative of an audio track of the video emitted by a speaker coupled to the display; a memory operably coupled to the at least one camera and the microphone and storing instructions executable by a processor, the memory including a buffer for storing the image data and the audio data; and at least one processor operably coupled to the at least one camera, the microphone, and the memory, wherein when the instructions executable by the processor are executed, the at least one processor With the at least one camera, while the video is being displayed on the display, obtaining image data of a viewing area in front of the display; With the at least one processor, estimating viewer data including the number of people present in the viewing area and the number of people involved in the video in the viewing area while the video is being displayed on the display, based at least in part on the image data; With the microphone, obtaining audio data representing the sound track of the video emitted by a speaker coupled to the display; With the at least one processor, determining identification information of the video based at least in part on the audio data; Quantifying the viewer engagement level with the video based at least in part on the viewer data; and Estimating a viewer count and a positive duration ratio based on the image data and demographic information for each of a plurality of households, the viewer count representing the number of people involved in each unique video, and the positive duration ratio representing the ratio of the total time spent by people within the plurality of households watching the unique video to the duration of the unique video; executing a method comprising; A system comprising.
16. The step of quantifying the viewer engagement level is performed by the at least one processor, Face tracking, Eye tracking, Face recognition, and Sentiment analysis, including performing one or more of; The system according to claim 15.
17. The system according to claim 16, wherein the step of performing sentiment analysis includes facial feature detection and sentiment analysis.
18. The system according to claim 16, wherein the step of performing sentiment analysis includes obtaining skeletal frame data.
19. The system according to claim 16, wherein the method further comprises, for each of a plurality of households, quantifying the viewer engagement level with the video based at least in part on the number of people present in the viewing area and the number of people involved in the video in the viewing area while the video is being displayed on the display.
20. The step of quantifying the viewer engagement level is estimating the attention rate for the video, the attention rate representing a ratio between the number of people present in the viewing area while the video is being displayed on the display and the number of people involved with the video in the viewing area, and, the system according to claim 15, comprising determining an attention index based on the attention rate of the video among the plurality of videos for each unique video among the plurality of videos.
21. The method further comprises determining that the video is a unique video among the plurality of videos, in response to determining that the video is a unique video, estimating (i) a viewing rate and (ii) a watching rate based on the image data and demographic information regarding each of a plurality of households, the viewing rate representing a ratio of the total number of people in the viewing area to the total number of displays on which the video is being displayed, and the watching rate representing a ratio of the total number of people in households having a display on which the video is being displayed to the total number of people in the plurality of households, and determining a visibility index based on the viewing rate and the watching rate, the system according to claim 15, comprising.
22. The method further comprises the system according to claim 15, wherein the at least one processor generates a plurality of commercial message curves including a pattern showing a time series curve of the watching rate between two scenes of one video.
23. A system for quantifying viewer engagement with a video being displayed on a display, the system comprising at least one camera arranged to image a viewing area in front of the display and acquire image data of the viewing area, a microphone arranged adjacent to the display and acquiring audio data representing a sound track of the video emitted by a speaker coupled to the display, a memory operably coupled to the at least one camera and the microphone and storing instructions executable by a processor, the memory including a buffer for storing the image data and the audio data, and at least one processor operably coupled to the at least one camera, the microphone, and the memory, wherein when the instructions executable by the processor are executed, the at least one processor With the at least one camera, while the video is being displayed on the display, obtaining image data of the viewing area in front of the display; With the at least one processor, estimating viewer data including the number of people present in the viewing area and the number of people involved in the video in the viewing area while the video is being displayed on the display, based at least in part on the image data; With the microphone, obtaining audio data representing the sound track of the video emitted by a speaker coupled to the display; With the at least one processor, determining identification information of the video based at least in part on the audio data; Quantifying the viewer engagement level of the video based at least in part on the viewer data; Estimating a viewer count and a positive duration ratio based on the image data and demographic information for each of a plurality of households, the viewer count representing the number of people involved in each unique video, and the positive duration ratio representing the ratio of the total time spent by people within the plurality of households watching the unique video to the duration of the unique video, and Determining identification information for each person present in the viewing area based at least in part on the image data; wherein the step of quantifying the viewer engagement level of the video includes quantifying the viewer engagement level for each identified person; Executing a method including; A system comprising. **Claim 24**: A system for quantifying viewer engagement with a video being displayed on a display, the system comprising: At least one camera arranged to image a viewing area in front of the display and obtain image data of the viewing area; A microphone arranged proximate to the display and obtaining audio data representing the sound track of the video emitted by a speaker coupled to the display; A memory operably coupled to the at least one camera and the microphone and storing instructions executable by a processor, the memory including a buffer for storing the image data and the audio data, and At least one camera, the microphone, and at least one processor operably coupled to the memory, wherein when instructions executable by the processor are executed, the at least one processor, acquiring, with the at least one camera, image data of a viewing area in front of the display while the video is being displayed on the display; estimating, with the at least one processor, viewer data including the number of people present in the viewing area and the number of people involved in the video in the viewing area while the video is being displayed on the display, based at least in part on the image data; acquiring, with the microphone, audio data representing a sound track of the video emitted by a speaker coupled to the display; determining, with the at least one processor, identification information of the video based at least in part on the audio data; quantifying, based at least in part on the viewer data, a viewer engagement level of the video; estimating a viewer count and a positive duration ratio based on the image data and demographic information for each of a plurality of households, the viewer count representing the number of people involved in each unique video, the positive duration ratio representing the ratio of the total time spent by people within the plurality of households watching the unique video to the duration of the unique video, and determining, based at least in part on the image data, identification information for each person present in the viewing area; wherein the step of quantifying the viewer engagement level of the video includes quantifying the viewer engagement level for each identified person, and the step of quantifying the viewer engagement level is performed by a remote server; executing a method including; A system comprising. **Claim 25**: A system for quantifying a viewer engagement level for a video being displayed on a display, the system comprising: at least one camera arranged to image a viewing area in front of the display and acquire image data of the viewing area; a microphone arranged proximate to the display and acquiring audio data representing a sound track of the video emitted by a speaker coupled to the display; A memory that is operably coupled to the at least one camera and the microphone and stores instructions executable by a processor, the memory including a buffer for storing the image data and the audio data, and At least one processor operably coupled to the at least one camera, the microphone, and the memory, wherein when the instructions executable by the processor are executed, the at least one processor Obtaining, by the at least one camera, image data of a viewing area in front of the display while the video is being displayed on the display Estimating, by the at least one processor, viewer data including the number of people present in the viewing area and the number of people involved in the video in the viewing area while the video is being displayed on the display, based at least in part on the image data Obtaining, by the microphone, audio data representing the sound track of the video emitted by a speaker coupled to the display Determining, by the at least one processor, identification information of the video based at least in part on the audio data Quantifying, based at least in part on the viewer data, the viewer engagement with respect to the video Estimating a viewer count and a positive duration ratio based on the image data and demographic information for each of a plurality of households, the viewer count representing the number of people involved in each unique video, and the positive duration ratio representing the ratio of the total time spent by people within the plurality of households watching the unique video to the duration of the unique video Determining, based at least in part on the image data, identification information for each person present in the viewing area Wherein the step of quantifying the viewer engagement with respect to the video includes quantifying the viewer engagement for each identified person, and the step of quantifying the viewer engagement is executed by a remote server, and Determining whether a predetermined video among a plurality of videos is being displayed on the display based at least in part on the audio data, wherein the step of quantifying the viewer engagement level is based at least in part on whether the predetermined video is being displayed. Executing a method including A system comprising