A system and method for collecting data for evaluating the validity of content
The system addresses the inaccuracy of active feedback by using AI-driven data collection and analysis to track passive emotional states and behavioral data, optimizing advertising campaigns for enhanced consumer engagement and campaign performance.
Patent Information
- Application Number
- JP2022508887
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-08-13
- Filing Date
- 2020-08-12
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2040-08-12
AI Technical Summary
Current methods for evaluating the performance of media content, such as advertisements, rely on active feedback that captures conscious thoughts rather than passive emotional states, leading to inaccurate predictions of consumer engagement and attention, and existing metrics fail to account for distractions and attention diversion.
A system and method for collecting and analyzing behavioral and physiological data from client devices using AI-driven algorithms to synchronize attention measurement criteria with context attributes, enabling real-time monitoring and optimization of advertising campaigns.
Provides accurate and scalable tracking of user attention and emotional states, allowing for optimized delivery of content that maximizes engagement and achieves campaign goals by predicting the effects of adjustments on performance.
Smart Images

Figure 0007717049000001 
Figure 0007717049000002 
Figure 0007717049000003
Abstract
Description
Technical Field
[0001] The present invention relates to a technique for collecting various data in real time from, for example, different sources or software within a network environment. Using the data, the performance of display content that may also be available via the network is evaluated. In particular, the present invention relates to a technique for collecting large amounts of data that enables measurement of a user's attention to display content. In this specification, the display content may be any information that a user can consume using a network-compatible device. For example, the content may be any of media content (e.g., video, music, images), advertising content, and web page information.
Background Art
[0002] Certain types of media content, such as advertisements, music videos, and movies, are intended to induce a change in the emotional state of consumers, for example, to attract a user's attention or otherwise enhance the user's attention. In the case of advertisements, it may be desirable to convert this change in emotional state into performance such as an increase in sales. For example, a television commercial may aim to increase the sales of a related product. There is a need for a tool that can evaluate the performance of media content before its release.
[0003] Active feedback, also called self-reported feedback, may be used in an attempt to judge or predict the performance of media content such as video commercials. In the case of active user feedback, the user provides feedback verbally or in writing after consuming the media content. For example, the user may fill out a questionnaire manually or provide voice feedback that can be recorded for analysis, such as by using an automated method with a speech recognition tool. The feedback may include things that indicate the emotional state experienced during the consumption of the media content. However, the active feedback of the user is obtained from a rationalized conscious thought process, rather than the actually experienced (passive) emotional state. It is known that the user's preferences are outside of conscious awareness and are strongly influenced by the passive emotional state. Therefore, the performance of media content cannot be accurately predicted using active feedback of the emotional state.
[0004] It is known that emotional state data can also be measured passively, for example, by collecting data indicating the user's behavioral characteristics or physiological characteristics, such as while the user is consuming media. In one example, facial reactions can be used as a passive indicator of the experienced emotional state. Video acquisition by a web camera can be used to monitor facial reactions by capturing image frames when the media content is consumed by the user. Therefore, by processing the video images, it is possible to capture the emotional state using a web camera.
[0005] Physiological parameters can also be good indicators of the experienced emotional state. Many physiological parameters cannot be consciously controlled, i.e., the consumer does not influence the physiological parameters. Thus, physiological parameters can be used to judge the true emotional state of a user consuming media content, which, in principle, can be used for an accurate prediction of the performance of media content. Examples of measurable physiological parameters include voice analysis, heart rate, heart rate variability, skin electrical activity (which may indicate arousal), respiration, body temperature, electrocardiogram (ECG) signals, and electroencephalogram (EEG) signals.
[0006] It is becoming increasingly common for users to own wearable or portable devices that can record physiological parameters of the above types. This expands the possibility of extending such physiological measurements to a large sample size, removing statistical variations (noise), and understanding the correlation with media content performance.
[0007] The emotional state information measured in this way has been found to correlate with media content performance, especially sales growth. The rapid increase in webcams on client devices means that the capture of this type of data can be extended to a large sample size.
[0008] The behavioral characteristics of a user can manifest in various forms. As used herein, "behavioral data" or "behavioral information" may refer to the visual aspects of a user's response. For example, behavioral information may include facial reactions, head and body gestures or postures, and eye tracking. In practice, it may be desirable to use a combination of raw data inputs, including behavioral data, physiological data, and self-reported data, to obtain emotional state information. A combination of raw data from two or three of the above sources may help identify "false" indicators. For example, when the emotional state data derived from all three sources overlap or coincide, the reliability of the acquired signal increases. If there is a discrepancy in the signals, it may indicate a false reading.
[0009] When behavioral characteristics are recorded for a user who is reacting to something other than the currently displayed media content, incorrect displays may occur. For example, while media content is being displayed, a user may be distracted by another person. In such a situation, the user's behavioral characteristics may be mainly affected by the conversation with the other person, and thus do not accurately reflect the user's response to the media content. Therefore, the user's attention or engagement with the media content is an important factor in determining the relevance of the collected user behavioral characteristics.
[0010] The rapid increase in web-enabled consumer devices means that it is becoming increasingly difficult for marketers to attract consumers' attention. For consumers to be influenced by an advertising message, it is essential for them to pay attention. The ease of distracting consumers means that it is increasingly desirable to accurately track the attention of viewers. Current measurement criteria, which may include the number of impressions, views, click-throughs, etc., do not provide this information. In particular, current measurement criteria do not provide information to help understand the reasons for the viewer's diverted attention. Summary of the Invention Means for Solving the Problems
[0011] Most generally, the present invention proposes a system and method for quickly and scalably tracking attention. This system includes means for collecting relevant data streams from multiple client devices while a consumer (user) is viewing content, means for analyzing the collected data using an AI-driven module that outputs one or more attention measurement criteria indicating actual attention, and means for synchronizing the collected data with the attention measurement criteria.
[0012] The system can be configured to aggregate data and generate meaningful reports regarding the effectiveness of content. In particular, the ability to synchronize attention measurement criteria with other data streams enables access to the reasons for promoting attention. Using this information, it becomes possible to generate recommendations that can direct the delivery of content to locations that optimize its effectiveness. The data may be aggregated for multiple consumers (e.g., a group of users with common demographic attributes or interests), or for multiple content items (e.g., videos with a common theme or different video ads from the same advertiser), or for a specific marketing campaign (e.g., data from various ads linked to a common advertising campaign), or for a brand (e.g., data from all content that mentions the brand or is otherwise linked to the brand).
[0013] The systems and methods of the present invention may be used to facilitate the optimization of advertising campaigns. The collected data enables the effective real-time monitoring of the share of attention of a given advertising campaign, or of the brand being displayed within several campaigns actually. The systems and methods of the present invention may provide the ability to report on the reasons for promoting attention, which may then help to determine the steps to be taken to optimize the advertising delivery strategy to achieve the campaign goals. The campaign goals may be set against measurable parameters in the system. For example, an advertising campaign may have the goal of maximizing the total attention time for a given budget. In another example, the campaign goal may be to maximize a particular type of attention, or attention in the context of a particular positive emotion, for example from a particular demographic group or within a particular geographical area. In another example, the campaign goal may be to reach a certain level of attention at the lowest cost. As will be explained in more detail below, the system can not only report on the performance against the campaign goals using the data, but can also predict how a particular additional action will affect that performance. Therefore, the system provides a tool for optimizing advertising campaigns by providing recommended actions supported by the predicted effects on the performance against the campaign goals.
[0014] Additionally or alternatively, the systems and methods of the present invention can report on the emotional states associated with attention. This may provide feedback on whether an advertisement or brand is being perceived positively or negatively.
[0015] According to the present invention, there is provided a computer-implemented method for collecting data to determine the attention paid to the display of content. The method includes displaying content on a client device, transmitting context attribute data indicating the interaction between the user and the client device during content display from the client device to an analysis server via a network, collecting the user's behavior data during content display on the client device, and applying the behavior data to a classification algorithm to generate the user's attention data, where the classification algorithm is a machine learning algorithm trained to map the behavior data to attention parameters, and the attention data indicates the variation of the attention parameters over time during content display, applying, and generating a validity dataset linking the temporal progression of the attention parameters to the corresponding context attribute data obtained during content display in the analysis server by synchronizing the attention data with the context attribute data, and storing the validity dataset in a data store.
[0016] In one example, the content to be displayed may include media content. Accordingly, the method may further include playing the media content using a media player application running on the client device, and the context attribute data further indicates the interaction between the user and the media player application during media content playback. The media player application may include an adapter module configured to transmit control analysis data of the media player application to the analysis server via a network, and the method includes executing the adapter module upon receiving the media content to be displayed.
[0017] The content to be displayed may be generated locally on the client device (e.g., by the software running thereon). For example, the content to be displayed may be related to a game or application running locally. Additionally or alternatively, the content to be displayed may be obtained from the web, e.g., by downloading, streaming, etc. Thus, the step of displaying the content may include accessing, by the client device via the network, a web page on a web domain hosted by a content server, and receiving, by the client device via the network, the content displayed by the web page.
[0018] Thus, this method can operate to collect two or more of the following types of data from the client device: (i) context attribute data from a web page, (ii) context attribute data from a media player application (if used), and (iii) behavioral data. Note that data is extracted from the collected data and all data is synchronized so that the cause or driving factor of the attention can be investigated.
[0019] In addition to the data collected from the client device, the analysis server may obtain additional information about the user from other sources. The additional information may include data indicating demographic attributes, user preferences, user interests, etc. The additional data may be incorporated into the validity dataset, e.g., as labels that allow the attention data to be filtered or sorted by demographic attributes, user preferences or interests, etc.
[0020] Additional data may be obtained in various ways. For example, an analytics server may communicate (either directly or via a network) with an advertising system such as a demand-side platform (DSP) for running programmatic ads. Additional information may be obtained from user profiles held by the DSP, or may be obtained directly from the user as feedback from a quiz, or through social network interactions. Additional information may also be obtained by analyzing images captured by a webcam on a client device.
[0021] Media content may be video, such as a video ad. Synchronization of the attention data and the context attribute data may be related to the timeline when the video is played in a media player application. The action data and the context attribute data may be timestamped to enable establishing temporal relationships between various data.
[0022] Display of media content on a web page may be triggered by accessing the web page or by performing some predetermined action on the web page. The media content may be hosted on a web domain, for example, directly embedded in the content of the web page. Alternatively, the media content may be obtained from another entity. For example, a content server may be a publisher that provides space on a web page to an advertiser. The media content may be an advertisement sent from an ad server (e.g., as a result of an ad bidding process) to fill the space on the web page.
[0023] Therefore, the media content may be outside the control of the content server. Similarly, the media player application on which the media content is played may not be software resident on the client device. Therefore, it may be necessary to obtain the context attribute data related to the web page independently of the context attribute data related to the media player application.
[0024] The classification algorithm may be arranged in the analysis server. Having a central location can facilitate the update process of the algorithm. However, it is also possible for the classification algorithm to be on the client device, in which case, instead of sending the behavior data to the analysis server, the client device is configured to send the attention and emotion data. The advantage of having the classification algorithm on the local device is that the user's privacy is improved because it is not necessary to send the user's behavior data from the user's computer. Executing the classification algorithm locally also means that much less processing power is required on the analysis server, and costs can be saved.
[0025] Accessing a web page may include obtaining a context data start script for execution on the client device. The context data start script may be machine-readable code and may be placed, for example, in a tag within the header of the web page.
[0026] Alternatively, the context data start script may be provided within the communication framework through which content is supplied to the client device. For example, if the content is a video advertisement, the communication framework typically includes an advertisement request from the client device and a video advertisement response sent from the advertisement server to the client device. The context data start script may be included in the video advertisement response. The video advertisement response may be formatted according to the Video Ad Serving Template (VAST) standard (e.g., VAST 3.0 or VAST 4.0), or may conform to any other advertisement response standard such as, for example, the Video Player Ad Interface Definition (VPAID), the Mobile Rich Media Ad Interface Definition (MRAID).
[0027] In a further alternative, the context data start script may be injected into the web page source code by an intermediary between the publisher (i.e., the sender of the web page) and the user (i.e., the client device). The intermediary may be a proxy server or, in some cases, a code injection component within the network router associated with the client. In these examples, the publisher does not need to incorporate the context data start script into its version of the web page. This means that it is not necessary to send the context data start script in response to every hit on the web page. Further, this technique may make it possible to include the script only in requests from client devices associated with users who have permitted the collection of behavioral data. In some examples, such users may form a panel for evaluating the effectiveness of web content before it is released to a wider audience.
[0028] This method may further include executing a context data start script on a client device to perform one or more preliminary operations before the content is displayed. The preliminary operations may include determining consent to send context attribute data and action data to an analysis server, determining the availability of the device for collecting action data, and checking whether the user has been selected for action data collection. This method may include ending the action data collection procedure when the client device using the context data start script determines that (i) consent to the transmission of action data is withheld, or (ii) the device for collecting action data is unavailable, or (iii) the user has not been selected for action data collection. If any one of these criteria is determined, the action data collection procedure may be ended. In this case, the client device may send only the context attribute data to the analysis server. As described below, the context attribute data can be used to predict attention data.
[0029] The collection of action data may include capturing an image of the user using a camera, such as a webcam or similar device. The captured image may be an individual image or a video. The image preferably captures the user's face and upper body, i.e., such that changes in posture, head position, etc. are observable. The context data start script may be configured to activate the camera.
[0030] Image or video data may be transmitted from a client device, e.g., by streaming or other transmission, using any suitable real-time communication protocol, e.g., WebRTC. This method may include loading code to enable a real-time communication protocol by a client device using a context data start script when it is determined that (i) consent has been given for the transmission of behavioral data, and (ii) a device for collecting behavioral data is available, and (iii) a user for behavioral data collection has been selected. To prevent the initial access to the web page from being slow, the code to enable the real-time communication protocol may not be loaded until all of the above conditions are determined.
[0031] Context attribute data may include web analytics data of a web page and control analysis data of a media player application. The analysis data may include information conventionally collected and communicated for a web page and a media player application, such as the visibility of any element, clickstream data, mouse movement (e.g., scroll, cursor position), keystrokes, etc.
[0032] Execution of the context data start script may be configured to trigger or initialize the collection of web analytics data. Analysis data from the media player application may be obtained using an adapter module that may be a plugin forming part of the media player application software, or a separate loadable software adapter that communicates with the media player application software. The adapter module may be configured to transmit control analysis data of the media player application to an analysis server via a network, and this method includes executing the adapter module when receiving the media content to be displayed. The adapter module may be activated or loaded as a plugin through the execution of the context data start script.
[0033] The context data start script may be executed as part of the execution of a web page, or as part of the execution of a mobile application for viewing content, or as part of the execution of a media player application. The control analysis data and the web analysis data may be sent from the entity where the context data start script is executed to an analysis server.
[0034] When the behavior data includes a plurality of images showing the user's reactions over time, the classification algorithm may operate to evaluate the attention parameter of each image in the plurality of images of the user captured during content display.
[0035] In addition to the attention data, the behavior data may be used to obtain the user's emotional state information. Thus, this method is to apply the behavior data to an emotional state classification algorithm to generate the user's emotional state data, where the emotional state classification algorithm is a machine learning algorithm trained to map the behavior data to the emotional state data, and the emotional state data further includes generating the user's emotional state data indicating the temporal variation of the probability that the user has a given emotional state during content display, and synchronizing the emotional state data with the attention data, and the effectiveness data set further includes the emotional state data.
[0036] The client device may be configured to respond locally to the detected emotional state and / or attention parameter data. For example, the content may be obtained and displayed by an application running on the client device, and the application is configured to determine an action based on the emotional state data and attention parameter data generated on the client device.
[0037] The functions described herein may be implemented as a software development kit (SDK) for use in creating an app or other program that can utilize the attention parameters or effectiveness data described above. The software development kit may be configured to provide a classification algorithm.
[0038] The methods described herein are scalable to a networked computing environment that includes multiple client devices, multiple content servers, and multiple different portions or types of content. Thus, the method may include receiving, by an analytics server, context attribute data and behavioral data from a plurality of client devices. The analytics server may operate to aggregate a plurality of effectiveness data sets obtained from the context attribute data and behavioral data received from the plurality of client devices, for example, according to the processes described above. The plurality of effectiveness data sets may be aggregated with respect to one or more common dimensions shared by the context attribute data and behavioral data received from the plurality of client devices, for example, with respect to a given media content, or a group of related media content (e.g., related to an advertising campaign), or by a web domain, by a website identity, by time, by type of content, or by any other suitable parameter.
[0039] The result of executing the above method is a data store with a rich and valid dataset that links the user's attention to other observable factors. The valid dataset may be stored in a data structure such as a database, from which queries can be made to create reports that can observe the relationships between attention data and other data. Accordingly, this method may further include receiving, by a reporting device via a network, a query for information from the validity dataset, extracting, by the reporting device, response data in response to the query from the data store, and transmitting, by the reporting device, the response data via the network. The query may be from a brand owner or publisher.
[0040] Aggregated data may be used to update the functionality of the client device. For example, if the content is obtained and displayed by an app running on the client device, this method may further include determining a software update for the app using the aggregated validity dataset, receiving the software update on the client device, and adjusting the functionality of the app by executing the software update.
[0041] As described above, this method may enable the acquisition of attention data even when behavioral data is not available. This can be done by using context attribute data to predict attention data. Accordingly, this method is to apply context attribute data to a prediction algorithm to generate predicted attention data for the user when it is determined that there is no available behavioral data from the client device, where the predicted attention data indicates fluctuations in the attention parameter over time during content display, and to synchronize the predicted attention data with the context attribute data to generate a predicted validity dataset that links the progression of the attention parameter over time to the corresponding context attribute data obtained during content display.
[0042] The prediction algorithm itself may be a machine learning algorithm trained to map context attribute data to attention parameters. Alternatively or additionally, the prediction algorithm may be rule-based.
[0043] In another aspect, the present invention can provide a system for collecting data for determining the attention paid to web-based content. This system includes a content server and an analysis server, and a plurality of client devices communicable via a network. Each client device accesses a web page of a web domain hosted by the content server, receives the content displayed by the web page, sends context attribute data indicating the interaction between the user of the client device during content display and the web page to the analysis server, and is configured to collect the user's behavioral data during content display. The system is further configured to apply the received behavioral data to a classification algorithm to generate the user's attention data. The classification algorithm is a machine learning algorithm trained to map behavioral data to attention parameters. The attention data indicates the variation of the attention parameters over time during content display. The analysis server synchronizes the attention data with the context attribute data to generate a validity dataset that links the temporal progression of the attention parameters to the corresponding context attribute data obtained during content display, and is configured to store the validity dataset in a data store. The features of the foregoing method can be similarly applied to the system.
[0044] As described above, the validity data generated by the system can be used to predict how a particular additional action affects the performance of a given piece of content or a given advertising campaign. In another aspect of the present invention, a method for optimizing an advertising campaign is provided. In this method, a programmatic advertising strategy is adjusted using the recommended actions supported by the predicted effects on performance for the campaign goals.
[0045] According to this aspect, a computer-implemented method for optimizing a digital advertising campaign may be provided. The method includes accessing an effectiveness dataset representing the temporal evolution of an attention parameter while playing advertising content belonging to a digital advertising campaign to a plurality of users, where the attention parameter is obtained by applying behavioral data collected from each user during the playback of the advertising content to a machine learning algorithm trained to map the behavioral data to the attention parameter; generating a candidate adjustment to a target audience strategy associated with the digital advertising campaign; predicting the effect of applying the attention parameter to the candidate adjustment; evaluating the predicted effect on the campaign goal of the digital advertising campaign; and updating the target audience strategy with the candidate adjustment if the predicted effect exceeds a threshold to improve the attention parameter. The update may be performed automatically, i.e., without human intervention. Thus, the target audience strategy may be automatically optimized.
[0046] The effectiveness dataset may be obtained using the method described above and thus may have any of the features described herein. For example, the effectiveness dataset may further include user profile information indicating the demographic attributes and interests of the user. In such an example, the candidate adjustment to the target audience strategy may change the demographic attributes or interest information of the target audience.
[0047] In practice, the method may generate and evaluate multiple candidate adjustments. The method may automatically implement all adjustments that lead to an improvement exceeding the threshold. Alternatively or additionally, the method may include presenting (e.g., displaying) all or a subset of the adjustments that lead to an improvement exceeding the threshold. The method may include selecting one or more of the adjustments used to update the target audience strategy, e.g., manually or automatically.
[0048] The step of automatically updating the target audience strategy may include communicating the revised target audience strategy to a demand-side platform (DSP). Thus, the method according to this aspect may be performed in a network environment and may include, for example, a DSP, the above-described analysis server, and a campaign management server. The DSP may operate in a conventional manner based on instructions from the campaign management server. The analysis server may be able to access the effectiveness dataset and may be an entity that performs optimization of campaign goals based on information from the campaign management server. Alternatively, the campaign goal optimization may be performed by the campaign management server, and the campaign management server may be configured to send a query to the analysis server, for example, to obtain and / or evaluate the predicted effect of candidate adjustments to the target audience strategy.
[0049] Embodiments of the present invention will be described in detail below with reference to the accompanying drawings.
Brief Description of the Drawings
[0050]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Mode for Carrying Out the Invention
[0051] Embodiments of the present invention relate to a system and method for collecting and utilizing behavioral data from a user while the user is consuming web-based content. In the following examples, the presented content is media content, such as video or audio. However, it should be understood that the present invention is applicable to any type of content that a website can present.
[0052] FIG. 1 is a schematic diagram of a data collection and analysis system 100 according to an embodiment of the present invention. In the following description, the system is described in the context of evaluating media content 104 in the form of a video advertisement, which may be created, for example, by a brand owner 102. However, it can be understood that the systems and methods of the present invention are applicable to any type of media content for which it is desirable to monitor a consumer's attention. For example, the media content may be a training video, a safety video, online learning materials, a movie, a music video, and the like.
[0053] System 100 is provided in a networked computing environment in which several processing entities are communicatively connected via one or more networks. In this example, system 100 includes one or more client devices 106 configured to play media content via, for example, a speaker or headphones and a software-based video player 107 on a display 108. Client device 106 may also or may be connected to behavior data capture devices such as a webcam 110, a microphone, etc. Exemplary client devices 106 include smartphones, tablet computers, laptop computers, desktop computers, and the like.
[0054] The client device 106 is communicatively connected via the network 112 such that it can receive content 115 provided for consumption from, for example, a content server 114 (e.g., a web host) that can operate under the control of the publisher to deliver content on one or more channels or platforms. The publisher may sell, on their channels, "space" for brand owners to display video advertisements via an advertising bidding process or by embedding the advertisements in the content.
[0055] Accordingly, the provided content 115 may include, for example, as a result of an advertising bidding process, media content 104 that is provided directly by the content server 114 or transmitted together with or separately from content provided by the advertising server 116. The brand owner 102 may supply the media content 104 to the content server 114 and / or the advertising server 116 in any conventional manner. The network 112 may be any type of network.
[0056] In this example, the provided content includes code for triggering the transmission of context attribute data 124 from the client device 106 to the analysis server 130 via the network 112. The code is preferably in the form of a tag 120 within the header of the main page loaded from the domain hosted by the content server 114. The tag 120 operates to load a bootstrap script that performs several functions to enable the delivery of information including the context attribute data 124 from the client device 106. These functions are described in more detail below. However, in the case of the present invention, the main function of the tag 120 is to trigger the delivery of the context attribute data 124 and, optionally, an action data stream 122 such as a webcam recording including video or image data from the camera 110 on the client device 106 to the analysis server 130.
[0057] The context attribute data 124 is preferably analysis data related to events that occur after the main page is loaded. The analysis data may include information conventionally collected and transmitted on the main page, such as the visibility, clicks, scrolls, etc. of any element. This analysis data may provide a control baseline for the influence of "increasing" emotions or attention when the associated media is displaying or playing the content 104.
[0058] As described above, the "behavior data" or "behavior information" in this specification may refer to the visual aspects of the user's responses. For example, the behavior information may include facial reactions, head and body gestures or postures, and eye tracking. In this example, the behavior data stream 122 transmitted to the analysis server 130 may include the user's facial reactions in the form of a set of videos or images of the user captured while consuming the media content 104.
[0059] In addition to the behavior data 122 and the context attribute data 124, the analysis server 130 is configured to receive a supplementary context attribute data stream 126 that includes the media content 104 itself and analysis data from the video player on which the media content is displayed. The media content 104 may be supplied to the analysis server 130 directly from the brand owner 102, or from the content server 114 or the client device 106. The supplementary context attribute data stream 126 may be obtained by loading an adapter for the video player 107 on which the media content 104 is displayed. Alternatively, the video player 107 may have a plugin to provide the same functionality in the native environment of the video player 107.
[0060] The supplemental context attribute data stream 126 is obtained for the purpose of synchronizing the action data 122 with the playback position within the media content, and thus provides brand measurement and creative level analysis. The supplemental context attribute data stream 126 may include visibility, playback events, clicks, and scroll data related to the video player.
[0061] Particularly when the rendering of the media content 104 is performed via a third-party advertising server 116, the video player 107 can be placed within an i-frame, so a separate mechanism for generating the supplemental context attribute data stream 126 is provided. In such a case, the adapter needs to be placed within the i-frame, where the adapter can cooperate with the function of the main tag 120 to record data and send it to the analysis server 130.
[0062] For example, the supplemental context attribute data stream 126 may include information related to user commands such as pause / start, stop, volume control, etc. Additionally or alternatively, the supplemental context attribute data stream 126 may include other information related to delays or interruptions in playback, such as due to buffering.
[0063] The context attribute data stream 124 and the supplemental context attribute data stream 126 are combined to provide the analysis server 130 with a rich background context that may be (and can actually be) related to the user's response to the media content obtainable from the action data stream 122.
[0064] The action data stream 122 is not necessarily obtained for all users who view the media content 104. This may be because the user has not consented to share information or does not have a suitable camera for recording action data. Even if permission to share information is given but action data has not been obtained, the main tag 120 may send the context attribute information 124, 126 to the analysis server 130. The attention information may be predicted from this information in the manner described below.
[0065] The bootstrap script may operate to determine whether the action data stream 122 should be obtained from a given client. This may include checking, for example, based on a random sampling technique and / or based on the restrictions of the publisher (e.g., because only feedback from a specific class of audience is required), whether the user has been selected to participate.
[0066] The bootstrap script may first operate to determine or obtain permission to share the context attribute data 124 and the supplementary context attribute data 126 with the analysis server 130. For example, if there is a consent management platform (CMP) in the domain in question, the script operates to check for consent from the CMP. It may also operate to check for a global opt-out cookie associated with the analysis server or a specific domain.
[0067] Next, the bootstrap script may operate to check whether the behavioral data stream 122 should be acquired. If it should be acquired (e.g., because the user was selected as part of the sample), the bootstrap script may check the permission API of the camera 110 for recording and transmitting the camera feed. Since the behavioral data stream 122 is transmitted together with the context attribute data from the primary domain page, it is important that the tag for executing the bootstrap script is in the header of the primary domain page rather than in an associated i-frame.
[0068] In one example, the behavioral data stream 122 is a complete video recording from the camera 110 that is transmitted to the analysis server 130 via a suitable real-time communication protocol such as WebRTC. To optimize the page load speed, the WebRTC recording and the code for tracking on the device are not loaded by the bootstrap script until the relevant permissions are confirmed. In another approach, the camera feed may be locally processed by the client device so that only the detected attention, emotions, and other signals are transmitted and no images or videos leave the client device. In this approach, some of the functions of the analysis server 130 described below are distributed to the client device 110.
[0069] Generally, the function of the analysis server 130 is to convert the basically free-form display data obtained from the client device 106 into a rich data set that can be used to determine the validity of the media content. As a first step, the analysis server 130 operates to determine the attention data of each user. The attention data can be obtained from the behavioral data stream 122 by using the attention classifier 132, which is an AI-based model that returns the probability that the face on a given webcam frame shows attention to the content on the screen.
[0070] Accordingly, the attention classifier 132 can output a time-varying signal indicating the progress of the user's attention while consuming the media content 104. This can be synchronized with the media content 104 itself to enable the detected states of attentiveness and distraction to match those experienced by the user when consuming the media content. For example, when the media content is a video advertisement, a brand may appear at a specific point or period within the video. The present invention enables these points or periods to be marked or labeled with attention information.
[0071] Similarly, the video's creative content can be represented as a stream of keywords associated with various points or periods within the video. Synchronizing the keyword stream with the attention signal can enable the recognition of the correlation between the keywords and attentiveness or distraction.
[0072] The attention signal may also be synchronized with the context attribute signal in a similar manner, thereby providing a rich dataset of context data synchronized with the progress of the user's attention. These datasets, which can be obtained from each user consuming the media content, are aggregated and stored in the data store 136, from which they can be queried and further analyzed to generate reports, identify correlations, and make recommendations. This will be explained below.
[0073] The context attribute data 124 may also be used, for example, to give confidence or trust that the output from the attention classifier 132 is applicable to the content to which it relates, by enabling a cross-check of what is visible on the screen.
[0074] The action data stream 122 may also be input to the emotional state classifier 135, which operates to generate a time-varying signal indicative of the user's emotions when consuming media content. Thus, this emotional state signal may also be synchronized with the attention signal, thereby enabling the evaluation and reporting of emotions related to attention (or distraction).
[0075] Even if the data received at the analysis server 130 from the client device 106 does not include the action data stream 122, the attention signal can be obtained by using the attention predictor 134. The attention predictor 134 is configured to generate or infer attention from the context attribute data 124 and the supplementary context attribute data 126. The attention predictor 134 may be a rule-based model that generates predictions based on a statistical modeling of the context attribute data for which attention data is known. For example, if the context attribute data indicates that a frame showing a video advertisement is not visible on the screen (e.g., because it is hidden behind another frame), the rule-based model can determine that no attention is being paid to the video advertisement.
[0076] Additionally or alternatively, the attention predictor may include an AI-based model that returns the probability that the user is paying attention to the content on the screen, based on the context attribute data 124 and the supplementary context attribute data 126. This model may be trained (and updated) using data from users who consume the same media content and for whom action data (or actual attention data) and related context attribute data are available. Such a model may provide enhanced attention recognition capabilities compared to conventional rule-based or statistical-based models.
[0077] In addition to generating the rich dataset described above, the analytics server 130 may be configured to determine specific attention measurement criteria for a given media content. An example of an attention measurement criterion is the amount of attention. The amount of attention may be defined as the average amount of attention responses paid to the media content by the respondents. For example, an attention amount score of 50% means that, on average, half of the viewers paid attention to the content throughout the video. The higher the number of seconds the video can attract the audience's attention, the higher this score will be. Another example of an attention measurement criterion is the quality of attention. The quality of attention may be defined as the percentage of the media content to which the respondents continuously paid attention on average. For example, a score of 50% means that, on average, the respondents paid attention to half of the video without interruption. This measurement criterion is different from the amount of attention because it determines the value of the score based not on the overall amount of attention but on how attention is distributed along with the viewing. If the respondents' attention period is short, the quality of attention decreases, indicating that the respondents are being distracted regularly.
[0078] The above measurement criteria, or other measurement criteria, may be to determine the degree to which attention was paid to a given viewed instance of media content delivered to the client device. This can be done, for example, by setting thresholds for the amount of attention and / or the quality of attention and determining that attention was paid to the viewed instance when one or both of the thresholds are exceeded. From the perspective of the brand owner or publisher, the advantage of this feature is that it can distinguish not only the number of impressions and views of a specific media content but also the views that attracted the user's attention and the views that the user was distracted from. Next, with the accompanying context attribute data, an attempt can be made to understand the levers that cause attention or distraction.
[0079] The system includes a report generator 138 configured to query a data store 136 to generate one or more reports 140 that can be provided to the brand owner 102, for example, directly or via the network 112. The report generator 138 may be a conventional computing device or server configured to query a database on a data store that includes the collected and synchronized data. Some examples of reports 140 will be described in more detail below with reference to FIGS. 4 and 5.
[0080] FIG. 2 is a flowchart showing the steps taken by the client device 106 and the analysis server 130 in a method 200 according to an embodiment of the present invention.
[0081] The method begins with step 202 of requesting and receiving web content by a client device via a network. Here, the web content means, for example, a web page that can be accessed and loaded from a domain hosted by the content server 114 as described above.
[0082] The header of the web page includes tags that include some preliminary checks to enable data collection from the client device and a bootstrap script configured to execute a process. Thus, the method continues with step 204 of executing the bootstrap script. One of the tasks performed by the script is to check for consent to share the collected data with the analysis server or to obtain permission. If applicable to the domain from which the web page was obtained, this may be done by referring to a content management platform (CMP). In this case, the bootstrap script is located after the code of the web page header that initializes the CMP.
[0083] This method continues with step 206 of checking or obtaining permission to share data. This can be done in a conventional manner by checking the current status of the CMP or by displaying a screen prompt. The permission is preferably requested at the domain level, so that, for example, repeated requests when accessing additional pages from the same domain are avoided. This method includes step 208 of checking the availability of the camera and obtaining consent to send the data collected from the camera to the analysis server. Even if there is no camera or the user does not consent to send data from the camera, the method may continue if the user consents to send the context attribute data. This is the scenario described above when there is no behavioral data stream.
[0084] If the camera is available and consent is given to send data from the camera, this method continues with step 210 of checking whether the user has been selected or sampled for behavioral data collection. In other embodiments, this step 210 may occur before step 208 of checking the availability of the camera.
[0085] Depending on the situation, all users with an available camera may be selected. However, in other examples, to ensure that an appropriate (e.g., random or pseudo-random) range of data is received by the analysis server 130 or to meet requirements set by the brand owner or publisher (e.g., to collect data from only one demographic sector), users may be selected. In another example, the ability to select users may be used to control the rate of data received by the analysis server. This may be useful when there are problems or limitations with network bandwidth.
[0086] If the user consents to sending action data from the camera, and if the user is selected, this method continues with step 212 of loading appropriate code to enable sharing of camera data via a web page. In one example, the sending of action data is done using the WebRTC protocol. It is preferable to delay loading the code for action data transmission until it is determined that the action data will actually be sent. By doing so, network resources (i.e., unnecessary traffic) are conserved and the quick loading of the initial page is facilitated.
[0087] After accessing the web page and executing the bootstrap script, this method may continue with step 214 of activating media content on the client device. Activating media content can mean starting the playback of media embedded in the web page or encountering an ad space on the web page that causes the playback of a video ad received from an ad server, for example, as a result of a conventional ad bidding process.
[0088] The playback of media content may be done, for example, by running a media player such as a video player. The media player may be embedded in the web page and configured to display the media content within an i-frame within the web page. Examples of suitable media players include Windows Media Player, QuickTime Player, Audacious, Amarok, Banshee, MPlayer, Rhythmbox, SMPlayer, Totem, VLC, and xine, or online video players such as JW Player, Flowplayer, VideoJS, and Brightcove.
[0089] As described above, it is desirable to send context attribute data related to the behavior and control of the media player, i.e., the analysis data of the media player, to the analysis server. To achieve this, the method continues with step 216 of loading an adapter for the media player configured to communicate media player analysis data to a web page (or, if it exists, executing a plugin of the media player), after which the media player analysis data can be sent to the analysis server.
[0090] This method continues with step 218 of sending context attribute data and, if applicable, step 220 of sending behavior data to the analysis server. If a camera is available and consent has been given, this means that the data sent to the analysis server is from the following three sources.
[0091] (1) Behavior data from the camera. This is usually an image or video from the camera itself. However, as described above, it is also possible for the client device itself to perform some preliminary analysis on the raw image data in order to measure attention and / or identify emotions. In this example, the behavior data sent to the analysis server may be data on attention and emotional state, and it is not necessary to send the image data.
[0092] (2) Context data from the web page. This is usually analysis data associated with the domain accessing the web page.
[0093] (3) Context data from the media player. This is usually analysis data associated with the media player on which the media content is being displayed.
[0094] Here, the method moves on to the actions taken at the analysis server and begins with step 222 of receiving the above data from the client device. The method also includes step 224 of obtaining, by the analysis server, media content that is the subject of the collected behavioral data and context attribute data. The analysis server may obtain the media content directly from the brand owner or content server, for example, based on an identifier transmitted by the client device. Alternatively, the analysis server may have a local store of media content.
[0095] This method is followed by step 226 of classifying the behavioral data regarding attention. In this step, individual images from the data captured by the camera on the client device are fed into an attention classifier, which evaluates the probability that the image shows a user paying attention to the media content. Thus, the output of the attention classifier may be a user's attention profile for the media content, where the attention profile shows the progression of attention over time during the period of the media content. In another example, the classifier may be binary, that is, it may generate an output of either "attentive" or "inattentive" for each frame. For such a two-state solution, an attention profile can also be generated. In another example, the classifier may be trained to include labels for input data where attention parameters cannot be obtained. For example, the classifier may be able to distinguish between a state where a user is present but the user's face cannot be read clearly enough to confirm whether they are paying attention and an unknown state. This can correspond to situations where no relevant signal is obtained from the camera. Thus, the classifier may output labels such as "attentive", "not attentive", "present", and "unknown".
[0096] The attention classifier or the analysis server may be configured to generate one or more attention measurement criteria for that particular viewing instance of the media content. The attention measurement criteria may be or may include the above attention quantity measurement criteria and attention quality measurement criteria.
[0097] This method follows step 228 of extracting emotional state information from the behavioral data stream. This may be done by an emotional state classifier and can be executed in parallel with step 226. The output of this step may be an emotional state profile showing the progression of one or more emotional states over time during the period of the media content.
[0098] As described above, the behavioral data stream may include image data captured by a camera, and the image data may be a plurality of image frames showing the face image of the user. When the image frames represent facial features such as the user's mouth, eyes, eyebrows, etc., the facial features may provide descriptor data points indicating the position, shape, orientation, sharing, etc. of a selected plurality of fiducial marks on the face. The descriptor data point of each facial feature may encode information indicating a plurality of fiducial marks on the face. The descriptor data point of each facial feature may be associated with each frame, for example, each image frame of a time-series of image frames. The descriptor data point of each facial feature may be a multi-dimensional data point, and each component of the multi-dimensional data point indicates each fiducial mark on the face.
[0099] The emotional state information may be obtained directly from the raw behavioral data input, or from the descriptor data points extracted from the image data, or from a combination of the two. For example, the plurality of fiducial marks on the face may be selected to include information that can characterize the user's emotions. In one example, the emotional state data may be determined by applying a classifier to the descriptor data points of one or more facial features within one image or over a series of images. In some examples, deep learning techniques can be utilized to generate emotional state data from the raw data input.
[0100] The user's emotional state may include one or more emotional states selected from anger, disgust, fear, happiness, sadness, and surprise.
[0101] If the data received by the analysis server does not include behavior data, steps 226 and 228 above may be omitted. Instead, the method includes step 230 of predicting attention data using context attribute data. In this step 230, the context attribute data is fed to an attention predictor, which evaluates the probability of an image depicting a user paying attention to media content. The attention classifier is an AI-based model trained with annotated face images (described in more detail below), and the attention predictor is an AI-based model trained with context attribute data for which attention data is available. Thus, the attention predictor can transform information related to the environment in which the viewed instance of the media content occurs and the interaction between the user and the client device during that viewed instance.
[0102] Accordingly, the output of the attention predictor may be similar to an attention classifier, e.g., a user's attention profile for media content, where the attention profile indicates the progression of attention over time during the period of the media content. The output from either or both of the attention classifier and the attention predictor may be weighted, e.g., according to the confidence associated with the collected data. For example, the output of the attention classifier may be weighted based on the detected angle of the camera from which the behavior data is collected. If the user is not facing the camera, the reliability of the output may be reduced.
[0103] The method continues with step 232 of synchronizing attention profile 232 with corresponding context attribute data and emotional state data to generate a rich "validity" dataset in which the context of periods of attention and distraction in the attention profile is associated with various elements of the relevant context or emotional state data.
[0104] This method follows step 234 of aggregating a validity dataset obtained for multiple viewed instances of media content from multiple client devices (e.g., different users). The aggregated data is stored in a data store and can be queried therefrom to generate reports of the type described below with reference to FIGS. 4 and 5.
[0105] FIG. 3 is a schematic diagram of a data collection and analysis system 300 for generating an attention classifier suitable for use in the present invention. It can be understood that the system of FIG. 3 shows components for collecting and annotating data, as well as components for the subsequent use of that data in generating and utilizing an attention classifier.
[0106] System 300 is provided in a networked computing environment where several processing entities are communicably connected via one or more networks. In this example, system 300 includes one or more client devices 302 configured to play media content, for example, via a speaker or headphones and a display 304. Client device 302 may also be equipped with, or connected to, behavior data capture devices such as a webcam 306, a microphone, etc. Exemplary client devices 302 include smartphones, tablet computers, laptop computers, desktop computers, and the like.
[0107] System 300 may also include one or more client sensor units, such as a wearable device 305, that collect physiological information from the user while the user is consuming media content on client device 102. Examples of physiological parameters that can be measured include voice analysis, heart rate, heart rate variability, skin electrical activity (which may indicate arousal), respiration, body temperature, electrocardiogram (ECG) signals, and electroencephalogram (EEG) signals.
[0108] The client device 302 is communicatively connected via the network 308 such that, for example, it can receive media content 312 to consume from a content provider server 310.
[0109] The client device 302 may be further configured to transmit the collected behavioral information via the network for analysis or further processing at a remote device such as the analytics server 318. As described above, "behavioral data" or "behavioral information" herein may refer to the visual aspects of a user's response. For example, the behavioral information may include facial reactions, head and body gestures or postures, and eye tracking.
[0110] In this example, the information transmitted to the analytics server 318 may include the user's facial reactions 316 in the form of, for example, a set of captured videos or images of the user while consuming media content. The information may also include the associated media content 315, or a link or other identifier that enables the analytics server 318 to access the media content 312 consumed by the user. The associated media content 315 may include information regarding how the media content was played back on the client device 302. For example, the associated media content 315 may include information related to user commands such as pause / play, stop, volume control, etc. Additionally or alternatively, the associated media content 315 may include other information related to delays or interruptions in playback, such as due to buffering. This information may correspond to (and be obtained in a similar manner as) the analysis data from the media player described above. Thus, the analytics server 318 may effectively receive a data stream that includes information regarding the user's response to the media content.
[0111] The information sent to the analysis server 318 may also include physiological data 114 obtained regarding the user while consuming media content. The physiological data 314 may be sent directly by the wearable device 305, or the wearable device 305 may be paired with one or more client devices 302 configured to send and receive data from the wearable device 305. The client device 302 may be configured to process raw data from the wearable device, such that the physiological data 314 sent to the analysis server 318 may include data that has already been processed by the client device 302.
[0112] In this example, the purpose of collecting information regarding the user's response to media content is to enable annotating that response with attention labels. In one example, this annotation process may include establishing a time series of attention scores that are mapped to one or more time series of behavioral characteristic parameters received at the analysis server 318. For example, the time series of attention scores may be associated with an image or video of the user collected while the user is consuming media content. Other behavioral characteristic parameters, such as emotional state information, physiological information, etc., may be synchronized with the image or video of the user. Thus, the output of the annotation process may be a rich data stream representing the user's behavioral characteristics, including attention, in response to the media content.
[0113] System 300 includes an annotation tool 320 that facilitates the execution of an annotation process. The annotation tool 320 can include a computer terminal that communicates (e.g., network communication) with an analysis server 318. The annotation tool 320 includes a display 322 that presents a graphical user interface to a human annotator (not shown). The graphical user interface can take various forms. However, the graphical user interface can usefully include several functional elements. First, the graphical user interface can present the collected user response data 316 (e.g., a set of face images or videos showing the movement of the user's face) in synchronization with the associated media content 315. In other words, the user's facial reaction is displayed simultaneously with the relevant media content that the consumer was viewing. In this way, the annotator can (consciously or unconsciously) recognize the context in which the user's response occurred. In particular, the annotator may be able to adjudicate attention based on the reaction to an event within the associated media content, or may be sensitive to external events that may distract the user.
[0114] The graphical user interface can include a controller 324 for controlling the playback of the synchronized response data 316 and the associated media content. For example, the controller 324 can enable the annotator to play, pause, stop, rewind, fast forward, backstep, forward step, scroll back, forward scroll, etc. through the presented material.
[0115] The graphical user interface may comprise one or more score applicators 326 for applying an attention score to one or more portions of the response data 316. In one example, the score applicator 326 may be used to apply an attention score to a period of a set of video or image frames corresponding to a given period of the user's response. The attention score may have any suitable format. In one example, the attention score is binary, i.e., indicates attention by a simple yes / no. In other examples, the attention score may be selected from a set number of predetermined levels (e.g., high, medium, low), or from a numerical range between limits representing no attention (i.e., none) and high attention, respectively (e.g., a linear scale).
[0116] From the perspective of increasing potential annotator resources, it may be desirable to simplify the annotation tool. The simpler the annotation process, the less training the annotator needs to participate in. In one example, the annotated data may be harvested using a crowdsourcing approach.
[0117] Accordingly, the annotation tool 320 may represent a device for receiving time series data indicative of the user's attention while consuming media content. The attention data may be synchronized with the response data 316 (e.g., depending on how the score is applied). The analysis server 318 may be configured to collate or otherwise combine the received data to generate attention-labeled behavior data 330 that can be stored in a suitable storage device 328.
[0118] Attention data from multiple annotators may be aggregated or otherwise combined to generate an attention score for a given response. For example, attention data from multiple annotators may be averaged across portions of the media content.
[0119] In one embodiment, the level of agreement among multiple annotators itself may be used as a way to quantify attention. For example, the annotation tool 320 may enable each annotator to score the response data with a binary option of either (a) attentive or (b) inattentive for the user. In other examples, the annotation tool may include states corresponding to the above "present" and "unknown" labels. The annotator tool 320 may present one or more reason fields where the annotator can provide reasons for the binary selection. There may be, for example, a drop-down list of predetermined reasons that can be entered into the field. The predetermined reasons may include common reasons for attention or inattention, such as "turning away from the face", "not looking at the screen", "talking", etc. The field may also allow free text input. The attention data from each annotator may include the results of binary selections for various periods within the response data, along with the associated reasons. The reasons may be used to evaluate situations where there is a high level of disagreement among annotators or where the attention model outputs results that do not match the observations. This may occur, for example, when similar facial movements correspond to different actions (such as talking / eating, etc.).
[0120] The analysis server 318 may be configured to receive attention data from multiple annotators. The analysis server 318 may generate combined attention data from different sets of attention data. The combined attention data may include an attention parameter indicating the level of positive correlation between the attention data from multiple annotators. In other words, the analysis server 318 may output a score that quantifies the level of agreement between the binary selections made by multiple annotators across the entire response data. The attention parameter may be a time-varying parameter, i.e., the score indicating agreement may vary over the period of the response data to indicate an increase or decrease in the correlation.
[0121] In the development of this concept, the analysis server 318 may be configured to determine and store a trust value associated with each annotator. The trust value may be calculated based on how well the individual scores of the annotator correlate with the attention data combined. For example, annotators who typically score in a direction opposite to that of the annotator group as a whole may be assigned a lower trust value than those who are more in agreement. For example, the trust value may be dynamically updated when more data is received from an individual annotator. The trust value may be used to weight the attention data from each annotator in the process of generating the combined attention data. Thus, the analysis server 318 may demonstrate the ability to "adjust" itself for more accurate scoring.
[0122] The attention-labeled behavior data 330 may include attention parameters. In other words, the attention parameters may be associated with events in the data stream or media content, for example, synchronized, or otherwise mapped or linked.
[0123] The attention-labeled behavior data 330 may include any one or more of the original data 316 collected from the client device 302 (e.g., raw video or image data, also referred to herein as response data), time-series attention data, time-series data corresponding to one or more physiological parameters from the physiological data 314, and emotional state data extracted from the collected data 316.
[0124] The collected data may be image data captured at each of the client devices 302. The image data may include a plurality of image frames showing the user's face image. Further, the image data may include a time series of image frames showing the user's face image.
[0125] The image frame shows the user's facial features, such as the mouth, eyes, eyebrows, etc. When each facial feature includes a plurality of landmarks on the face, the behavioral data may include information indicating the position, shape, orientation, shadow, etc. of the landmarks on the face of each image frame.
[0126] The image data may be processed on each client device 302 or streamed to the analysis server 318 via the network 308 for processing.
[0127] The facial features may provide descriptor data points indicating the position, shape, orientation, sharing, etc. of a selected plurality of landmarks on the face. The descriptor data points of each facial feature may encode information indicating a plurality of landmarks on the face. The descriptor data points of each facial feature may be associated with each frame, for example, each image frame in a time-series of image frames. The descriptor data points of each facial feature may be multi-dimensional data points, and each component of the multi-dimensional data point indicates each landmark on the face.
[0128] The emotional state information may be obtained directly from raw data inputs, or from the extracted descriptor data points, or from a combination of the two. For example, a plurality of landmarks on the face may be selected to include information that can characterize the user's emotions. In one example, the emotional state data may be determined by applying a classifier to the descriptor data points of one or more facial features within one image or across a series of images. In some examples, deep learning techniques can be utilized to generate emotional state data from raw data inputs.
[0129] The user's emotional state may include one or more emotional states selected from anger, disgust, fear, happiness, sadness, and surprise.
[0130] The creation of the attention-labeled behavioral data represents the first function of the system 300. The second function described below is to use that data later to generate and utilize an attention model for the attention classifier 132.
[0131] System 300 may include a modeling server 332, which is configured to communicate with a storage device 328 and access attention-labeled action data 330. The modeling server 332 may be directly connected to the storage device 328 as shown in FIG. 3 or connected via a network such as network 308.
[0132] The modeling server 332 is configured to apply machine learning techniques 334 to a training set of the attention-labeled action data 330 to establish a model 336 that scores attention from unlabeled response data, such as response data 316 initially received by the analysis server 318. The model may be established as an artificial neural network trained to recognize patterns of collected response data that exhibit a high level of attention. Thus, this model can be used to automatically score the collected response data with respect to attention without human input. The advantage of this technique is that the model is fundamentally based on a direct measurement of attention that is sensitive to measurements or involvements or contextual factors that may be missed by a particular predetermined proxy.
[0133] In one example, the model 336 combines two types of neural network architectures, namely, a convolutional neural network (CNN) and a long short-term memory neural network (LSTM).
[0134] The CNN portion is trained with images of respondents obtained from individual video frames. The final layer representation of the CNN is then used to generate a time sequence for training the LSTM.
[0135] By combining these two architectures, a model is constructed that (i) learns useful spatial information extracted from face and upper body images using a CNN, and (ii) learns useful temporal patterns of facial expressions and gestures using an LSTM, which helps the model determine whether it is looking at an attentive face or a distracted face.
[0136] In one example, the attention-labeled behavioral data 330 used to generate the attention model 336 may also include information about the media content. This information may be related to how the media content is manipulated by the user, such as being paused or otherwise controlled. Additionally or alternatively, the information may include data about the subject of the media content being displayed, for example, to provide context to the collected response data.
[0137] As used herein, media content may be any type of user-consumable content for which information about user feedback is desirable. The present invention may be particularly useful when the media content is commercial (e.g., a video commercial or advertisement) and the user's engagement or attention is likely to be closely related to performance, such as an increase in sales. However, the present invention is applicable to any type of content, such as video commercials, audio commercials, movie trailers, movies, web advertisements, animated games, images, and the like.
[0138] FIG. 4 is a screenshot of a report dashboard 400 that includes a presentation of rich validity data stored in the data store 136 of FIG. 1 for various different media content, such as a group of advertisements in a common field. The common field may be indicated by a main heading 401 shown as "Sports Apparel" in FIG. 4, but may be changed, for example, by a user selecting from a drop-down list.
[0139] Dashboard 400 includes an impression categorization bar 402, which is the relative ratio of the total impressions supplied that were (i) visible (i.e., seen on the screen) and (ii) seen by users with an attention score exceeding a predetermined threshold (i.e., "attentive viewers"). A reference may be marked on the bar to indicate how the visibility-to-attention ratio compares to the expected performance.
[0140] Dashboard 400 may further include a relative emotional state bar 404 that indicates the relative strength of the emotional states detected from attentive viewers.
[0141] Dashboard 400 further includes a driver indicator bar 406, which, in this example, indicates the relative amounts correlated to the attention where different context attribute categories are detected. Each of the context attribute categories (e.g., creative, brand, audience, and context) may be selectable to provide a more detailed breakdown of the factors contributing to that category. For example, the "creative" category may be related to the information presented in the media content. The context attribute data may include a content stream that describes the major items visible at any point in the media content. In FIG. 4, driver indicator bar 406 shows the correlation of the categories to the attention. However, other features may be selectable where the relative strength of the correlation to a category, such as a particular emotional state, is of interest.
[0142] Dashboard 400 further includes a brand attention chart 408, which shows the time - based progression of the levels of attention achieved by various brands in the common field shown in the main heading 401.
[0143] Dashboard 400 further includes a series of charts that divide impression categorization by context attribute data. For example, chart 410 breaks down impression categorization by displaying device type, and chart 412 breaks down impression categorization using gender and age information.
[0144] Dashboard 400 further includes a map 414 where relative attention is indicated using location information from context attribute data.
[0145] Dashboard 400 further includes a domain comparison chart 416 that compares the amount of attention associated with the web domains from which impressions were obtained.
[0146] Finally, dashboard 400 may further include a summary panel 418 that classifies campaigns covered by common fields according to a predetermined attention threshold. In this example, the threshold is 10%, which means that 10% of the impressions are detected as having attentive viewers.
[0147] FIG. 5 is a screenshot of an advertising campaign report 500 that includes a presentation of rich effectiveness data stored in the data store 136 of FIG. 1 for a particular advertising campaign 501 that can be represented by a single media content (e.g., video advertisement) or a group of related media content.
[0148] Advertising campaign report 500 may include an impression categorization bar 502 that indicates the relative ratio of the total impressions delivered under a selected campaign that were (i) visible (i.e., seen on the screen), and (ii) had an attention score exceeding a predetermined threshold (i.e., "attentive viewers"). A baseline may be marked on the bar to show how the visibility and attention ratio compares to the expected performance.
[0149] Advertising campaign report 500 may further include a chart 504 showing the progress of the impression categorization bar over time.
[0150] Advertising campaign report 500 may further include a relative emotional state bar 506 showing the relative strength of the emotional states detected from attentive viewers.
[0151] Advertising campaign report 500 may further include a driver indicator bar 508, which, in this example, shows the relative amount correlated with the attention where different context attribute categories are detected. Each of the context attribute categories (e.g., creative, brand, audience, and context) may be selectable to provide a more detailed breakdown of the factors contributing to that category. For example, the "creative" category may be related to the information presented in the media content. The context attribute data may include a content stream that describes the major items visible at any point in the media content. In FIG. 5, the driver indicator bar 508 shows the correlation of the category with attention. However, other features may be selectable where the relative strength of the correlation with a category, such as a particular emotional state, is of interest.
[0152] Advertising campaign report 500 may further include a recommendation panel 510 where various proposals for adapting or maintaining the campaign strategy are provided. Each proposal includes the associated cost and the predicted effect on the attention to the campaign. The prediction is made using the information detected in that campaign. The proposals may be driven by pre-determined campaign optimization targets.
[0153] Advertising campaign report 500 may further include a prediction panel 512 that tracks the past performance of the campaign and shows the effect of executing the proposals from the recommendation panel 510.
[0154] Finally, the advertising campaign report 500 may further include a keyword display panel 514 where data from the contextual attribute data is displayed. The data may include segment data used to identify various user types and / or common terms that appear in the contextual attribute data.
[0155] The advertising campaign report 500 may be used to control a programmatic advertising campaign. The control may be performed manually, for example, by adapting instructions to the DSP based on the recommendations provided in the report. However, it may be particularly useful to implement an automatic adjustment of programmatic advertising instructions in order to effectively establish an automatic feedback loop that optimizes the programmatic advertising strategy to achieve the campaign goals.
[0156] The term "programmatic advertising" is used herein to refer to an automated process for purchasing digital advertising space, such as on a web page, an online media player, etc. Typically, the process includes real-time bidding for each advertising slot (i.e., each impression of an available advertisement). In programmatic advertising, the DSP operates to automatically select bids in response to impressions of available advertisements. The bids are selected based in part on a determined level of correspondence between the campaign strategy provided by the advertiser to the DSP and the contextual information regarding the advertisement impression itself. The campaign strategy identifies the target audience, and the bid selection process operates to maximize the likelihood that an advertisement will be delivered to a portion within that target audience.
[0157] In this context, the present invention can be used as a means to adjust, in real time and preferably in an automated manner, the campaign strategy provided to the DSP. In other words, the definition of the target audience of a given advertising campaign may be adjusted using the recommendations output from the analysis server.
[0158] FIG. 6 is a flowchart of a method 600 for optimizing a digital advertising campaign. This method is applicable to programmatic advertising technology. In programmatic advertising technology, a digital advertising campaign has a defined goal and a target audience strategy aimed at achieving that goal. The target audience strategy may form an input to a demand-side platform (DSP) that is responsible for the ad content delivered to users in a way that achieves the defined goal.
[0159] Method 600 begins at step 602 of accessing an effectiveness dataset that represents the time-dependent evolution of attention parameters during the playback of ad content belonging to a digital advertising campaign to a plurality of users. The effectiveness dataset may be of the type described above, and the attention parameters are obtained by applying behavioral data collected from each user during the playback of the ad content to a machine learning algorithm trained to map the behavioral data to the attention parameters.
[0160] This method is followed by step 604 of generating candidate adjustments to the target audience strategy associated with the digital advertising campaign. With the candidate adjustments, any applicable parameters of the target audience strategy can be changed. For example, the candidate adjustments may change the demographic attributes or interests of the target audience. A plurality of candidate adjustments may be generated. The candidate adjustments may be generated based on information from the effectiveness dataset of the digital advertising campaign. For example, the candidate adjustments may aim to increase the influence of a part of the target audience with relatively high attention parameters or decrease the influence of a part of the target audience with relatively low attention parameters.
[0161] This method is followed by step 606 of predicting the effect on the attention parameters of applying the candidate adjustments. This may be done in the manner described above with reference to FIG. 5.
[0162] This method follows step 608 of evaluating the predicted effect on the campaign goals of a digital advertising campaign. This may also be done in the manner described above with reference to FIG. 5. The goals of the campaign may be quantified by one or more parameters. Thus, in the evaluation step, the predicted values of these parameters are compared with the current values of the digital advertising campaign. In one example, the goal of the campaign may relate to maximizing attention, and thus, improving the target audience strategy would manifest as an increase in the attention parameter.
[0163] This method follows step 610 of updating the target audience strategy using candidate adjustments when the predicted effect improves the performance against the campaign goal beyond a threshold. In the above example, this may be an improvement in the attention parameter (e.g., attention share achieved by the advertising campaign) beyond the threshold. The update may be performed automatically, i.e., without human intervention. Thus, the target audience strategy may be automatically optimized.
[0164] As described above, the present invention may be used when measuring the effectiveness of advertisements. However, it may also be used in other fields.
[0165] For example, the present invention may be used for evaluating online teaching materials such as video lectures and webinars. It may also be used for measuring attention to locally displayed text, survey questions, etc. In this context, it can be used, for example, to evaluate the effectiveness of the content itself or individual learners, such as when sufficient attention has been paid to the training materials before being permitted to take a test.
[0166] In another example, the present invention may be used in a game application that is executed locally on a client device or executed online with single or multiple participants. Any aspect of gameplay may provide display content where attention is measurable. For example, the present invention may be used to understand whether a particular episode of a game results in the desired or required level of attention or emotional reaction. Further, the present invention may be used as a tool for indicating and measuring the effectiveness of changes to gameplay.
Claims
**Claim 1**: A method for collecting data to determine the attention paid to the display of content, which is executed by a plurality of computers, wherein the plurality of computers includes a content server, an analysis server, and a plurality of client devices communicable with each other via a network, displaying the content on the client device, transmitting, from the client device to the analysis server via the network, context attribute data indicating the interaction between the user and the client device during the display of the content, collecting, on the client device, the user's behavior data during the display of the content, applying the behavior data to a classification algorithm to generate the user's attention data, wherein the classification algorithm is a machine learning algorithm trained to map the behavior data to attention parameters, and the attention data indicates the variation of the attention parameters over time during the display of the content, synchronizing, on the analysis server, the attention data with the context attribute data to generate a validity dataset linking the temporal progression of the attention parameters to the corresponding context attribute data obtained during the display of the content, storing the validity dataset in a data store, including, when it is determined that the behavior data from the client device is unavailable, applying the context attribute data to a prediction algorithm to generate the user's predicted attention data, wherein the prediction algorithm is a machine learning algorithm trained to map the context attribute data to attention parameters, and the predicted attention data indicates the variation of the attention parameters over time during the display of the content, including synchronizing the predicted attention data with the context attribute data to generate a predicted validity dataset linking the temporal progression of the attention parameters to the corresponding context attribute data obtained during the display of the content. **Claim 2** wherein the displayed content includes media content, and the method Further comprising playing the media content using a media player application executed on the client device, wherein the context attribute data further indicates an interaction between the user and the media player application during playback of the media content, The method according to claim 1. **Claim 3** The method according to claim 2, wherein the media player application comprises an adapter module configured to transmit control analysis data of the media player application to the analysis server via the network, and the method comprises executing the adapter module when receiving the media content to be displayed. **Claim 4** Displaying the content comprises accessing, by the client device via the network, a web page on a web domain hosted by a content server, receiving, by the client device via the network, the content displayed by the web page The method according to any one of claims 1 to 3. **Claim 5** The method according to claim 4, wherein the context attribute data further indicates an interaction between the user and the web page during display of the content. **Claim 6** The method according to claim 4 or 5, wherein accessing the web page comprises obtaining a context data start script for execution on the client device, and the method further comprises executing the context data start script on the client device. **Claim 7** The method according to claim 6, further comprising injecting, by an intermediator on the network between the content server and the client device, the context data start script into the source code of the web page. **Claim 8** Obtaining the context data start script comprises sending, by the client device, an advertisement request, receiving, from an advertisement server, a video advertisement response in response to the advertisement request, wherein the context data start script is included in the video advertisement response, The method according to claim 6. **Claim 9** When the context data start script is executed, the method comprises Determining consent to send the context attribute data and the behavior data to the analysis server; Determining the availability of a device for collecting the behavior data; Checking whether the user has been selected for behavior data collection further comprising; The method is performed by the client device using the context data start script, (i) consent to the transmission of behavior data is withheld, or (ii) a device for collecting the behavior data is unavailable, or (iii) the user has not been selected for behavior data collection If it is determined that the behavior data collection procedure is terminated, further comprising; The method according to any one of claims 6 to 8.
10. When the method determines, by the client device using the context data start script, (i) that consent has been given to transmit the behavior data, and (ii) that a device for collecting the behavior data is available, and (iii) that the user has been selected for behavior data collection, further comprising loading a real-time communication protocol for transmitting the behavior data from the client device to the analysis server. The method according to claim 9.
11. The method according to any one of claims 4 to 10, wherein the context attribute data includes web analysis data of the web page.
12. Applying the behavior data to the classification algorithm is performed on the client device, and the method further comprises transmitting the attention data to the analysis server via the network by the client device. The method according to any one of claims 1 to 11.
13. Further comprising transmitting the behavior data to the analysis server via the network by the client device, and applying the behavior data to a classification algorithm is performed on the analysis server. The method according to any one of claims 1 to 11.
14. Collecting the user's behavior data on the client device includes capturing an image of the user using a camera. The method according to any one of claims 1 to 13.
15. The method according to claim 14, wherein the classification algorithm operates to evaluate the attention parameter of each of a plurality of images of the user captured during playback of the content.
16. Applying the behavioral data to an emotional state classification algorithm to generate emotional state data of the user, wherein the emotional state classification algorithm is a machine learning algorithm trained to map behavioral data to emotional state data, and the emotional state data indicates a temporal variation in the probability that the user has a given emotional state during playback of the content, generating the emotional state data of the user; Synchronizing the emotional state data with the attention data; Further comprising, wherein the validity dataset further comprises the emotional state data. The method according to any one of claims 1 to 15.
17. The content is acquired and displayed by an application executed on the client device, and the method Determining an action based on the emotional state data and the attention parameter data by the application executed on the client device. The method according to claim 16, further comprising.
18. The method according to claim 17, wherein the application includes a software development kit configured to provide the classification algorithm.
19. Receiving a query for information of the validity dataset by a report generator via the network; Extracting response data in response to the query from the data store by the report generator; Transmitting the response data via the network by the report generator. The method according to any one of claims 1 to 18, further comprising.
20. Receiving context attribute data and behavioral data from a plurality of client devices by the analysis server; Aggregating a plurality of validity datasets obtained from the context attribute data and the behavioral data received from the plurality of client devices by the analysis server. The method according to any one of claims 1 to 19, further comprising.
21. The method according to claim 20, wherein the plurality of validity data sets are aggregated with respect to one or more common dimensions shared by the context attribute data and the behavior data received from the plurality of client devices. **Claim 22** The method according to claim 21, wherein the common dimension includes any one of a web domain, a website identity, a time, and a content type. **Claim 23** The content is acquired and displayed by an application executed on the client device, and the method includes using the aggregated validity data set to determine a software update for the application; receiving the software update on the client device; and executing the software update to adjust the function of the application. The method according to any one of claims 20 to 22, further comprising. **Claim 24** When it is determined that behavior data consented to by a user is available from the client device, applying the context attribute data to a second prediction algorithm to generate prediction attention data for the user, wherein the prediction attention data indicates a change over time of the attention parameter during display of the content; synchronizing the prediction attention data with the context attribute data to generate a prediction validity data set that links the progress of the attention parameter over time with corresponding context attribute data acquired during display of the content. The method according to any one of claims 1 to 23, further comprising. **Claim 25** The method according to claim 24, wherein the second prediction algorithm is a machine learning algorithm trained to map context attribute data to an attention parameter. **Claim 26** A system for collecting data for determining attention paid to displayed content, the system including a plurality of client devices communicable via a network with a content server and an analysis server, each client device is configured to display content; send context attribute data indicating an interaction between the user and the client device during display of the content to the analysis server; and collect behavior data of the user during display of the content. The system is configured as follows. The system includes Further configured to apply the received behavior data to a classification algorithm to generate the user's attention data, the classification algorithm being a machine learning algorithm trained to map behavior data to attention parameters, the attention data indicating the variation of the attention parameters over time during the display of the content, Further configured to apply the context attribute data to a prediction algorithm to generate the user's predicted attention data when the behavior data from the client device is unavailable, the prediction algorithm being a machine learning algorithm trained to map the context attribute data to attention parameters, the predicted attention data indicating the variation of the attention parameters over time during the display of the content, The analysis server, To generate a validity dataset that links the temporal progression of the attention parameters to the corresponding context attribute data obtained during the display of the content, synchronize the attention data with the context attribute data, and Store the validity dataset in a data store The system is further configured as such.
Citation Information
Patent Citations
Ad selection via viewer feedback
JP2014520329A
Affect based political advertisement analysis
US20130218663A1
Method in support of video impression analysis including interactive collection of computer user data
US20160198238A1
Method of collecting and processing computer user data during interaction with web-based content
US20170139802A1
Method of targeting web-based advertisements
US20170249663A1