Media device on / off detection using return path data
By training machine learning algorithms and combining meter data and return path data from ordinary residences, the problem that return path data cannot directly provide the on/off status of media devices was solved, thus improving the accuracy of audience measurement and the representativeness of the sample.
Patent Information
- Application Number
- CN202080058757.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-11-27
- Filing Date
- 2020-06-16
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2040-06-16
AI Technical Summary
In existing technologies, the returned path data cannot directly provide information about the on/off operating status of media devices connected to the set-top box, resulting in insufficient accuracy in audience measurement.
By training a model using machine learning algorithms and combining it with meter data and return path data from ordinary residences, the on/off status of media devices connected to set-top boxes can be predicted.
It improves the accuracy of audience measurement, reduces data sample bias based on group residences, and enhances the ability to judge the actual on/off status of media devices.
Smart Images

Figure CN114258556B_ABST
Abstract
Description
[0001] Related Applications
[0002] This patent is a continuation of U.S. Patent Application Serial No. 16 / 698,167, filed November 27, 2019, entitled “MEDIA DEVICE ON / OFF DETECTION USING RETURN PATH DATA,” which claims the benefit of U.S. Provisional Application Serial No. 62 / 863,131, filed June 18, 2019, entitled “MEDIA DEVICE ON / OFF DETECTION USING RETURN PATH DATA.” Priority is claimed to U.S. Patent Application Serial No. 16 / 698,167 and U.S. Provisional Application Serial No. 62 / 863,131. U.S. Patent Application Serial No. 16 / 698,167 and U.S. Provisional Application Serial No. 62 / 863,131 are hereby incorporated by reference in their respective entireties. TECHNICAL FIELD
[0003] The present invention relates generally to media device on / off detection, and more particularly, to media device on / off detection using return path data. BACKGROUND
[0004] Set-top boxes (STBs) in cable and satellite subscriber homes access subscriber viewing data, including the user’s television tuning data, on a second-by-second basis. The viewing data can include the programs viewed by the user, while the tuning data can include the location of the user’s home, changes in channels, times of access to programs, and the like. The STBs report return path data (RPD) that includes such television tuning data and viewing data back to the multi-channel video program distribution provider (e.g., cable and satellite providers). BRIEF DESCRIPTION OF DRAWINGS
[0005] Figure 1 is a block diagram illustrating an example operating environment for a media device on / off detector to implement media device on / off detection using return path data in accordance with the teachings of the present invention.
[0006] Figure 2 is a block diagram of an example implementation of the media device on / off detector of Figure 1
[0007] Figure 3 is a flowchart of example computer-readable instructions that can be executed to implement the media device on / off detector of Figure 1 in accordance with the teachings of the present invention to perform media device on / off detection using return path data.
[0008] Figure 4 is a flowchart representing example computer-readable instructions that can be executed by a media device on / off detector to train a machine learning algorithm using return path based training data.
[0009] Figures 5A-5B includes example validation metrics indicating that example machine learning algorithm training according to the teachings of this disclosure can result in improved accuracy as compared to a reference ordinary residence tuning method.
[0010] Figures 6A-6B includes examples of changes in tuning minutes and residual tuning minute percentages when using example machine learning algorithms trained based on ordinary residence return path data and panel meter data according to the teachings of this disclosure.
[0011] Figure 7 is a block diagram of an example processor platform of a media device on / off detector configured to execute Figure 3 and / or Figure 4 example computer-readable instructions to implement Figure 1 and / or Figure 2 .
[0012] The drawings are not to scale. In general, the same reference numbers in the drawings have been used for common or similar items, elements, or functions. In some instances, one described feature, structure, or element can be
[0013] When identifying multiple elements or components that can be individually referred to, the descriptors "first," "second," "third," etc. are used herein. Unless otherwise specified or understood based on their usage context, such descriptors are not intended to convey any priority or order in time, but merely serve as labels to refer to multiple elements or components, respectively, for ease of understanding the disclosed examples. In some examples, the descriptor "first" can be used to refer to an element in the detailed description, while the same element can be referred to in the claims with a different descriptor (e.g., "second" or "third"). In such instances, it will be understood that the use of such descriptors is merely for ease of referencing the multiple elements or components. DETAILED DESCRIPTION
[0014] Example technical solutions are disclosed for performing media device on / off detection using return path data. Such example technical solutions can include one or more of a method, apparatus, system, article of manufacture (e.g., a physical storage medium), etc. for performing media device on / off detection using return path data according to the teachings of this disclosure.
[0015] Many residential entertainment systems include a set-top box (STB) for receiving media from a service provider and displaying the media on a media device, such as a television. Examples of service providers include cable television providers, satellite television providers, over-the-top (OTP) service providers, Internet service providers, and the like. Audience measurement entities (AMEs), such as The Nielsen Company (US), LLC, monitor viewing of media presented by such media devices. For example, an AME can extrapolate ratings metrics and / or other audience measurement data for total television viewers from a relatively small panel of residential samples. The panel homes can be well-studied and are typically selected to be representative of the entire audience universe. However, accurately representing the geographic distribution and demographic diversity present in the total audience population with a small panel of residential samples remains a challenge. Incorporating additional streams of information about media exposure into the total audience population can fill in any gaps or biases inherent to a statistical sample.
[0016] To help supplement panel data, an AME, such as The Nielsen Company (US), LLC, can enter into agreements with pay television provider companies to obtain television tuning information obtained from STBs and / or other devices / software, which is referred to herein and in the industry as return path data. STB data includes all data collected by the STB. The STB data can include, for example, tuning data related to tuning events and / or commands received by the STB (e.g., power on, power off, change channel, change input source, record a presentation of media, volume up / down, etc.). The STB data can also include viewing data related to the type of media content accessed by the user (e.g., advertisement, movie, etc.) and the time at which the media content was accessed (e.g., time / date the presentation of media began, time the presentation of media completed, when the presentation of media was paused, etc.). The STB data can additionally or alternatively include commands sent by the STB to the content provider (e.g., switch input source, record a presentation of media, delete a recorded presentation of media, etc.), heartbeat signals, and the like. The STB data can additionally or alternatively include a household identification (e.g., household ID) and / or a STB identification (e.g., STB ID).
[0017] Return path data includes any data receivable at a media service provider (e.g., a cable television service provider, a satellite television service provider, a streaming media service provider, a content provider, etc.) through a return path from a media consumer site to the service provider. As such, return path data includes at least a portion of STB data. Return path data can additionally or alternatively include data from any other consumer device with network access capabilities (e.g., through a cellular network, the Internet, other public or private networks, etc.). For example, return path data can include any or all linear real-time data from the STB, guide user data from a guide server, clickstream data, tuning data associated with key stream data (e.g., any clicks on the remote - volume, mute, etc.), viewing data associated with interactive activities (e.g., video on demand), and any other additional data (e.g., data from an intermediary device). RPD data can additionally or alternatively be from the network (e.g., through a digital exchange software) and / or any cloud-based data from the cloud (e.g., a remote server DVR).
[0018] RPD can provide insight into media exposure associated with a larger audience group. However, RPD can not directly provide information about the operational state of the media device(s) connected to the STB reporting the RPD, such as the on / off operational state of the media device connected to the STB. Determining the operational state of the media device connected to the STB can be important to accurately judge exposure associated with media output from the STB. For example, the media device connected to the STB can be turned off, while the STB remains powered on and outputs media, either unintentionally or intentionally. For example, while televisions can be turned off, STBs remain in an on state, with approximately 10% of STBs never being turned off for more than a month (e.g., approximately 30% of STBs remain on for 24 hours on any given day). In such examples, knowledge of the operational state of the media device can help the AME accurately judge whether media output from the STB was actually presented by the media device.
[0019] The example technical solutions disclosed herein predict the on / off operating state of media devices connected to STBs from RPDs reported by the STBs. The disclosed example technical solutions utilize ordinary household data to train one or more machine learning algorithms (e.g., random forest, neural network, etc.) to predict the operating state of media devices connected to those STBs from features extracted from RPDs reported from the STBs. Ordinary household data refers to panel households that (i) are monitored by an AME using one or more meters and (ii) also have STBs that report RPDs received by the AME (e.g., directly or indirectly from service providers of the STBs). The audience measurement entity meter data obtained by the AME for the ordinary households produces a set of true viewing data that identifies media presented in each ordinary panel household, reflecting the operating state of the monitored media devices and STBs in those households over a monitoring period. The meter data for each ordinary household is then linked to RPDs from the same household to produce training RPDs that either have matching panel meter viewing data (which indicates that the media devices in that ordinary household were on) or do not have any matching meter viewing data (which indicates that the media devices in that ordinary household were off) (e.g., because the STB is reporting RPDs, but the panel meters did not report any corresponding viewing data). The training RPDs are used to train the machine learning algorithms to predict whether a given ordinary household’s training RPD has matching meter data (corresponding to a media device on state) or does not have matching meter data (corresponding to a media device off state). The disclosed example technical solutions then employ the trained machine learning algorithms to process RPDs reported from STBs to predict whether the media devices connected to that STB are on or off.
[0020] Figure 1 is a block diagram illustrating an example operating environment of a media device on / off detector using return path data to implement media device on / off detection in accordance with the teachings of this disclosure. Figure 1 The example operating environment 100 includes example user(s) 101, example media device(s) 102 associated with the user(s) 101, and an example set-top box (STB) 103. In the illustrated example, the user(s) 101 are not panel members of an AME. Figure 1The example operating environment 100 also includes one or more example team members 104, one or more example media devices 105 and one or more example set-top boxes (STBs) 106 associated with the one or more team members 104. The operating environment 100 also includes one or more example meters 107 to collect data from the one or more media devices 105 and / or STBs 106, the example network 108, one or more example media service providers 110, and the example audience measurement entity (AME) 120. The one or more example media service providers 110 includes an example return path data storage 112. The example audience measurement entity (AME) includes example meter data 122 and an example media device on / off detector 124.
[0021] One or more users 101 include any individual who accesses media content on one or more media devices 102 and is not associated with and / or registered in the AME 120 group (e.g., without one or more AME-based meters 107). One or more users 101 include individuals who subscribe to services provided by one or more media service providers 110 and use those services through their one or more media devices 102.
[0022] The media devices 102 associated with (one or more) non-group member users 101 may be fixed or portable computers, handheld computing devices, smartphones, internet devices, and / or any other type of device capable of displaying media from (one or more) media service providers 110. Figure 1 In the example shown, one or more media devices 102 may include, for example, a television, a tablet computer (e.g., iPad TM Motorola TM Xoom TM (etc.), desktop computers, cameras, internet-compatible TVs, smart TVs, etc. Figure 1 One or more media devices 102 are used to access (e.g., request, receive, present and / or broadcast) media, for example, provided by one or more media service providers 110 via example network 108.
[0023] An STB 103 associated with one or more media devices 102 may include, for example, an STB associated with a home entertainment system. The home entertainment system may receive media from one or more media service providers 110 and display the media on one or more media devices 102 (e.g., a television). STB data includes some or all of the data collected by a given STB 103, including tuning events and / or commands received by the STB 103 (e.g., power on, power off, channel change, input source change, start media presentation, pause media presentation, record media presentation, volume increase / decrease, etc.). STB data may additionally or alternatively include commands sent by the STB 103 to one or more media service providers 110 (e.g., switch input source, record media presentation, delete recorded media presentation, start media presentation time / date, end media presentation time, etc.), heartbeat signals, etc. STB data may include a home identifier (e.g., home ID) and / or an STB identifier (e.g., STB ID) for the STB 103.
[0024] One or more group members 104 include users who are part of the AME group residence, causing users to have media access and / or exposure to create media impressions (e.g., watching advertisements, movies, etc.). For example, one or more group members 104 may include users who provided their demographic information when registering with example AME 120. When one or more example group members 104 access media via example network 108 using example media device 105, AME 120 (e.g., AME server) stores group member activity data associated with the group members' demographic information (e.g., in group residence meter data 122) via one or more meters 107.
[0025] The media devices 105 associated with group member 104 may be fixed or portable computers, handheld computing devices, smartphones, internet devices, and / or any other type of device capable of displaying media from media service provider 110. Figure 1 In the example shown, one or more media devices 105 may include, for example, a television, a tablet computer (e.g., iPad TM Motorola TM Xoom TM (etc.), desktop computers, cameras, internet-compatible TVs, smart TVs, etc. Figure 1The media device(s) 105 are used to access (e.g., request, receive, present, and / or play out) media, such as provided by the media service provider(s) 110 over the example network 108. The media device(s) 105 can interact with the meter(s) 107 to provide viewing data (e.g., programs contacted by the panelist(s) using the media device(s) 105) to the AME 120.
[0026] The STB(s) 106 associated with the media device(s) 105 can include, for example, STBs associated with a residential entertainment system. The residential entertainment system can receive media from the media service provider(s) 110 and display the media on the media device(s) 105 (e.g., televisions, etc.). The STB data includes some or all data collected by a given STB 106, including tuning events and / or commands (e.g., power on, power off, change channel, change input source, start presenting media, pause presentation of media, record presentation of media, volume up / down, etc.) received by the STB 106. The STB data can additionally or alternatively include commands sent by the STB 106 to the media service provider(s) 110 (e.g., switch input source, record media presentation, delete recorded media presentation, time / date to start media presentation, time to complete media presentation, etc.). The STB data can include a household identification (e.g., household ID) and / or a STB identifier (e.g., STB ID) for the STB 106. The STB 106 can also interact with the meter(s) 107 to provide STB data (e.g., tuning data and / or viewing data) directly to the meter(s) 107.
[0027] The meter(s) 107 include hardware and / or software provided by the AME 120 when or after the panelist(s) 104 associated with the media device(s) 105 agree to be monitored. The meter(s) 107 can include, for example, a set-top box (STB) and / or a metering device (e.g., a smart TV, a smart phone, a tablet, a computer, etc.). The meter(s) 107 can be configured to collect data from the media device(s) 105 and / or the STB(s) 106 associated with the media device(s) 105. The meter(s) 107 can be configured to collect data from the media device(s) 105 and / or the STB(s) 106 associated with the media device(s) 105. Figure 1In the example of FIG. 1, meter(s) 107 collect monitoring information, such as media device-panelist interaction, content accessed on the media device, media device state, user selections, user inputs, location information, image information, etc. Periodically and / or aperiodically, meter(s) 107 transmit the monitoring information to an AME server (e.g., AME 120). Meter(s) 107 can also collect information from STB(s) 106, which can include tuning data and / or viewing data, in order to transmit such data to AME 120. In this context, assuming that meter(s) 107 can provide both media device-based data and STB-based data (e.g., using media device(s) 105 and STB(s) 106), panelist(s) 104 are part of a panel household, referred to herein as a common household, indicating that the panel household is not only monitored by AME 120 using meter(s) 107, but also includes STB(s) 106 that report return path data that is subsequently received by AME 120.
[0028] Network 108 can be implemented using any suitable wired and / or wireless network(s), including, for example, one or more wired provider networks, one or more satellite provider networks, one or more local area networks (LANs), one or more wireless LANs, one or more cellular networks, the Internet, etc. As used herein, the phrase "in communication," including variations thereof, encompasses direct communication and / or indirect communication, through one or more intermediary components, and does not require direct physical (e.g., wired) communication and / or continuous communication, but additionally includes selective communication at periodic intervals or aperiodic intervals and one-time events. An audience measurement entity (AME), such as The Nielsen Company (US), LLC, monitors viewing of media presented by such media devices.
[0029] The media service provider(s) 110 can include a cable television service provider, a satellite television service provider, a streaming media service provider, an OTP service provider, an Internet service provider, a content provider, and the like. The media service provider(s) 110 can include a database that stores return path data (e.g., return path data 112) received from the STB 106. For example, the return path data 112 can include any data receivable at the media service provider(s) 110 through a return path 110 from the media consumer site to the media service provider(s) 110. For example, the return path data 112 can include STB data from the STB(s) 103 and / or at least a portion of the STB data from the STB(s) 106. The return path data 112 can also include data from any other consumer that has network access capabilities (e.g., through a cellular network, the Internet, other public or private networks, and the like). For example, the return path data 112 can include any or all linear real-time data from the STB(s) 103 and / or the STB(s) 106, guide user data from a guide server, clickstream data, keystream data (e.g., any clicks on a remote control - volume, mute, and the like), interactive activity (e.g., video on demand), and any other data (e.g., data from an intermediary device). The return path data 112 can be received from the STB(s) 103 and / or the STB(s) 106 through the network 108 (e.g., through an exchange of digital software) and / or can be cloud-based (e.g., associated with a remote server DVR) data received from a cloud service (e.g., a return path data cloud service that collects, processes, and analyzes cloud-based data).
[0030] The AME 120, e.g., Nielsen (America) Inc., operates as an independent party to measure and / or verify audience measurement information related to media accessed by users. The AME 120 can have an agreement with a pay television provider company, e.g., the media service provider(s) 110, to obtain television tuning information, e.g., return path data 112, from the STB(s) 103 and / or the STB(s) 106 and / or other devices / software. This allows the AME 120 to augment panelist data, e.g., tuning and / or viewing data collected from the panelist(s) 104, with non-panelist data, e.g., tuning and / or viewing data collected from the STB(s) 106 associated with the user(s) 101. In some examples, the AME 120 utilizes common household data to enable the combination of return path data 112 with meter data 122. Common household data relates to a panel household, e.g., the household of the panelist(s) 104, that is monitored by the AME, e.g., the AME 120, using one or more meters, e.g., the meter(s) 107, and also has a STB, e.g., the STB(s) 106, that reports return path data, e.g., the return path data 112, received by the AME 120, e.g., directly or indirectly from the media service provider(s) 110 of the STB(s) 106.
[0031] The meter data 122 includes meter data for common households, e.g., households that have the AME panelist(s) 104 and provide return path data 112 to the media service provider(s) 110, obtained by the AME 120, and meter data obtained from households that include AME-based meters but do not include STBs. As such, the meter data 122 is collected from various meters, e.g., people meters, etc., that are used as audience measurement tools to measure, e.g., television and cable television audience viewing habits, e.g., of the panelist(s) 104. The meter data can include, for example, demographic information of media viewers, e.g., the panelist(s) 104, and their viewing status, e.g., media content being viewed by the panelist(s) 104. In Figure 2-3 In examples, the AME 120 can use the meter data 122 to produce a ground truth set of viewing data that identifies media presented in common panel households to reflect the operational status of the monitored media devices, e.g., the media device(s) 105, and STBs, e.g., the STB(s) 106, in these households, during a monitoring period using the media device on / off detector 124.
[0032] The media device on / off detector 124 links panel meter data 122 for each common household (e.g., a household with panelist(s) 104 that provide return path data 122 through STB(s) 106) to return path data 112 from the same household. The media device on / off detector 124 uses the linked information to create a return path data set for training a machine learning algorithm, as described in detail in the detailed description. For example, the media on / off detector creates a training return path data set that includes (1) matching panel meter viewing data 122 (e.g., indicating that a media device(s) 105 in a common household is on), and / or (2) no matching panel meter viewing data 122 (e.g., indicating that a media device(s) 105 is off), as described in detail in the detailed description. Figure 2 Figure 2 The training return path data generated by the media device on / off detector 124 trains a machine learning algorithm to predict whether return path data 112 has matching meter data 122, and thus the media device on / off detector 124 uses this information to determine whether a media device(s) 105 connected to STB(s) 106 is on or off. The media device on / off detector 124 can then use the trained algorithm to evaluate data received from STBs (e.g., STB(s) 103) that are not associated with a common household, and thus can determine whether a media device(s) 102 has been on or off during a particular viewing event (e.g., viewing a particular channel). This allows the AME 120 to use the media device on / off detector 124 to determine whether data for a particular viewing event is actually associated with a media device 102 being on, or whether a STB 103 was on when a media device 102 was off, in which case the viewing event was not a true viewing segment that can be used to obtain audience measurement data. In this way, the media device on / off detector 124 applies a machine learning algorithm to households that do not include a panelist but do have a user(s) 101 associated with a STB(s) 103 that provides return path data after being trained using common household data (e.g., data from panelist(s) 104 that provide both meter data from meter(s) 107 and return path data from STB(s) 106) that can be identified as reporting a true viewing event (e.g., a media device 102 is determined to be on) or a viewing event that is not a true viewing event (e.g., a media device 102 is determined to be off while a STB 103 remains on).
[0033] Figure 1 is a block diagram of an example implementation of a media device on / off detector 124. The media device on / off detector 124 includes an example data store 202, an example identifier 204, an example classifier 206, an example generator 208, an example trainer 210, and an example on / off determiner 212.
[0034] The data store 202 stores return path data 112 and metering data 122 for the media device(s) 105, as well as return path data associated with the media device(s) 102. For example, the data store 202 stores data retrieved from the media service provider(s) 110 (e.g., return path data 112) and data available to the AME 120 (e.g., panel metering data 122). For example, the data retrieved from the media service provider(s) 110 can include at least a portion of STB 103 and / or STB 106 data, and / or data from any other consumer device that has network access capabilities (e.g., through a cellular network, the Internet, other public or private networks, etc.). In some examples, the data can include linear real-time data from the STB(s) 103 and / or STB(s) 106, guide user data from a guide server, clickstream data, keystroke stream data (e.g., any clicks on a remote control - volume, mute, etc.), interactive activity (e.g., video on demand), and any other data (e.g., data from an intermediary device). The data stored by the data store 202 can include panel metering data 122, such as demographic information for media viewers (e.g., the panel member(s) 104) and their viewing status (e.g., media content being viewed by the panel member(s) 104). In some examples, the data of the data store 202 includes data retrieved for a common residence (e.g., a residence with the panel member(s) 104 that are both AME 120 panel members and have a STB 106 that provides return path data 112 to the media service provider(s) 110). In some examples, such data can include panel metering data 122 from a set meter (SM) and code reader (CR) meter, and / or data from a national meter (NPM) (e.g., an audience measurement entity reader or audience measurement entity meter 107). In such examples, return path data 112 and panel metering data 122 obtained from the common residence (e.g., a residence with the media device(s) 105) can be used to train a machine learning algorithm using the panel metering data (e.g., data from the meter(s) 107) as a ground truth set, such that the algorithm is trained to identify whether the media device(s) 105 are on or off. Once the algorithm is trained, it can be used to determine whether the media device(s) 102 are on or off (e.g., using return path data from the STB(s) 103 associated with a non-panel residence), thereby identifying true viewing events and / or viewing segments.From this, a machine learning algorithm trained on public housing data (e.g., from STB(s) 106 and meter(s) 107) can be used to infer the state (e.g., on / off) of media devices 102 associated with non-panel households (e.g., the household(s) of user(s) 101).
[0035] Non-panel household return path data thus supplements existing panel meter data 122 to increase sample size and representative panel base per market (e.g., increasing the number of households (HHs) that can be included in data reporting based on audience measurement). For example, adding return path data (RPD) 112 can reduce the number of quarter-hour zero ratings (QHs) in data based on AME 120 (e.g., reducing the number of times and networks that have no viewing data 122 based on panel members during a day). Data store 202 can be implemented by any storage device and / or disk for storing data, such as flash memory, magnetic media, optical media, etc. Further, the data stored in data store 202 can be in any data format, such as binary data, comma-delimited data, tab-delimited data, structured query language (SQL) structures, etc. While in the illustrated example, data store 202 is shown as a single database, data store 202 can be implemented by any number and / or type(s) of database(s).
[0036] The identifier 204 can access common household data from the data store 202 for one or more groups of common households (e.g., return path data 112 from STB(s) 106 that are also monitored by AME monitor(s) 107, and panel meter data 122 from meter(s) 107). In some examples, the identifier 204 groups the common household data into viewing segments (e.g., half-hour segments). The viewing segments can correspond to a particular viewing time (e.g., between 4 a.m. and 5 a.m. on Monday through Friday) when the panel member(s) 104 view media provided by media service provider(s) 110. In some examples, a common household group can include households located within a particular geographic area of interest (e.g., identified by the same zip code). The identifier 204 can group common households in any manner that is of interest for evaluating data related to improving market coverage and individual audience estimates (e.g., improving presentation of local markets). In some examples, the identifier 204 also identifies additional data available from the panel meter data 122, such as audience data for tuning events, household characteristics, and composition data derived from household tuning (e.g., by STB 106), third party (e.g., media service provider(s) 110) data, and known panel information (e.g., meter data 122). In some examples, the identifier 204 compares the panel meter data 122 and the return path data 112 tuned for each common household. In some examples, such a comparison can include a minute-by-minute comparison of each set of data (e.g., RPD 112 and panel meter data 122) tuned for each common household.
[0037] The classifier 206 classifies the viewing segments (e.g., one-half hour segments) identified using the identifier 204 to determine labeled viewing segments based on whether the RPD 112 for each viewing segment has matching panel meter data 122. For example, the classifier 206 can classify viewing segments as "matched" or "otherwise" to determine labeled viewing segments. In such examples, a given viewing segment can be classified as "matched" if the RPD 112 data (e.g., tuning data) for the viewing segment has matching panel meter data 122 (e.g., viewing data) for the given viewing segment. For example, a match can occur when it is determined that the same tuning event occurred for both the RPD 112 and the panel meter data 122 (e.g., return path data 112 from the STB 106 indicates that a particular channel was tuned for a total of 3 hours, and the panel meter data 122 from the meter 107 confirms that the channel was actually active and media was presented at the panelist's 104 site for the full 3 hours). In some examples, the classifier 206 classifies a viewing segment as "otherwise" if the RPD 112 tuning data in the viewing segment does not have matching viewing data from the panel meter data 122. In some examples, the classifier 206 classifies some viewing segments as partially "matched" or partially "otherwise." For example, return path data 112 can indicate that a channel was tuned for 3 hours, but the panel meter data 122 indicates that the channel was active and media was presented at the panelist's 104 site for 1.5 hours of the 3 hours reported by the return path data 112 of the STB 106, such that some viewing segments are classified as "matched" when the panel meter data 122 corresponds to the return path data 112, and other viewing segments are classified as "otherwise" when the panel meter data 122 does not correspond to the return path data 112. In some examples, the classifier 206 classifies partially "matched" and partially "otherwise" viewing segments as "matched" viewing segments. For example, partially "matched" and / or partially "otherwise" viewing segments (e.g., 30 minute long viewing segments) can be classified as "matched" using both return path data 112 and meter data 122 when a majority of the viewing segment (e.g., at or above a first threshold) is "matched" (e.g., 20 minutes of a 30 minute long viewing segment). In some examples, partially "otherwise" and / or partially "matched" viewing segments (e.g., 30 minute long viewing segments) can be classified as "otherwise" when a majority of the viewing segment (e.g., at or below a second threshold) does not include a match between the return path data 112 and the meter data 122 (e.g., 20 minutes of a 30 minute long viewing segment). In some examples, the classifier 206 classifies partially "matched" and partially "otherwise" viewing segments as "otherwise" viewing segments.
[0038] The generator 208 generates features from the labeled viewing segments (e.g., “matched” and / or “additional” viewing segments). For example, the generator 208 can generate features from the labeled viewing segments of the public residential data to create training data for training a machine learning algorithm using the training data. For example, the features generated by the generator 208 using the labeled viewing segments can include, but are not limited to: day of the month, viewing segment index (e.g., “viewing segment index” corresponding to where the viewing segment occurred in relation to the event), viewing segment duration (e.g., “viewing segment duration”) corresponding to the length of time of the given viewing segment, event duration corresponding to the length of time of viewing the particular media content, number of minutes since the event started, day of the week, weekday / weekend, STB model type, time zone, event type (e.g., live viewing, time-shifted viewing (TSV), etc.), average event duration for the particular household on the particular day, number of events for the household in the day, number of viewing segments for the household in the day, average event duration for the particular device on the particular day, number of events for the device in the day, number of viewing segments for the device in the day, ratio of event duration to device average event duration for the day, etc. In some examples, the generator 208 generates other types of features, such as features specified by a user-based configuration or input or as specified by a machine learning algorithm based on the training data.
[0039] The trainer 210 trains a machine learning algorithm included in the on / off determiner 212 based on the features generated by the generator 208 that form the training data. For example, the trainer 210 uses the training data to iteratively train and tune a machine learning algorithm, which can be, for example, a neural network. In some examples, the machine learning algorithm can be a random forest or random decision forest learning method (e.g., a supervised classification algorithm). For example, using a random forest learning method allows for inputting a training data set with targets and features into a decision tree, allowing the algorithm to formulate a set of rules that in turn are used to form a prediction. Also, using a random forest learning method allows for inputting data that can be missing values. In some examples, a random forest classification algorithm can be used as the selected machine learning algorithm in order to capture the non-linear behavior of the training data and due to its ability to classify based on a wide range of parameter settings. For example, the trainer 210 can use a random forest learning method to estimate the probability that an observation falls into a given class. In some examples, the trainer 210 can use a random forest classifier to train the data (e.g., using a collection of trees grown randomly, whose final prediction is an aggregation of the predictions from the individual trees). In some examples, once the trainer 210 fits a classification random forest to the training data, the conditional class probabilities for a test point can be inferred by computing the proportion of “trees” in the “forest” that voted for a certain class. When two classifiers in a set are highly correlated, the estimated probabilities converge to 0 or 1.
[0040] In some examples, the trainer 210 utilizes one or more threshold values to convert probability values output from the machine learning algorithm into a “matched” or “otherwise” classification, where the threshold(s) are adjusted to meet one or more performance goals. For example, it is important to select a probability threshold (e.g., p-value) that actually classifies a category as “matched” or “otherwise,” and it can not always be the case that a p-value of 0.5 is the default. In some examples, an adjusted probability threshold (e.g., p-value = x) can be used to reclassify those probability values greater than x as “matched” and those less than x as “otherwise” (e.g., a threshold adjusted based on whether the machine learning algorithm correctly identifies a media device as “otherwise” or “matched”). For example, a probability value of 0.995 returned by a machine learning algorithm (e.g., random forest) predicts that a data set is likely to be “matched” (e.g., all RPD 112 data (e.g., tuning data) in a viewing segment has matching panel meter data 122). Conversely, a probability value of 0.004 predicts that a data set is likely to be “otherwise” (e.g., all RPD 112 data (e.g., tuning data) in a viewing segment does not have matching panel meter data 122). However, a prediction value of 0.6 is not clearly either “matched” or “otherwise.” Thus, a probability threshold is defined to determine that a probability value above a certain threshold x indicates that a data set is “matched,” while a probability value below a certain threshold x indicates that the data set is “otherwise.” This allows for the use of data sets that can include missing values or lack of features, as the final probability value will be compared to the threshold probability value to determine whether a given data set is “matched” or “otherwise.” In some examples, the threshold is selected to ensure that the model post-RPD tuning can be comparable to national meter (NPM) tuning. Once the model is trained and a classification probability threshold is selected, the model can be applied to a full RPD set (e.g., RPD 112). For example, a full RPD set includes RPD 112 derived from STB(s) 103 that are not associated with a panel home (e.g., user(s) 101 are not AME panel members). By training a machine learning algorithm to identify when a media device is turned on or off based on public home data (e.g., meter(s) 107 data and STB(s) 106 return path data), the algorithm can be applied to RPD 112 data to determine whether a media device(s) 102 associated with a user(s) 101 that is not a panel member is turned on or off based on the provided return path data 112 associated with STB(s) 103.From this, data evaluation of the viewing segments can be performed, for example, using the complete RPD set, which includes not only the common household data associated with the STB(s) 106, but also the non-panel household data associated with the STB(s) 103.
[0041] The on / off determiner 212 determines whether the media device associated with the reported return path data is on or off. For example, once the trainer 210 has trained the machine learning algorithm as described above, the reported return path data (e.g., new return path data provided by the media service provider(s) 110 with which the AME 120 is cooperating) is applied to the trained machine learning algorithm. The algorithm predicts a“matched” or“otherwise” classification for each viewing segment and each RPD household represented by the reported RPD (e.g., RPD 112 from STB(s) 103 in a non-panel member user 101 household), which translates to predicting whether each viewing segment of each non-panel member RPD household is associated with an on or off media device (e.g., one or more media devices 102). For example, a“matched” classification would indicate that the media device is on, while an“otherwise” classification would indicate that the media device is off (e.g., STB 103 reports RPD 112 that indicates that media content was viewed on the media device(s) 102, but using the trained algorithm, the on / off determiner 212 can determine that the media device 102 was off during the length of time that the RPD 112 reports that media content was viewed, thereby removing that viewing event from a true viewing event). As such, the trainer 210 optimizes the algorithm to predict matched RPD 112 and panel data 122 (e.g., corresponding to the media device 105 being on) or otherwise RPD 112 data (e.g., corresponding to the media device 105 being off). For example, the algorithm can take RPD 112 as input, and once it has been trained to recognize the differences between RPDs that correspond to the media device on / off state, it can output a prediction based on the RPD 112. In some examples, the common household RPD 112 input to the algorithm results in an output such that the trainer 210 compares the prediction generated by the algorithm to the corresponding common household panel data (e.g., from meter(s) 107) such that the trainer 210 can train the algorithm to achieve a desired level of accuracy in predicting whether the media device (e.g., media device 105) is on or off. Thus, when the on / off determiner 212 receives RPD 112 (e.g., provided by STB(s) 103) from a non-panel household, the on / off determiner outputs a“matched” or“otherwise” prediction such that“matched” corresponds to the media device 102 being on and“otherwise” corresponds to the media device 102 being off. In some examples, the on / off determiner identifies the prediction based on features that the trained algorithm was taught to recognize as being associated with matched panel data.Assuming that RPD 112 may not directly provide information about the media devices (e.g., media devices 102) connected to STB 103 reporting RPD 112, this allows for improved accuracy in determining the exposure associated with media output from the STB (e.g., the on / off operating status of the media devices connected to STB 103). For example, media devices 102 connected to STB 103 may be turned off, while STB 103 remains unintentionally or intentionally powered on and outputting media through media devices 102. In some examples, once the classification performed by the algorithm based on the RPD tuning training dataset provided by RPD 112 is correlated with the tuning provided by the National Meter (NPM) (e.g., by establishing a classification threshold that ensures the RPD tuning data is comparable to data obtained using NPM), the on / off determiner 212 uses a machine learning algorithm trained using trainer 210.
[0042] Although Figure 2 and Figure 1 The example shown is how to implement the media device on / off detector 124, but Figure 2 and Figures 1-2 One or more elements, methods, and / or devices shown may be combined, divided, rearranged, omitted, eliminated, and / or implemented in any other way. Furthermore, Figures 1-2 The example data storage 202, example recognizer 204, example classifier 206, example generator 208, example trainer 210, example on / off determiner 212, and / or more generally, the example media device on / off detector 124 can be implemented by hardware, software, firmware, and / or any combination of hardware, software, and / or firmware. Therefore, for example, Figure 1The example data store 202, the example identifier 204, the example classifier 206, the example generator 208, the example trainer 210, the example on / off determiner 212, and / or, more generally, any of the example media device on / off detector 124 can be implemented by one or more analog or digital circuits, logic circuits, programmable processor(s), programmable controller(s), graphics processing unit(s) (GPU(s)), digital signal processor(s) (DSP(s)), application specific integrated circuit(s) (ASIC(s)), programmable logic device(s) (PLD(s)), and / or field programmable logic device(s) (FPLD(s)). When reading any claims into the present patent, any apparatus or system claims encompass both software and / or firmware implementations of at least one of the example data store 202, the example identifier 204, the example classifier 206, the example generator 208, the example trainer 210, and / or the example on / off determiner 212, where the software and / or firmware implementations are defined as including non-transitory computer readable storage devices or storage disks, such as memory, digital versatile disks (DVDs), compact discs (CDs), Blu-ray discs, etc., that include software and / or firmware. Moreover, the example media device on / off detector 124 can include one or more elements, methods, and / or devices (in addition to or instead of Figure 2 and Figures 3-4 the elements, methods, and / or devices shown), and / or can include more than one or all of any of the elements, methods, and devices shown. As used herein, the phrase "in communication," including variations thereof, encompasses direct communication and / or indirect communication, through one or more intermediary components, and does not require direct physical (for example, wired) communication and / or continuous communication, but additionally includes selective communication at periodic intervals, scheduled intervals, non-periodic intervals, and / or one-time events.
[0043] Figure 7 Flow diagrams representative of example hardware logic, machine readable instructions, hardware implemented state machines, and / or any combination thereof for implementing the example solution techniques disclosed herein are shown in FIGS. 1-3. In this example, the machine readable instructions can be one or more executable programs or portions of an executable program for execution by a computer processor such as the processor 712 shown in the example processor platform 700 discussed below in connection with FIG. 7. The one or more programs or portions thereof can implement the example data store 202, the example identifier 204, the example classifier 206, the example generator 208, the example trainer 210, the example on / off determiner 212, and / or any of the example media device on / off detector 124. Figures 3-4 The one or more programs or portions thereof can be embodied in a computer program product, which can include one or more computer program storage media. The one or more computer program storage media can be configured to store data, which can be accessed by the computer processor. The one or more computer program storage media can include volatile and / or non-volatile memory. The one or more computer program storage media can include database(s) or other data storage TMor memory associated with processor 712), but the entire program(s) and / or portions thereof can alternatively be executed by a device other than processor 712 and / or embodied in firmware or dedicated hardware. Also, although the example procedures are described as being implemented in the context of a single processor, it should be appreciated that the procedures can instead be implemented in the context of multiple processors. Figures 3-4 Figures 3-4 described, and / or some of the described blocks can be combined, eliminated or modified, and / or some blocks can be subdivided into multiple blocks. Additionally or alternatively, any or all of the blocks can be implemented by one or more hardware circuits (e.g., discrete and / or integrated analog and / or digital circuitry, FPGAs, ASICs, comparators, operational amplifiers (op-amps), logic circuits, etc.) structured to perform the corresponding operations without running software or firmware.
[0044] The machine-readable instructions described herein can be stored in one or more of a compressed format, an encrypted format, a dispersed format, a packaged format, etc. The machine-readable instructions as described herein can be stored as data (e.g., portions of instructions, code, representations of code, etc.) that can be used to create, manufacture, and / or produce machine-executable instructions. For example, the machine-readable instructions can be dispersed and stored on one or more storage devices and / or computing devices (e.g., servers). The machine-readable instructions can require one or more of installation, modification, adaptation, updating, combination, supplementation, configuration, decryption, decompression, unpacking, distribution, redistribution, etc. to make it directly readable and / or executable by a computing device and / or other machine. For example, the machine-readable instructions can be stored in multiple portions that are individually compressed, encrypted, and stored on separate computing devices, where the portions, when decrypted, decompressed, and combined, form a set of executable instructions that implement a program as described herein. In another example, the machine-readable instructions can be stored in a state that is readable by a computer, but require the addition of a library (e.g., a dynamic link library), a software development kit (SDK), an application programming interface (API), etc. to execute the instructions on a particular computing device or other device. In another example, the machine-readable instructions can require configuration (e.g., stored settings, data input, recorded network addresses, etc.) before the machine-readable instructions and / or the corresponding program(s) can be executed in whole or in part. Accordingly, the disclosed machine-readable media and / or corresponding program(s) are intended to encompass such machine-readable instructions and / or programs regardless of the particular format or state of the machine-readable instructions and / or programs when stored or when in transit.
[0045] As described above, Figure 3 The example process(es) can be implemented using executable instructions stored on non-transitory computer and / or machine readable media such as a hard disk drive, flash memory, read-only memory, optical disc, digital versatile disc, cache, random access memory, and / or any other storage devices or storage disks in which information is stored for any duration (e.g., for extended period of time, permanently, for brief instances, temporarily, and / or for caching and / or buffering). As used herein, the term non-transitory computer readable and / or machine readable storage medium is expressly defined to include any type of computer and / or machine readable storage device and / or storage disk and to exclude propagating signals and propagating media.
[0046] "Comprising" and "including" (and all their forms and tenses) are used herein as open-ended terms. Therefore, whenever a claim uses any form of "comprising" or "including" (e.g., including (comprises, include, comprising, including), having, etc.) as a preamble, or in any kind of claim statement, it should be understood that additional elements, terms, etc., may be present without exceeding the scope of the corresponding claim or statement. As used herein, when the phrase "at least" is used as a transitional term, for example, in the preamble of a claim, it is open-ended in the same way as the terms "comprising" and "including". For example, when the term "and / or" is used in the form of A, B, and / or C, it refers to any combination or subset of A, B, C, such as (1) A alone, (2) B alone, (3) C alone, (4) A with B, (5) A with C, (6) B with C, and (7) A with B with C. As used herein in the context of describing structures, components, items, objects, and / or things, the phrase "at least one of A and B" means an implementation including (1) at least one A, (2) at least one B, and (3) at least one A and at least one B. Similarly, as used herein in the context of describing structures, components, items, objects, and / or things, the phrase "at least one of A or B" means an implementation including (1) at least one A, (2) at least one B, and (3) at least one A and at least one B. As used herein in the context of describing the execution or operation of processes, instructions, actions, activities, and / or steps, the phrase "at least one of A and B" means an implementation including (1) at least one A, (2) at least one B, and (3) at least one A and at least one B. Similarly, as used herein in the context of describing the execution or operation of processes, instructions, actions, activities and / or steps, the phrase “at least one of A or B” means an implementation including (1) at least one A, (2) at least one B, and (3) any one of at least one A and at least one B.
[0047] Figure 3 This is a flowchart 300 illustrating example computer-readable instructions executable according to the teachings of the present invention to perform media device on / off detection using return path data. Referring to the preceding figures and the associated written description, Figure 2 Example program 300 begins execution in box 305, where, Figure 4The identifier 204 accesses the RPDs 112 and corresponding panel meter data 122 for a set of common households, collectively referred to as common household data. For example, the identifier 204 can access the data store 202 to obtain minute-level RPD 112 tuning data and corresponding panel meter 122 viewing data for a set of common households. At block 310, the identifier 204 groups the common household data into one-hour segments, referred to as viewing segments. Thus, at block 310, the identifier 204 divides the minute-level RPD tuning for a given common household into a plurality of one-hour viewing segments and links to corresponding panel meter viewing data and viewing segments for that common household. At block 315, the classifier 206 classifies the viewing segments as “matched” or “otherwise” to determine labeled viewing segments for the common household data. In this example, the classifier 206 classifies a viewing segment as “otherwise” if none of the RPD 112 tuning data in the viewing segment has a match with the panel meter 122 viewing data for that viewing segment. Conversely, the classifier 206 classifies a viewing segment as “matched” if all of the RPD 112 tuning data in the viewing segment has a match with the panel meter 122 viewing data for that viewing segment. In some examples, the classifier 206 can group the common household data into viewing segments of 15-minute duration, which results in almost all viewing segments being classified as “matched” or “otherwise” with relatively few viewing segments being partially “matched” or partially “otherwise.” In some examples, at block 315, the classifier 206 classifies viewing segments that are partially “matched” or partially “otherwise” as “matched.”
[0048] At block 320, the generator 208 generates features (e.g., based on RPD 112 tuning data and possibly other available RPD included in the labeled viewing segments) from the labeled viewing segments to determine training data to be used to train a machine learning algorithm (e.g., random forest, neural network, etc.) to predict whether an input viewing segment of RPD 112 tuning data is likely to be classified as "matched" (and thus likely associated with an open media device) or possibly classified as "otherwise" (and thus likely associated with a closed media device). Example features generated by the generator 208 from the labeled viewing segments include, but are not limited to: day of the month, viewing segment index, viewing segment duration, event duration, number of minutes since the start of the event, event type, number of events for the household in a day, etc. Other features that can be generated from the labeled viewing segments include, but are not limited to: household id, device id, event type (live, dvr, etc.), play delay, station code, etc. In some examples, feature selection is based on an evaluation of the percentage of "matched" viewing segments and the percentage of "otherwise" viewing segments removed such that if the use of certain features results in overfitting (e.g., the training data is modeled too well such that when a new data set is applied, the model's learning of the details and noise in the training data set negatively impacts the performance of the model), those features can not be included. In some examples, the model can be trained on at least a month (or some other monitoring interval) of common household data before applying the model to reported return path data from a given common household to determine media display on / off status. In some examples, the model can be retrained and tested monthly (or at some other rate).
[0049] At block 325, trainer 210 uses the training data generated at block 320 to iteratively train and adjust a machine learning algorithm (e.g., random forest, neural network, etc.) implemented by on / off determiner 212. The machine learning algorithm outputs a prediction classifying an input viewing segment of RPD 112 tuning data as one of 2 labels, namely “matched” (corresponding to a determination that a media device 105 associated with the input viewing segment of RPD 112 tuning data is on) or “otherwise” (corresponding to a determination that a media device 104 associated with the input viewing segment of RPD 112 tuning data is off). In some examples, trainer 210 converts probability values output from the machine learning algorithm to “matched” or “otherwise” classifications using one or more thresholds, where the threshold(s) are adjusted to meet one or more performance targets. At block 330, trainer 210 applies reported RPDs from RPD households (e.g., not public households) to the trained machine learning algorithm (e.g., random forest, neural network, etc.), which predicts a “matched” or “otherwise” classification for each viewing segment and each RPD household represented by the reported RPDs, which translates to predicting whether each viewing segment of each RPD household is associated with an on or off media device.
[0050] Figure 4This is a flowchart 325 representing example computer-readable instructions that can be executed by the media device on / off detector 124 to train a machine learning algorithm using training data based on return paths. Using public housing return path data and group meter data accessed by the recognizer 204 (represented by box 405) and features generated from the return path data by the generator 208 (represented by box 410), the trainer 210 trains a model such that the model outputs a prediction (e.g., the digits 0-1) (box 415). The trainer 210 uses a classification threshold to classify the model predictions as “other” or “matching”. In some examples, the classification threshold is selected to allow the trainer 210 to train the algorithm such that the final “matching” prediction is associated with NPM public housing tuning data that can be used as a reference during the training process (box 420). In some examples, the classification threshold can be selected such that the model’s post-RPD public housing tuning is no greater than 20% when compared to the NPM public housing tuning used as a reference. The trainer 210 uses the classification threshold to identify the model predictions as “matching” or “other” (box 425). For example, if the model used is based on a random forest machine learning algorithm, a probability value of 0.995 (or some other relatively high probability value) returned by the model is likely a “matching” viewing segment, while a probability value of 0.004 (or some other relatively low probability value) predicts a dataset that is likely “alternative” (e.g., none of the RPD 112 data (e.g., tuning data) in the viewing segment has a matching group meter data 122). However, a prediction value of 0.6 (or some other probability value relatively close to 0.5) is not clearly a “matching” or “alternative” one, so a threshold is needed to determine how such probability values should be classified. In some examples, the classification threshold can be adjusted based on whether the “matching” and / or “alternative” predictions are correct compared to the group meter data 122 obtained from group dwellings as part of the public housing data used during algorithm training (box 428). Establishing a classification threshold enables the algorithm to predict whether media devices are on or off with high accuracy (e.g., the algorithm accurately identifies the media device status). If the model prediction is classified as "matched" (box 430) based on the model output, then trainer 210 identifies the media display status as "on" (box 435). If the model prediction is classified as "other" (box 440), then trainer 210 identifies the media display status as "off" (box 445). In some examples, the training data output generated by the model that identifies whether the media device is on or off is compared with the NPM data to determine whether the classification threshold should be adjusted (e.g., ensuring that the post-model RPD tuning is no more than 20% larger than the NPM tuning, while also minimizing the amount of matched tuning removed). Once based on Figures 5A-5BThe example instruction completion training period, the media device on / off detector 124 determines the media display on / off state for a given RPD residence using the on / off determiner 212 based on the trained machine learning algorithm.
[0051] Figure 5A Example validation metrics including the use of the techniques described herein to indicate media device on / off state determinations use public residence return path data and panel meter data with improved accuracy compared to reference on / off determination techniques. In Figure 5A In the example table 500, when comparing a non-machine learning algorithm training data set (e.g., prod (product)) to a machine learning algorithm training data set (e.g., new (new)), example amounts of “additional” tuning percentages (%) and removed “matched” tuning percentages (%) of removal are shown for three broadcast service providers. For example, return path data (e.g., return path data 112) captures STB tuning (e.g., STB 106), but does not reveal when a media device (e.g., a television) is on or off. Thus, modeling on / off times using a machine learning algorithm training data set can ensure that tuning is not exaggerated. In the example of table 500, for each of the three example broadcast service providers (e.g., 510, 520, and 530), the removed “additional” (e.g., television off) tuning percentage is evaluated in comparison to the removed “matched” (e.g., television on) tuning percentage. For broadcast service provider 510, the trained model results in a greater removed “additional” tuning percentage (e.g., 52% compared to 43%), and a reduced removed “matched” tuning percentage (e.g., 15% compared to 21%). In some examples, the removed “additional” tuning percentage can not increase significantly, but the removed “matched” tuning percentage can decrease significantly. For example, for broadcast service provider 520, using the trained model results in a slight increase in removed “additional” tuning (e.g., 71% compared to 70%), and a general decrease in removed “matched” tuning percentage (e.g., 12% compared to 25%). Likewise, in the example of broadcast service provider 530, the removed “additional” tuning is slightly reduced (e.g., from 78% to 75%), but the removed “matched” tuning is significantly reduced (e.g., from 36% to 18%). Thus, in some examples, the trained model enables a greater percentage of “additional” tuning to be removed (e.g., new (new) on / off model for broadcast service provider 510), such that a greater number of public residences identified as having an off media (e.g., designated as “additional” tuning) are removed from the overall tuning data, while in other examples, the trained model enables a greater percentage of “matched” tuning to be removed (e.g., prod (product) on / off model for broadcast service provider 520), such that a greater number of public residences identified as having an on media (e.g., designated as “matched” tuning) are removed from the overall tuning data. Figure 5BIn the presented example (e.g., the new on / off model for broadcast service providers 510, 520, and 530), the removed "matched" tuning percentage is reduced, indicating that more data can be included in the tuning count because the "matched" data indicates that the return path data and the panel meter data for the common household confirm that the media device (e.g., television) is on, thus allowing the tuning data to be included in the total count. This allows the tuning data to be more accurate, more reliable, and representative of the use of the media devices in the common household.
[0052] In Figures 6A-6BIn the example table 550, a previous on / off model (e.g., without using a machine learning based training algorithm) and an on / off model using a machine learning based training algorithm described herein (e.g., designated as a new on / off model) are compared to general measures (e.g., established audience measurement methods, such as measures from set meters and / or code meters, but not including RPD). For example, training for an on / off model is consistent with the methods disclosed herein, aiming to improve the accuracy of the model, which can be measured against a reference (e.g., a general reference determined using National Panel Meter (NPM) panel data). For example, model results can be compared to data obtained using set meters and / or code readers to obtain household ratings 555 and person-based ratings 585 (e.g., specific to demographic data, such as 18-24 year old people, 25-54 year old people, and 55 and older people). In the example table 550, comparisons are made between general measures 560 and either a previous on / off model 565 or a new on / off model 570, resulting in a previous model versus general measure comparison 575 and a new model versus general measure comparison 580. For set meter data and code reader data, the new on / off model (e.g., using a machine learning based training algorithm) improves the accuracy (e.g., based on data science validation and analysis) of the results for household ratings 555. For example, the percentage difference between tuned data when comparing the new model to general measures (e.g., 0.2% and -0.7%) is lower than the percentage difference between tuned data when comparing the previous model to general measures (e.g., -0.6% and -1.6%) compared to general measures obtained using set meters and code readers. For person-based ratings 585, the comparison data for set meters using a non-machine learning training model (e.g., -3.4%, -3.3%, and -0.6%) is higher than the difference percentage for general measures compared to the algorithm training model (e.g., -2.9%, -2.7%, and 0%). Likewise, the comparison data for code meters using a non-machine learning training model is also higher than the difference percentage for general measures compared to the algorithm training model (e.g., -3.2%, -5.9%, and -1.1%) compared to the algorithm training model (e.g., -2.7%, -5.1%, and -0.4%). In some examples, data accuracy can increase according to other variables, such as the frequency of channel changes (e.g., measures of household ratings for 3 hours or more without a channel change can be more accurate).
[0053] Figure 6A Examples including changes in tuned minutes and percentage of remaining tuned minutes when using a machine learning based training algorithm described herein based on public residential return path data and panel meter data. Figure 6BA chart 600 of tuning minutes 620 recorded for a household in a given set of quarter-hour 630 data includes data using a prior on / off model 610 (e.g., a model that does not include return path data), a product on / off model 615 (e.g., a model that does not include a machine learning based training algorithm), and a new on / off model 605 (e.g., a model that includes return path data and uses training of a machine learning based algorithm). The number of tuning minutes counted per quarter hour is much higher using the prior on / off model 610 (e.g., that does not include return path data) compared to the product on / off model 615 and the new on / off model 605. Overall, the number of tuning minutes counted using a machine learning trained model that utilizes return path data is higher compared to when using return path data without additional training (e.g., product on / off model 615), except for the early morning time period where the tuning minute readings are nearly identical. The training aspect of the algorithm of the new on / off model allows for improved accuracy of the tuning minute count without eliminating minutes that should otherwise be included in the quarter hour tuning minute count. As further shown by chart 650, the remaining tuning minute percentage 655 for the on / off model 615 that does not include machine learning based training is lower compared to the on / off model 605 that includes such training based on return path data and panel meter data. For example, using the new on / off model 605, the remaining tuning minute percentage is higher such that more of the original tuning minute numbers provided by the reference evaluation (e.g., national meter data) are used compared to the number of minutes available when using the untrained model 615. Overall, the impact of the on / off model can vary by broadcast service provider and can need to be evaluated individually for each provider to determine the degree of applicability of a given model. Figure 7 As further shown by chart 650, the remaining tuning minute percentage 655 for the on / off model 615 that does not include machine learning based training is lower compared to the on / off model 605 that includes such training based on return path data and panel meter data. For example, using the new on / off model 605, the remaining tuning minute percentage is higher such that more of the original tuning minute numbers provided by the reference evaluation (e.g., national meter data) are used compared to the number of minutes available when using the untrained model 615. Overall, the impact of the on / off model can vary by broadcast service provider and can need to be evaluated individually for each provider to determine the degree of applicability of a given model.
[0054] Figures 3-4 are configured to execute Figures 1-2 are configured to execute Figure 2 is a block diagram of an example processor platform of an example media device on / off detector 124 of TM , for example, a server, a personal computer, a workstation, a self-learning machine (e.g., a neural network), a mobile device (e.g., a cell phone, a smart phone, a tablet computer such as an iPad
[0055] The processor platform 700 of the illustrated example includes a processor 712. The processor 712 of the illustrated example is hardware. For example, the processor 712 can be implemented by one or more integrated circuits, logic circuits, microprocessors, GPUs, DSPs, or controllers from any desired family or manufacturer. The hardware processor 712 can be a semiconductor (e.g., silicon) based device. In this example, the processor 712 implementsFigures 3-4 Example recognizer 204, example classifier 206, example generator 208, example trainer 210, and / or example on / off determiner 212 of the illustrated example.
[0056] The processor 712 of the illustrated example includes local memory 713 (e.g., a cache). The processor 712 of the illustrated example, can, through a link 718, be further connected to a main memory including volatile memory 714 and non-volatile memory 716. The link 718 can be implemented by at least one of a bus, one or more point-to-point connections, etc., or a combination thereof. The volatile memory 714 can be implemented by synchronous dynamic random access memory (SDRAM), dynamic random access memory (DRAM), RAMBUS dynamic random access memory (RAMBUS DRAM), and / or any other type of random access memory device. The non-volatile memory 716 can be implemented by flash memory and / or any other desired type of memory device. Access to the main memory 714, 716 is controlled by a memory controller. Dynamic random access memory and / or any other type of random access memory device. The non-volatile memory 716 can be implemented by flash memory and / or any other desired type of memory device. Access to the main memory 714, 716 is controlled by a memory controller.
[0057] The processor platform 700 of the illustrated example also includes interface circuitry 720. The interface circuitry 720 can be implemented by any type of interface standard, such as an Ethernet interface, a universal serial bus (USB), a Bluetooth® interface, a near field communication (NFC) interface, and / or a PCI-express (serial bus) interface. Interface, near field communication (NFC) interface, and / or PCI-express (serial bus) interface.
[0058] In the illustrated example, one or more input devices 722 are connected to the interface circuitry 720. The input device(s) 722 allow a user to input data and / or commands to the processor 712. The input device(s) can be implemented by, for example, an audio sensor, a microphone, a camera (still or video), a keyboard, buttons, a mouse, a touchscreen, a trackpad, a trackball, a trackstrip (such as a point-of-impulse), a voice recognition system, and / or any other human interface. In addition, many systems, such as the processor platform 700, can allow a user to control the computer system and provide data to the computer using physical gestures (such as, but not limited to, hand or body movements, facial expressions, and facial recognition).
[0059] One or more output devices 724 are also connected to the interface circuitry 720 of the illustrated example. The output devices 724 can be implemented by, for example, a display device (e.g., a light emitting diode (LED), an organic light emitting diode (OLED), a liquid crystal display (LCD), a cathode ray tube display (CRT), an in-place switching (IPS) display, a touchscreen, etc.), a haptic output device, a printer, and / or speaker(s). Thus, the interface circuitry 720 of the illustrated example typically includes a graphics driver card, a graphics driver chip, and / or a graphics driver processor.
[0060] The interface circuit 720 of the illustrated example also includes a communication device such as a transmitter, a receiver, a transceiver, a modem, a home gateway, a wireless access point, and / or a network interface to facilitate exchange of data with external machines (e.g., computing devices of any kind) via network 726. Communication can be facilitated via, for example, Ethernet
[0061] The processor platform 700 of the illustrated example also includes one or more mass storage devices 728 for storing software and / or data. Examples of such mass storage devices 728 include floppy disk drives, hard drive disks, compact disk drives, Blu-ray disk drives, redundant array of independent disks (RAID) systems, and digital versatile disk (DVD) drives.
[0062] Machine-executable instructions 732 corresponding to the software components of the example of FIG. 7 can be stored in mass storage device 728, volatile memory 714, non-volatile memory 716, local memory 713, and / or removable storage media (e.g., CD or DVD) 736.
[0063] From the foregoing, it will be appreciated that the example systems, methods, and apparatuses allow for predicting on / off operating states of media devices connected to set-top boxes (STBs) based on return path data (RPD) reported by the STBs. The disclosed example technical solutions utilize public housing data to train one or more machine learning algorithms, such as random forests, neural networks, etc., to predict operating states of media devices connected to STBs based on features extracted from RPD reported from the STBs. Meter data for each public housing is linked to RPD from the same housing to produce training RPD with matching meter viewing data (e.g., media device viewing data) that indicates a media device in the public housing is on or without any matching meter viewing data that indicates a media device in the public housing is off. In the examples disclosed herein, the training RPD is used to train a machine learning algorithm to predict whether a given public housing’s training RPD has matching meter data (corresponding to a media device on state) or does not have matching meter data (corresponding to a media device off state). The disclosed example technical solutions then employ the trained machine learning algorithm to process RPD reported from STBs to predict whether a media device connected to the STB is on or off.
[0064] Although certain example methods, apparatus and articles of manufacture have been disclosed herein, the scope of coverage of this patent can not be so limited. Rather, this patent covers all methods, apparatus and articles of manufacture falling within the scope of the claims.
[0065] The appended claims are hereby incorporated into this DETAILED DESCRIPTION by reference, each claim as if individually incorporated by reference herein.
Claims
1. A method of performing media device on / off detection using return path data, the method comprising: obtaining first return path data from a media service provider, the first return path data being from a set-top box in a common household monitored by an audience measurement entity using a meter, the first return path data comprising television tuning information reported by the set-top box in the common household, wherein the set-top box provides media to a media device in the common household and the meter monitors the media device; accessing the first return path data and panel meter data obtained from the meter at a data store of the audience measurement entity; classifying individual viewing segments associated with the common household as "matched" or "otherwise" based on whether the first return path data in the viewing segments has matching panel meter data for the viewing segments, wherein a given viewing segment is classified as "matched" if the first return path data for the given viewing segment has matching panel meter data for the given viewing segment, and the given viewing segment is classified as "otherwise" if the first return path data for the given viewing segment does not have matching panel meter data for the given viewing segment; generating a first set of features from the classified viewing segments, wherein the first set of features comprises: day of week, viewing segment duration, event duration, and average event duration for the common household on the day of week, wherein the event duration refers to a length of time that the media device is tuned to a particular channel; training a machine learning algorithm based on the first set of features to output on / off determinations for media devices; obtaining second return path data for a group of households from the media service provider, the group of households not being associated with and / or not registered with the audience measurement entity and not having a meter; applying the second return path data to the machine learning algorithm trained based on the first set of features to output first on / off determinations associated with media devices represented in the second return path data; further training the machine learning algorithm based on a second set of features, the second set of features being generated from the classified viewing segments; and applying the second return path data to the machine learning algorithm trained based on the second set of features to output second on / off determinations associated with media devices represented in the second return path data.
2. The method of claim 1, wherein, the machine learning algorithm classifies media devices associated with the second return path data as on based on a classification probability threshold.
3. The method of claim 2, wherein, the machine learning algorithm predicts whether media devices represented in the second return path data will have matching panel meter data based on the classification probability threshold.
4. The method of claim 1, further comprising generating a viewing segment index or a viewing segment duration, the viewing segment index corresponding to a position at which the viewing segment occurs in a tuning event.
5. The method of claim 1, further comprising: accessing, at a data store of the audience measurement entity, common household data for the plurality of common households, the common household data including the first return path data and panel meter data.
6. The method of claim 5, further comprising: grouping the common household data into viewing segments based on quarter hour segments.
7. The method of claim 6, wherein, the media device represented in the second return path data is different from any one of the respective media devices of the plurality of common households.
8. A computing system to perform media device on / off detection, the computing system comprising: a processor; and a non-transitory computer-readable storage medium having stored thereon program instructions that, when executed by the processor, are capable of performing a set of operations comprising: obtaining, from a media service provider, first return path data from set-top boxes in common households monitored by an audience measurement entity using meters, the first return path data including television tuning information reported by the set-top boxes, wherein the set-top boxes provide media to media devices in the common households and the meters monitor the media devices; accessing, at a data store of the audience measurement entity, the first return path data and panel meter data obtained from the meters; classifying respective viewing segments associated with the common households as "matched" or "otherwise" based on whether the first return path data for the viewing segments has matching panel meter data for the viewing segments, wherein a given viewing segment is classified as "matched" if the first return path data for the given viewing segment has matching panel meter data for the given viewing segment, and the given viewing segment is classified as "otherwise" if the first return path data for the given viewing segment does not have matching panel meter data for the given viewing segment; generating a first set of features from the classified viewing segments, wherein the first set of features includes: day of week, viewing segment duration, event duration, and average event duration for the common household on the day of week; wherein the event duration refers to a length of time that the media device is tuned to a particular channel; training a machine learning algorithm based on the first set of features to output on / off determinations for media devices; obtaining, from the media service provider, second return path data for a group of households that are not associated with and / or not registered with the audience measurement entity and do not have meters; applying the second return path data to the machine learning algorithm trained based on the first set of features to output first on / off determinations associated with media devices represented in the second return path data; further training the machine learning algorithm based on a second set of features generated from the classified viewing segments; and applying the second return path data to the machine learning algorithm trained based on the second set of features to output second on / off determinations associated with media devices represented in the second return path data.
9. The computing system of claim 8, wherein, the machine learning algorithm classifies the media device associated with the second return path data as on based on a classification probability threshold.
10. The computing system of claim 9, wherein, the machine learning algorithm predicts whether the media device represented in the second return path data will have matching panel meter data based on the classification probability threshold.
11. The computing system of claim 8, wherein, the operations further include generating a watch segment index or a watch segment duration, the watch segment index corresponding to a location of the watch segment occurrence in a tuning event.
12. The computing system of claim 8, wherein, the operations further include accessing, at a data store of the audience measurement entity, common household data for the plurality of common households, the common household data including the first return path data and the panel meter data.
13. The computing system of claim 12, wherein, the operations further include grouping the common household data into watch segments based on a quarter hour segment.
14. The computing system of claim 8, wherein, the media device represented in the second return path data is different from any one of the respective media devices of the plurality of common households.
15. A non-transitory computer-readable storage medium comprising computer-readable instructions that cause one or more processors to at least: obtaining first return path data from a media service provider, the first return path data being from set-top boxes in a common household monitored by an audience measurement entity using meters, the first return path data including television tuning information reported by the set-top boxes in the common household, wherein, the set-top box provides media to a media device in the common household and the meter monitors the media device; accessing, at a data store of the audience measurement entity, the first return path data and panel meter data obtained from the meter; classifying respective watch segments associated with the common household as "matched" or "otherwise" based on whether the first return path data for the watch segment has matching panel meter data for the watch segment, wherein a given watch segment is classified as "matched" if the first return path data for the given watch segment has matching panel meter data for the given watch segment; and the given watch segment is classified as "otherwise" if the first return path data for the given watch segment does not have matching panel meter data for the given watch segment; generating a first set of features from the classified watch segments, wherein the first set of features includes: day of week, watch segment duration, event duration, and average event duration for the common household on the day of week; wherein the event duration refers to a length of time that the media device is tuned to a particular channel; training a machine learning algorithm based on the first set of features to output on / off determinations for media devices; obtaining, from the media service provider, second return path data for a set of households that are not associated with and / or not registered with the audience measurement entity and do not have a meter; applying the second return path data to the machine learning algorithm trained based on the first set of features to output first on / off determinations associated with media devices represented in the second return path data; further training the machine learning algorithm based on a second set of features, the second set of features generated from the classified watch segments; and applying the second return path data to a machine learning algorithm trained based on the second set of features to output a second on / off determination result associated with a media device represented in the second return path data.
16. The storage medium of claim 15, wherein, the machine learning algorithm classifies the media device associated with the second return path data as on based on a classification probability threshold.
17. The storage medium of claim 16, wherein, the machine learning algorithm predicts whether the media device represented in the second return path data will have matching panel meter data based on the classification probability threshold.
18. The storage medium of claim 15, wherein, the instructions, when executed, further cause the one or more processors to generate a watch segment index or a watch segment duration, the watch segment index corresponding to a location at which the watch segment occurs in a tuning event.
19. The storage medium of claim 18, wherein, the instructions, when executed, further cause the one or more processors to access, at a data store of the audience measurement entity, common household data for the plurality of common households, the common household data including the first return path data and the panel meter data.
20. The storage medium of claim 19, wherein, the instructions, when executed, further cause the one or more processors to group the common household data into watch segments based on quarter hour segments.
Citation Information
Patent Citations
Media summarization
CN108737857A
Methods and apparatus to determine synthetic respondent level data using constrained markov chains
US20180376198A1