Integration of event reports and aggregated digest reports from privacy sandbox attribution reports

By removing false positives from the summary reports received by the API and reducing noise, generating a denoised aggregated summary report, the problem of interaction data deviation in the API report is solved, and more accurate interaction performance measurement is achieved.

CN120226310APending Publication Date: 2025-06-27GOOGLE LLC
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202480004806.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-07-14
Filing Date
2024-05-07
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The prior art has deviation problems in collecting interactive data from application programming interface (API) reports, including anonymization, aggregation, information truncation, and noise addition, resulting in inconsistent with the explicit action data and the interactive data.

Method used

By receiving a summary report of aggregated event data from the API, the false positive portion is removed, noise is reduced, and a denoised aggregated summary report is generated to provide more accurate measurement of interaction performance.

Benefits of technology

It realizes more accurate measurement of interactive performance, reduces the impact of noise, and improves the authenticity and reliability of data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120226310A_ABST
    Figure CN120226310A_ABST
Patent Text Reader

Abstract

The present disclosure generally describes methods, software, and systems for interactive performance analysis. An aggregated summary report including aggregated event data is received from the application programming interface, the aggregated event data being collected by the application programming interface during events corresponding to interactions with the user interface. The aggregated event data is aggregated using a hierarchical structure corresponding to the data type. A first portion of the aggregated event data identified as false positive is removed from the aggregated summary report to maintain true positive aggregated event data. At each stage of the hierarchical structure, noise from the true positive aggregated event data is reduced to generate a denoised aggregated summary report. The de-noised aggregated summary report is used to determine an operation. Instructions are provided to the asset provider system to activate at least one of the operations using the de-noised aggregated digest report.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to computer-implemented methods, software, and systems for combining data from event-level reports and aggregated summary reports of attribution reports from application programming interfaces (APIs). Background Art

[0002] An application programming interface (API) provides an interface that can be used in a computer application to access other systems and associated functionality. In some cases, an API can be used to collect and measure digital activities regarding digital components provided by a platform (e.g., a content provider). However, the interaction data reported by the API (including ad data generated to indicate user activities such as clicks corresponding to presented ads) can deviate from explicit action data (conversion data), which, in the context of digital components, refers to performing an action regarding the digital component at the time of the initial interaction with the digital component. The deviation can be due to anonymization, aggregation, information truncation, and the addition of noise. For example, event-level reports can be added with statistical noise, where attribution conversions are truncated within a pre-specified limit and reported with limited conversion metadata only within one, two, or three time windows. Summary of the Invention

[0003] Implementations of the present disclosure relate to techniques and tools for interaction performance analysis. More specifically, implementations of the present disclosure are directed to integrating data from event-level reports and aggregated summary reports of attribution reports to provide a measurement of interaction performance.

[0004] In some implementations, a method includes: receiving, from an application programming interface, an aggregated summary report including aggregated event data collected by the application programming interfaces during events corresponding to interactions with a user interface, the aggregated event data being aggregated using a hierarchical structure corresponding to a data type; removing, from the aggregated summary reports, a first portion of the aggregated event data identified as false positives to maintain true positive aggregated event data; reducing noise from the true positive aggregated event data at each level of the hierarchical structure to generate a denoised aggregated summary report; determining an operation using the denoised aggregated summary reports; and providing instructions to an asset provider system to activate at least one of the operations using the denoised aggregated summary reports.

[0005] The present disclosure also provides a computer-implemented system, comprising: a memory that stores application programming interface (API) information; and a server that performs operations including: receiving, from an application programming interface, an aggregated summary report including aggregated event data collected by the application programming interfaces during events corresponding to interactions with a user interface, the aggregated event data being aggregated using a hierarchical structure corresponding to a data type; removing, from the aggregated summary reports, a first portion of the aggregated event data identified as false positives to maintain true positive aggregated event data; reducing noise from the true positive aggregated event data at each level of the hierarchical structure to generate a denoised aggregated summary report; determining operations using the denoised aggregated summary reports; and providing instructions to an asset provider system to activate at least one of the operations using the denoised aggregated summary reports.

[0006] The present disclosure also provides a non-transitory computer-readable medium encoded with a computer program, the computer program comprising instructions that, when executed by one or more computers, cause the one or more computers to perform operations including: receiving, from an application programming interface, an aggregated summary report including aggregated event data collected by the application programming interfaces during events corresponding to interactions with a user interface, the aggregated event data being aggregated using a hierarchical structure corresponding to a data type; removing, from the aggregated summary reports, a first portion of the aggregated event data identified as false positives to maintain true positive aggregated event data; reducing noise from the true positive aggregated event data at each level of the hierarchical structure to generate a denoised aggregated summary report; determining operations using the denoised aggregated summary reports; and providing instructions to an asset provider system to activate at least one of the operations using the denoised aggregated summary reports.

[0007] The foregoing and other implementations may each optionally individually or in combination include one or more of the following features. Specifically, an implementation may include all of the following features: The aggregated summary reports include hierarchically structured event attribution configuration data as nodes distributed across multiple levels. Given the hierarchical structure of these aggregated summary reports, weighted averages of parent and child nodes may be generated to minimize the variance of the estimate of the parent node. The event attribution configuration data includes truncated configurations aggregated into data slices at one or more levels. The noise includes Laplace noise added to each of these data slices. Processing these aggregated summary reports to reduce the noise includes: applying a denoising transform to correct both the truncated data and the randomized response noise achieved by the Laplace noise. Applying the denoising transform includes using a set of summary statistics for aspects of the index configuration type. These aggregated summary reports include metadata corresponding to the configuration type applied to the structured event attribution configuration data.

[0008] Other implementations of this aspect include corresponding systems, devices, and computer programs configured to perform the actions of the method, encoded on a computer storage device.

[0009] The present disclosure also provides a computer-implemented method, including: receiving, from an application programming interface, an event-level report including branches that define records of events that include interactions with the application programming interfaces, the event-level reports paired with metadata corresponding to the configuration of the application programming interfaces; performing an identification of portions of the metadata that include invalid metadata; estimating, for each of these branches, the probability of being true or noised by using the identification of these portions of the metadata that include the invalid metadata, classifying portions of the events as being on true branches; determining configuration parameters for these true branches, the configuration parameters including the number of average configurations for each event; generating original application programming interface data by applying a debiasing model that uses these configuration parameters, by jointly considering multiple reporting windows and configuration types; determining an operation by using the original application programming interface data; and providing instructions to an asset provider system to activate at least one of these operations by using the original application programming interface data.

[0010] The present disclosure also provides a computer-implemented system, comprising: a memory that stores application programming interface (API) information; and a server that performs operations including: receiving, from the application programming interface, an event-level report including branches that define records of events including interactions with the application programming interfaces, the event-level reports being paired with metadata corresponding to the configurations of the application programming interfaces; performing an identification of portions of the metadata that include invalid metadata; estimating, for each of the branches, a probability of being true or noisy by using the identification of the portions of the metadata that include the invalid metadata, classifying portions of the events as being on true branches; determining configuration parameters for the true branches, the configuration parameters including a number of average configurations for each event; generating raw application programming interface data by applying a debiasing model using the configuration parameters by combining multiple reporting windows and configuration types; determining operations using the raw application programming interface data; and providing instructions to an asset provider system to activate at least one of the operations using the raw application programming interface data.

[0011] The present disclosure also provides a non-transitory computer-readable medium encoded with a computer program, the computer program comprising instructions that, when executed by one or more computers, cause the one or more computers to perform operations including: receiving, from the application programming interface, an event-level report including branches that define records of events including interactions with the application programming interfaces, the event-level reports being paired with metadata corresponding to the configurations of the application programming interfaces; performing an identification of portions of the metadata that include invalid metadata; estimating, for each of the branches, a probability of being true or noisy by using the identification of the portions of the metadata that include the invalid metadata, classifying portions of the events as being on true branches; determining configuration parameters for the true branches, the configuration parameters including a number of average configurations for each event; generating raw application programming interface data by applying a debiasing model using the configuration parameters by combining multiple reporting windows and configuration types; determining operations using the raw application programming interface data; and providing instructions to an asset provider system to activate at least one of the operations using the raw application programming interface data.

[0012] The foregoing and other implementations can each optionally individually or in combination include one or more of the following features. Specifically, an implementation can include all of the following features: These configuration parameters include a configuration window and a configuration count for each configuration type. The computer-implemented method further includes: when truncated for each configuration type, determining a truncated average as the expected total configuration count for each such configuration type and window. Determining the truncated average includes: determining a total truncated average by aligning aggregated item configuration counts with the application programming interface data using the display dates; and determining a sliced truncated average by aligning the total aggregated item configuration counts with the application programming interface data using the display dates and a latency window level. These total truncated averages include the total configuration count truncated at the configured number. These total truncated averages include edge cases that do not exist in the event simulation but occur in the total aggregated item configuration count. These sliced truncated averages include a truncation rate per data slice. Estimating the probability that each of these branches is true or noisy includes generating a ratio of the total interaction count of display dates obtained from the interaction log to the total number of conditioning interactions.

[0013] The present disclosure also provides a computer-implemented method, including: receiving, from an application programming interface, an event-level report including a biased record of events, the events including interactions with the application programming interfaces, the event-level reports being paired with metadata corresponding to configurations applied by the application programming interfaces; generating an original event-level report from the event-level reports by applying a debiasing model using the configuration parameters of the application programming interfaces to remove spurious events from the event-level reports; receiving, from the application programming interfaces, an aggregated summary report, the aggregated summary reports including an aggregated record of the events corresponding to the interactions with the application programming interfaces included in the event-level reports; generating an original aggregated summary report from the aggregated summary reports by using a metadata mapping to remove false positive events from the event-level reports; generating statistical data by matching the original event-level reports with the original aggregated summary reports according to event scenarios; determining an operation using the statistical data; and providing instructions to an asset provider system to activate at least one of the operations using the statistical data.

[0014] The present disclosure also provides a computer-implemented system, comprising: a memory that stores application programming interface (API) information; and a server that performs operations including: receiving, from the application programming interfaces, event-level reports including biased records of events, the events including interactions with the application programming interfaces, the event-level reports being paired with metadata corresponding to the configurations applied by the application programming interfaces; generating, from the event-level reports, raw event-level reports by applying a debiasing model using the configuration parameters of the application programming interfaces to remove spurious events from the event-level reports; receiving, from the application programming interfaces, aggregated summary reports that include aggregated records of the events corresponding to the interactions with the application programming interfaces included in the event-level reports; generating, from the aggregated summary reports, raw aggregated summary reports by using metadata mapping to remove false positive events from the event-level reports; generating statistical data by matching the raw event-level reports with the raw aggregated summary reports according to event scenarios; determining operations using the statistical data; and providing instructions to an asset provider system to activate at least one of the operations using the statistical data.

[0015] The present disclosure also provides a non-transitory computer-readable medium encoded with a computer program, the computer program comprising instructions that, when executed by one or more computers, cause the one or more computers to perform operations including: receiving, from the application programming interfaces, event-level reports including biased records of events, the events including interactions with the application programming interfaces, the event-level reports being paired with metadata corresponding to the configurations applied by the application programming interfaces; generating, from the event-level reports, raw event-level reports by applying a debiasing model using the configuration parameters of the application programming interfaces to remove spurious events from the event-level reports; receiving, from the application programming interfaces, aggregated summary reports that include aggregated records of the events corresponding to the interactions with the application programming interfaces included in the event-level reports; generating, from the aggregated summary reports, raw aggregated summary reports by using metadata mapping to remove false positive events from the event-level reports; generating statistical data by matching the raw event-level reports with the raw aggregated summary reports according to event scenarios; determining operations using the statistical data; and providing instructions to an asset provider system to activate at least one of the operations using the statistical data.

[0016] The foregoing and other implementations may each optionally include, individually or in combination, one or more of the following features. Specifically, an implementation may include all of the following features: The biased recording of events includes anonymizing, aggregating, information truncating, and noise injecting the interaction data to protect the privacy of users performing these interactions with these application programming interfaces. These aggregated summary reports include hierarchically structured event attribution configuration data as nodes distributed across multiple levels, and the event attribution configuration data includes truncated configurations aggregated as data slices at one or more levels. The hierarchical structure of the aggregated summary reports can be used to generate weighted averages of parents and children to minimize the variance of the estimates of the parent nodes. The noise includes Laplace noise added to each of these data slices. Post-processing these raw aggregated summary report data includes two steps: removing false positives in the data and taking full advantage of the hierarchical structure to reduce the noise. Processing these raw aggregated summary reports to reduce the noise includes: applying a denoising transform to correct both the truncated configurations and the randomized response noise implemented by the Laplace noise. The denoising transform uses a set of summary statistics indexing aspects of the configuration type. These aggregated summary reports include metadata corresponding to the configuration types applied to the structured event attribution configuration data. These configuration parameters include a configuration window and a configuration count for each configuration type. The computer-implemented method further includes: when the configuration is truncated, determining the truncated average as the expected total configuration count for each configuration type and window. Determining the truncated average includes: determining the total truncated average by aligning the aggregated item configuration counts with the application programming interface data; and determining the sliced truncated average by aligning the total aggregated item configuration counts with the application programming interface data. These total truncated averages include the total configuration count truncated at the number of configurations. These total truncated averages include edge cases. These sliced truncated averages include a per-data-slice truncation rate. Estimating the probability that each of these branches is true or noisy includes generating a ratio of the total interaction count of the display date obtained from the interaction log relative to the total conditional interactions.

[0017] It is understood that the methods according to the present disclosure may include any combination of the aspects and features described herein. That is, the methods according to the present disclosure are not limited to the combinations of aspects and features specifically described herein, but also include any combination of the provided aspects and features.

[0018] Specific implementations of the subject matter described in this specification can achieve one or more of the following advantages. The techniques described in this specification provide protection for user data privacy and security. The adaptability of data processing to multiple types of API configurations enables the flexibility of technology integration. Statistical data including interaction measurements can be generated and consumed faster than in traditional systems, where separate and different protocols are applied. By leveraging the combined use of API data (which provides better measurement fidelity compared to using either reporting type alone), the accuracy of interaction measurements is improved by generating statistical data through the merger of event-level data and aggregated summary reports. A constructed machine learning model can be used to support and optimize the merger of event-level data and aggregated summary reports to optimize the interaction measurement derivation process. As an interaction proceeds, the API returns some information about conversions that may have occurred (or not occurred) within a predefined duration after the interaction. Typical use cases for event-level reports can include model training. Conditional on the characteristics of the interaction, the trained information can be used to predict conversions, conversion rates, or conversion values. Predicting conversions is a key input to automated bidding models, as it serves as a (random) signal for the bidding algorithm indicating which events are likely to result in conversions acceptable to the content provider. The bidding algorithm algorithmically determines its bid by taking the interaction as input, and thus the quality of the model prediction directly affects the quality of bidding optimization. To protect user privacy, the API does not return event-level data with full fidelity. Instead, a small subset of interactions is randomly selected by the browser / platform and assigned random conversion metadata. Additionally, there are limitations on how much metadata can be extracted from conversions.

[0019] Other advantages of the described implementations are associated with eventification, which refers to the process of extracting event-level explicit action data from event-level reports and from aggregated summary reports. One advantage of eventification is that, despite using completely different measurement techniques, eventified logs are also structurally similar to event-level data logged by third-party cookies. The structural similarity facilitates inserting eventified logs into existing data pipelines and other modeling infrastructure. The compatibility of eventified logs with existing data pipelines optimizes the use of computing resources by reducing technical debt, thus facilitating the transition to systems including the ARA API. Another advantage of eventification is that the same eventified logs can be used for many use cases, including the top one (or two) main reports and bidding use cases. For reporting, eventified logs can be appropriately aggregated, for example, at the campaign level, according to the slices used for the report. For bidding, eventification facilitates using conversions (or conversion values) as labels and training machine learning models characterized by interactive event features. The training of machine learning models uses a training data set in units of a combination of interactive events and attributed conversions (or conversion values) as input, matching the structure of eventified logs. Building both reports and bidding based on the same log can reduce processing complexity and automatically provide consistency across use cases.

[0020] Details of one or more implementations of the subject matter of the specification are set forth in the accompanying drawings and the following description. Other features, aspects, and advantages of the subject matter will become apparent from the description, the drawings, and the claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Figure 1 is a block diagram of an example system that can be used to execute implementations of the present disclosure.

[0022] Figure 2A is a block diagram of another example system according to some implementations of the present disclosure.

[0023] Figure 2B is according to some implementations of the present disclosure Figure 2A block diagram of the data flow within an example system of.

[0024] Figure 3 is a flowchart of an example denoising process according to some implementations of the present disclosure.

[0025] Figure 4A is an example aggregated summary report according to some implementations of the present disclosure.

[0026] Figure 4B is example aggregated summary report data according to some implementations of the present disclosure.

[0027] Figure 4CAn aggregated summary report of example denoising according to some implementations of the present disclosure.

[0028] Figure 5 A flowchart of an example bias removal process according to some implementations of the present disclosure.

[0029] Figure 6A Example raw event-level data according to some implementations of the present disclosure.

[0030] Figure 6B Example event application programming interface data according to some implementations of the present disclosure.

[0031] Figure 7A A graph of example event branch classification according to some implementations of the present disclosure.

[0032] Figure 7B A graph of an example transformation based on event branch classification according to some implementations of the present disclosure.

[0033] Figure 8 A flowchart of an example process according to some implementations of the present disclosure.

[0034] Figure 9 A graph of example probability estimation according to some implementations of the present disclosure.

[0035] Figure 10A A graph of example fake probability estimation according to some implementations of the present disclosure.

[0036] Figure 10B A graph of example jointly truncated data according to some implementations of the present disclosure.

[0037] Figure 10C A graph of example jointly truncated average data according to some implementations of the present disclosure.

[0038] Figure 11A An example of explicit action data according to some implementations of the present disclosure.

[0039] Figure 11B An example of event-level report data according to some implementations of the present disclosure.

[0040] Figure 11C An example of an eventified log after event denoising data according to some implementations of the present disclosure.

[0041] Figure 11D An example of event-level application programming interface output data according to some implementations of the present disclosure.

[0042] Figure 11EAn example of debiased event-level application programming interface result data according to some implementations of the present disclosure.

[0043] Figure 11F An example of probability estimation on a false branch according to some implementations of the present disclosure.

[0044] Like reference numerals and names in the various figures indicate like elements. Detailed Implementations

[0045] Implementations of the present disclosure are directed to techniques and tools for interactive performance analysis. More specifically, implementations of the present disclosure are directed to integrating data from event-level reports and aggregated summary reports from attribution reports to provide a measurement of interactive performance. The event-level report includes filtered interactive event data corresponding to interactive events, which are generated according to multiple conversion types that can be reported in a truncated format. The event-level report is received from application programming interfaces (APIs) of different source systems, which may have different ways of exposing metadata corresponding to the APIs and events. The aggregated summary report includes data aggregations generated by grouping event attribution explicit action data aggregated to the slice level. The aggregated summary report is configured based on predefined slices through which the interactive provider system plans to understand conversion activities. The event-level report and the aggregated summary report represent two different views of the same underlying interactive data. The nature of the data generated by both depends on how each transforms the same underlying data to protect user privacy. The described process includes deriving user interaction measurements based on the applied data privacy transformation. For example, the described techniques consider two aspects of operating API data: conversion truncation and noise consideration, and how these aspects differ for each API according to the corresponding configuration. The term "event" in the event-level report corresponds to an interactive event. That is, the event-level report includes reports with a granularity defined by an interaction (such as a click or a view).

[0046] Collecting interaction information from API reports is cumbersome due to deviations. Deviations can be due to, for example, anonymization, aggregation, information truncation, and addition of noise. For example, event-level reports can have statistical noise added, where attributed conversions are truncated within pre-specified limits and reported with limited conversion metadata only within one, two, or three time windows. Event-level reports include deviations as biased records of events indicating interactions with the user interface, biasing the recording of events to protect the security and privacy of user data. Event-level reports can be paired with metadata corresponding to the API configuration. Event-level reports can be configured according to the corresponding API configuration. The configured event-level reports can be processed to generate the original event-level reports according to the configured event-level reports by applying a debiasing model that uses the configuration parameters of the application programming interface to reduce (or remove) false events in the event-level reports.

[0047] In addition to generating event-level reports, the API can also generate an aggregated summary report including event attribution configuration data. The event-level reports can be configured as the configured aggregated summary reports according to the corresponding API configuration. The configured aggregated summary reports can be processed by using metadata mapping to remove false positive events from the event-level reports to generate the original aggregated summary reports. The original event-level reports are matched with the original aggregated summary reports according to the event scenario (defining the interaction data deviation policy, such as the truncation process) to generate statistical data. The statistical data can be used to determine one or more matching operations, and the one or more matching operations can be filtered and sorted according to the set selection criteria to provide instructions to the asset provider system to activate at least one of the operations using the statistical event attribution data.

[0048] As interaction monitoring systems (e.g., the advertising ecosystem) make a significant pivot towards ways to improve the protection of user privacy, the use of privacy-enhancing technologies (PETs) such as the Attribution Reporting API (ARA) can become increasingly relevant for interaction measurement. PETs focused on measurement can facilitate interaction measurement while protecting the user's cross-site and cross-application identifiers from being revealed to interaction service provider systems such as content provider systems, service providers (e.g., advertisers), publishers, and other entities (hereinafter, collectively referred to as the general term "content provider systems"). The particularities of data privacy implementation methods can include one or more differences and at least one common feature. The common feature of PETs is that, before publishing data to the asset provider system and / or content provider system through a secure distribution system, the information content of campaign performance data is restricted by using some combination of anonymization, aggregation, information truncation, and noise injection on the data. The common feature of the PET process can provide privacy guarantees while continuing to support service (e.g., testing or advertising) use cases in a sustainable manner.

[0049] The systems and processes described in the present disclosure provide techniques for deriving statistical data that can be used by content provider systems and / or asset provider systems, which consume interaction data and utilize the statistical data for operations related to various interaction use cases. The described techniques for deriving statistical data address several new data and modeling related problems that did not previously exist in third-party cookies (3PCs). These problems stem from changes introduced by API data for protecting user data privacy. Due to anonymization, aggregation, information truncation, and addition of noise, the data reported by the API to the content provider system can deviate from the explicit action data measured by the 3PC (hereinafter referred to as 3PC conversions). For ARA, event-level reports can have statistical noise added, where attributed conversions are truncated within a pre-specified limit and reported with limited conversion metadata only within one, two, or three time windows. Summary reports that can be obtained from the aggregated summary reports can be at the aggregated slice level, can have statistical noise added, and can have conversions truncated within a preset limit. The described systems enable content provider systems and / or asset provider systems to consume and process the ARA data (event-level reports and summary reports) received from the API to generate statistical data that can be used for service (e.g., testing or advertising) use cases.

[0050] To address the limitations of API data deviation, the integration protocol described in the present disclosure enables bundling resources across different protocols without using third-party cookies. If third-party cookies are used, control of user privacy and system security may fail. For example, third-party tracking requests can use nested scripts to avoid detection and user interaction data tracking can deliberately block its own referral URL. In comparison to third-party user interaction data tracking, the described implementation presents user interaction data tracking as aggregated summaries and event-level reports generated by the API. The configuration settings of the API can be designed to protect user privacy and enable control of system security. Another advantage of the described implementation is that even though API-based user interaction data tracking utilizes completely different measurement techniques, the generated reports can be structurally similar to the user interaction data recorded by third-party cookies. The structural similarity facilitates inserting evented logs into existing data pipelines and other modeling infrastructures built for third-party cookie data with few changes while still protecting user privacy and system security. The integration protocol described in the present disclosure includes a set of methodologies through which data from event-level reports and aggregated summary reports can be merged together to facilitate interaction measurement with high utility. The described process of combining and merging data sets can provide guidance on how to utilize the API to improve service (e.g., advertising) measurement.

[0051] In addition to protecting user data privacy and security, the interaction measurement protocol described in this disclosure can also provide adaptability of data processing to multiple types of API configurations that enables flexibility in technical integration. As another technical advantage of the described technology, statistical data including interaction measurements and consumption statistics can be generated faster than in traditional systems where separate different protocols are applied. By leveraging the combined use of API data (which provides better measurement fidelity compared to using either reporting type alone), the accuracy of interaction measurement is improved by generating statistical data through the merger of event-level data and aggregated summary reports. Machine learning models can be constructed to support and optimize the merger of event-level data and aggregated summary reports to optimize the interaction measurement derivation process. As the interaction progresses, the API returns some information about conversions that may have occurred (or not occurred) within a predefined duration after the interaction. A typical use case for event-level reporting can include model training. Conditional on the characteristics of the interaction, the trained information can be used to predict conversions, conversion rates, or conversion values. Predicting conversions is a key input to the automated bidding model as it serves as a (stochastic) signal for the bidding algorithm indicating which events are likely to result in conversions acceptable to the content provider. The bidding algorithm algorithmically determines its bid by taking the interaction as an input, and thus the quality of the model prediction directly affects the quality of bidding optimization. To protect user privacy, the API does not return event-level data with full fidelity. Instead, a small subset of the interactions is randomly selected by the browser / platform and assigned random conversion metadata. Additionally, there are limitations on how much metadata can be extracted from the conversions.

[0052] Figure 1 FIG. 100 is a block diagram showing an example system 100 for deriving interaction measurements from API reports. Specifically, the example system 100 shown includes a server system 102, a client device 104, a content provider system (and / or asset provider system) 106, an API provider system 110, and a network 108 or is communicatively coupled thereto. Although shown separately, in some implementations, a single system or server can provide the functions of two or more systems or servers. In some implementations, the functions of one of the shown systems, servers, or components can be provided by multiple systems, servers, or components.

[0053] In Figure 1In the example of , the server system 102 is intended to represent various forms of servers, including but not limited to web servers, application servers, proxy servers, network servers, and / or server pools. Typically, the server system 102 accepts requests for application services (such as testing services, advertising services, experimental services), and provides these services to any number of client devices 104 (for example, client devices 104 on a network 108). In accordance with implementations of the present disclosure, and as noted above, the server system 102 can host a solution environment, which can be a cloud environment that provides software applications, systems, and services (such as content within an application that is displayed on a client device 104 and can be consumed as a service by an entity). Interactions generated in response to the services provided can be measured and provided to a content provider system (and / or an asset provider system) 106. In some cases, the server system 102 can support configuring different types of APIs, as well as different types of services that are integrated into user privacy settings (scenarios) and support process execution, such as reference Figure 3 , Figure 5 , Figure 8 and Figures 10A to 10C as described.

[0054] The server system 102 includes a processor 112A, a memory 114A, and an interface 116A. The memory 114A may include event-level reports 120A, aggregated summary reports 120B, and metadata 122. The event-level reports 120A, aggregated summary reports 120B may include documents defining events (e.g., interactions with a user interface) recorded by a resource (API) provided by the API provider system 110. The metadata 122 provides additional information related to ad interactions and / or conversions. In some implementations, the metadata 122 may include encoded data pointing to an API configuration (e.g., defining the type of conversion applied by the corresponding API).

[0055] The client device 104 and the API provider system 110 can each be any computing device operable to connect to or communicate in the network 108 using a wired connection or a wireless connection. Typically, each of the client device 104 and the API provider system 110 includes an electronic computer device operable to receive, transmit, process, and store information related to the network 108. Figure 1Any suitable data corresponding to system 100. Each of client device 104 and API provider system 110 is generally intended to include any computing device such as: laptop / notebook computer, wireless data port, smart phone, personal data assistant (PDA), tablet computing device, one or more processors within these devices, or any other suitable processing device. Client device 104 and API provider system 110 respectively include interfaces 116B, 116C, processors 112B, 112C, memories 114B, 114C, and graphical user interfaces (GUI) 124A, 124B.

[0056] Client device 104 may include one or more client applications 126. Client application 126 may be any type of application that allows the client device (e.g., an Internet browser) to request and view content on the client device. In some implementations, client application 126 may correspond to an API that can record user parameters, metadata, and other API event information, and may process this other API event information in a specific manner to protect user privacy before transmitting data (event-level report 120A, aggregated summary report 120B) to server 102. In some cases, client application 126 may be a proxy-side or client-side version of one or more enterprise applications running on an enterprise server (not shown). The memory 114C of target API provider system 110 may include API client 132, API resource 134, and event resource 136 that can be used to integrate dependencies.

[0057] The client device 104 and / or the API provider system 110 may each include a computer that includes: an input device, such as a keyboard, a touch screen, or other device that can accept user information; and an output device or GUI 124A, 124B that conveys information corresponding to the operation of the server 102 or the client device itself, the information including digital data, visual information. The GUIs 124A, 124B are each interfaced with at least a portion of the system 100 for any suitable purpose, including generating visual representations of the client application 126 or the management application 133, respectively. Specifically, the GUIs 124A, 124B may each be used to view and navigate various web pages. Generally, the GUIs 124A, 124B each provide the user with an efficient and user-friendly presentation of the object data (metadata) provided by or communicated within the system. The GUIs 124A, 124B may each include a plurality of customizable frames or views that have interactive fields, drop-down lists, and buttons that are operated by the user during recordable events, which may be included in the data collected by the API (e.g., event-level reports 120A, aggregated summary reports 120B). The GUIs 124A, 124B each contemplate any suitable graphical user interface, such as a combination of a general web browser, an intelligent engine, and a command line interface (CLI) that processes information and presents the results intuitively and efficiently to the user.

[0058] The content provider system (and / or asset provider system) 106 may include multiple systems present in a multi-system landscape. For example, an organization may use different types of systems to run the organization. The content provider system (and / or asset provider system) 106 may include systems from the same entity or different entities. The content provider system (and / or asset provider system) 106 may each include at least one of an interface 116D, a processor 112D, and an interactive data integration system 128. The interactive data integration system 128 may include an implementation of operations associated with statistical data indicative of interaction measurements. The operation implementation capabilities include a set of criteria to select and trigger the automatic implementation of operations based on statistical event attribution data. The interactive data integration system 130 may filter the entity landscape from multiple asset provider systems 106 based on the API configuration to identify suitable operation targets, and may automatically select the identified API provider system 110 to establish a connection with either the client device 104 and / or the API provider system 110 via the network 108.

[0059] In some implementations, network 108 may include a large computer network that connects any number of communication devices, mobile computing devices, fixed computing devices, and server systems, such as a local area network (LAN), a wide area network (WAN), the Internet, a cellular network, a telephone network (PSTN), or a suitable combination thereof. Data exchanged through network 108 is transmitted using any number of network layer protocols, such as Internet Protocol (IP), Multiprotocol Label Switching (MPLS), Asynchronous Transfer Mode (ATM), Frame Relay, etc. Additionally, in implementations where network 108 represents a combination of multiple sub-networks, different network layer protocols are used at each underlying sub-network. In some implementations, network 108 represents one or more interconnected internetworks, such as the public Internet.

[0060] Each of the processors 112A, 112B, 112C, 112D included in the client device 104, the content provider system (and / or asset provider system) 106, or the API provider system 110 may be a central processing unit (CPU), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), or another suitable component. Generally, each of the processors 112A, 112B, 112C, 112D included in the client device 104 or the API provider system 110 executes instructions and manipulates data to perform the operations of the client device 104 or the API provider system 110, respectively. Specifically, each of the processors 112A, 112B, 112C, 112D included in the client device 104 or the API provider system 110 performs functions that are used to send requests to the server 102 and receive and process responses from the server 102. Each of the processors 112A, 112B, 112C, 112D may be a central processing unit (CPU), a blade processor, an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), or another suitable component. Each of the processors 112A, 112B, 112C, 112D executes instructions and manipulates data to perform the operations of the respective system (server system 102, client device 104, API provider system 110, and content provider system (and / or asset provider system) 106). Specifically, each of the processors 112A, 112B, 112C, 112D performs functions that are used to receive and respond to requests from the respective system (server system 102, client device 104, API provider system 110, and content provider system (and / or asset provider system) 106).

[0061] Interfaces 116A, 116B, 116C, and 116D are used by server 102, client device 104, schema system 106, and API provider system 110, respectively, for communicating with other systems in a distributed schema connected to network 108, including system 100. Generally, interfaces 116A, 116B, 116C, and 116D each include logic encoded in software and / or hardware in a suitable combination and operable to communicate with network 108. More specifically, interfaces 116A, 116B, 116C, and 116D may each include software that supports one or more communication protocols corresponding to the communication, such that network 108 or the hardware of the interface is operable to communicate physical signals both inside and outside the illustrated system 100.

[0062] Memories 114A, 114B, 114C may include any type of memory or database engine and may take the form of volatile and / or non-volatile memory, including but not limited to magnetic media, optical media, random access memory (RAM), read-only memory (ROM), removable media, or any other suitable local or remote memory component. Memories 114A, 114B, 114C may store various objects or data, including caches, classes, frameworks, applications, backup data, objects, jobs, web pages, web page templates, database tables, database queries, repositories storing entity information and / or dynamic information, and any other suitable information, where the any other suitable information includes any parameters, variables, algorithms, instructions, rules, constraints, or references thereto corresponding to the purposes of server system 102, client device 104, API provider system 110, or schema system 106, respectively.

[0063] There may be any number of client devices 104 and API provider systems 110 corresponding to or external to system 100 for collecting and processing interaction event data. Additionally, there may be one or more additional client devices external to the illustrated portion of system 100 capable of interacting with system 100 via network 108. Further, the terms "client", "client device", and "user" may be used interchangeably, as appropriate, without departing from the scope of the present disclosure. Moreover, although a client device may be described as being used by a single user, the present disclosure contemplates that many users may use one computer, or one user may use multiple computers. As used in the present disclosure, the term "computer" is intended to encompass any suitable processing device. For example, although Figure 1shows a single server 102, a single client device 104, and a single API provider system 110, but system 100 can be implemented using a single stand-alone computing device, two or more servers 102, or multiple client devices. Server system 102, client device 104, and API provider system 110 can include any computer or processing device, such as a blade server, a general-purpose personal computer (PC), a workstation, a UNIX-based workstation, or any other suitable device. In other words, the present disclosure contemplates computers other than general-purpose computers and computers without a conventional operating system. Further, server 102, client device 104, and API provider system 110 can be adapted to execute any operating system or runtime environment, including Linux, UNIX, Windows, Mac Java TM 、Android TM 、iOS, BSD (Berkeley Software Distribution), or any other suitable operating system. According to one implementation, server 102 can also include an email server, a web server, a cache server, a streaming data server, and / or another suitable server or be communicatively coupled thereto.

[0064] Regardless of the specific implementation, "software" can include computer-readable instructions, firmware, wired and / or programmed hardware, or any combination thereof (temporary or non-temporary, as appropriate) on a tangible medium that, when executed, is operable to perform at least the processes and operations described herein. In fact, each software component can be written or described in whole or in part in any suitable computer language, including C, C++, Java TM 、 Visual Basic, assembly language, ABAP (Advanced Business Application Programming), ABAP OO (Object Oriented), any suitable version of a fourth-generation programming language, and other languages. Although Figure 1 the software portions shown in

[0065] Figure 2Ais a block diagram of an example system 200A for debiasing interactive measurement data using a secure distribution system 202. The example system 200 shown includes a secure distribution system 202, a client device 204, a content provider system 206, an asset provider system 210, and a network 208, or is communicatively coupled thereto. The secure distribution system 202 may be included in a server system (e.g., the server system 102 described with reference to Figure 1 Although shown separately, in some implementations, the secure distribution system 202 may be included in the client device 204 (e.g., the client device 104 described with reference to Figure 1 ), the content provider system 206 (e.g., the system 106 described with reference to Figure 1 ), or the asset provider system 210 (e.g., the system 106 described with reference to Figure 1 ), or may be communicatively coupled to any one of the client device 204, the content provider system 206, and the asset provider system 210 via the network 208 (e.g., the network 108 described with reference to Figure 1 ).

[0066] The client device 204 may include an application 205, such as a web browser and / or a native application, to facilitate sending and receiving data over the network 208. A native application is an application developed for a specific platform or a specific device (e.g., a mobile device with a specific operating system). Although operations may be described as being performed by the client device 204, such operations may be performed by the application 205 running on the client device 204. The application 205 may present electronic resources, such as web pages, application pages, or other application content, to a user of the client device 204. The electronic resources may include digital component slots for presenting digital components together with the content of the electronic resources. A digital component slot is an area of an electronic resource (e.g., a web page or an application page) for displaying a digital component. A digital component slot may also refer to a portion of an audio and / or video stream (which is another example of an electronic resource) for playing a digital component.

[0067] Electronic resources are also referred to as resources in this document. For the purposes of this document, a resource can refer to a web page, an application page, application content presented by a native application, an electronic document, an audio stream, a video stream, or other suitable types of electronic resources that can be used to present digital components. As used throughout this document, the phrase "digital component" refers to discrete units of digital content or digital information (e.g., video clips, audio clips, multimedia clips, images, text, or another content unit). Digital components can be stored electronically as a single file or as a collection of files in a physical memory device, and digital components can take the form of video files, audio files, multimedia files, image files, or text files and include service (e.g., test or advertisement) information such that the interaction is of the digital component type. For example, a digital component can be content that is intended to complement the content of a web page or other resource presented by Application 205. More specifically, a digital component can include digital content related to the resource content (e.g., the digital component can relate to the same topic as the web page content, or to a related topic). The configuration of the digital component can complement and generally enhance the web page or application content.

[0068] In response to Application 205 loading a resource that includes a digital component slot, Application 205 can generate a digital component request 225 that requests a digital component for presentation in the digital component slot. In some implementations, the digital component slot and / or the resource can include code (e.g., a script) that causes Application 205 to request a digital component from Content Provider System 206 (which can be logged as interaction event data by API 207).

[0069] The interaction event data logged by API 207 can include data related to the user of Client Device 204 and / or non-sensitive data such as a query string. Data related to the user can include, for example, data that identifies a user group of which the user is a member. User groups can include interest-based groups. Each interest-based group can include a topic of interest and a set of members identified (e.g., determined or predicted) as being interested in that topic. User groups can also include, for example, groups of users who have performed a particular action on a publisher's electronic resource (e.g., a website or native application). For example, user groups can include users who have browsed a website, requested more information about an item, interacted with a particular digital component (e.g., selected a particular digital component), and / or added an item to a virtual shopping cart for potential acquisition of the item. Data related to the user can also include user profile data and / or attributes of the user.

[0070] In addition to the description throughout this document, controls (e.g., user interface elements with which a user can interact) can be provided to a user, allowing the user to make choices regarding whether and when the systems, programs, or features described herein can enable the collection of user information (e.g., information about a user's social network, social actions, or activities, occupation, user preferences, or user's current location) and whether to send content or communications to the user from a server. Additionally, specific data can be processed in one or more ways before being transmitted for storage or use in the digital component repository 212 of the secure distribution system such that personally identifiable information is truncated (at least partially removed) and noise is added to hide private data. For example, a user's identity can be processed such that the user's personally identifiable information cannot be determined, or in the case of obtaining location information, the user's geographical location can be generalized (such as to a city, zip code, or state level) such that the user's specific location cannot be determined. A user can have control over what information about the user is collected, how that information is used, and what information is provided to the content provider system 206 and the asset provider system 210.

[0071] Interaction event data logged by the API 207 can also include context data that is generally considered non-sensitive. The context data can describe the environment in which the selected digital component is presented. The context data can include, for example: approximate location information indicating the approximate location of the client device 204 that sent the digital component request; a resource (e.g., a website or native application) with which the selected digital component will be presented; the spoken language settings of the application 205 or the client device 204; the number of digital component slots in which the digital component is presented with the resource; the type of digital component slot; and other appropriate context information.

[0072] The secure distribution system 202 can be implemented using one or more server computers (or other suitable computing devices) that can be distributed across multiple locations. Generally, the secure distribution system 202 receives a request for a digital component from the client device 204, selects a digital component based on the data included in the request, and sends the selected digital component to the client device 204. In some implementations, the secure distribution system 202 can be operated and maintained by an independent trusted party (e.g., a party different from the user of the client device, an operating supply-side platform (SSP) and a demand-side platform (DSP), and a digital component provider) to ensure the security and privacy of the data. For example, the secure distribution system 202 can be operated by an industry group or a government group.

[0073] The secure distribution system 202 may include a digital component repository 212, a metadata mapping engine 214, an event API preprocessor 216, an interaction aggregator 218, a parameter estimator 220, a debiasing engine 222, and an interaction usage engine 224. The digital component repository 212 may be a database configuration tool for storing data, which includes data received from APIs, such as metadata 226, event-level reports 228, summary reports 230, and interaction data 232. The event-level reports 228 and summary reports 230 may include eventified and modeled data logs, which provide two types of reports at different granularity levels (one more detailed and one more in summary form) to reflect the same information content. As referenced Figures 4A to 4C and Figure 6A and Figure 6B as described, the event-level reports 228 and summary reports 230 include tabulated logs, where rows represent interaction events and columns represent the results attributed to the interaction events (conversion counts and conversion values).

[0074] The metadata mapping engine 214 may access (obtain or retrieve) the metadata 226 from the digital component repository 212 and provide the output of metadata processing to the event API preprocessor 216. The metadata mapping engine 214 may filter out some conversions by looking up a metadata mapping table that may be stored in the digital component repository 212. If the metadata 226 for an identified log entry is not registered in the metadata mapping table, it is determined that the conversion of that log entry is on a false branch.

[0075] The event API preprocessor 216 may be configured to process the input received from the metadata mapping engine 214 and the event-level reports 228 retrieved from the digital component repository 212 to generate an output provided to the parameter estimator 220 and the debiasing engine 222. The interaction aggregator 218 may access (obtain or retrieve) the interaction data 232 from the digital component repository 212 and provide the output of interaction data processing to the parameter estimator 220.

[0076] The debiasing engine 222 can be a data processing pipeline that uses a log aggregator (e.g., Flume C++) and protocol buffers. The debiasing engine 222 can include a debiasing layer to run the pipeline periodically and generate debiased event API data by processing the inputs received from the event API preprocessor 216 and the parameter estimator 220. For example, the data exported by the event API preprocessor 216 from the event-level reports 228 can be processed by the debiasing engine 222 to recover the underlying 3PC transformation applied to the interaction data using the aggregated information from the aggregated summary reports 230 while preserving the event-level nature of the interaction data. Due to the privacy protection characteristics of the API, the recovered data is not the same as the event-level 3PC explicit action data, but the output of the debiasing engine 222 includes actionable event-level data logs that can still be used by the interaction usage engine 224 for interaction use cases.

[0077] In some implementations, the debiasing engine 222 performs a post-mapping process based on a (trainable) machine learning model that maps transformed metadata values to transformed types or bidability information. A training dataset in units of a combination of interaction events and attributed conversions (or conversion values) can be used as input to train the machine learning model to match the structure of the evented logs. Building both reports and bid generation (e.g., initiating bids to a set of service providers) based on the same logs can reduce processing complexity and automatically provide consistency across use cases. For example, the debiasing engine 222 can be used to train a machine learning model on denoised aggregates and event-level report data to predict values for the evented logs, resulting in a trained model. Conversions (or conversion values) can be provided as input features to the training system, and interaction event features contain values that can be provided as target outputs to the training system. During the training phase, the debiasing engine 222 can select the type of machine learning model to be trained, e.g., pick a pre-defined or default type of machine learning model, or analyze the input features and target outputs to identify a specific type of machine learning model. For example, the type of machine learning model can include gradient boosting tree models, generalized linear models, support vector machines, decision tree models, or neural network models such as multi-layer perceptrons (MLPs). Machine learning training algorithms such as minimizing error, computing gradients, or performing backpropagation can be used to train the machine learning model. In some implementations, the training system can use metadata corresponding to the denoised aggregates and event-level report data to preprocess the values of the interaction event features for provision to the training system. For example, by using metadata that identifies the data type for a value, the system can preprocess the value such that the training system can more accurately interpret the value. In other words, the training system can map the conversion value in a cell to an encoded representation that can be used as an input feature for the training of the machine learning model. For example, the system can convert each pair of denoised aggregates and event-level report data into a predicted value for the evented logs (such as for use cases that can be identified by the interaction usage engine 224) in a format that can be interpreted by the content provider system 206 and / or the asset provider system 210 (such as Figure 6A as shown).

[0078] The debiased data is sent to the interaction usage engine 224 and optionally to the content provider system 206 and the asset provider system 210. The interaction usage engine 224 can process the debiased data to identify the use cases associated with the debiased data. The interaction usage engine 224 can send the use cases associated with the debiased data or control commands associated with one or more use cases to the content provider system 206 and the asset provider system 210.

[0079] As used in this specification, eventification refers to the process of extracting event-level 3PC explicit action data from event-level reports 228 and aggregated summary reports 230. Eventification has multiple advantages. One advantage of eventification is that, despite using completely different measurement techniques, eventified logs are also structurally similar to event-level data recorded by third-party cookies. The structural similarity facilitates inserting eventified logs into existing data pipelines and other modeling infrastructures built for 3PC data with few changes. The compatibility of eventified logs with existing data pipelines reduces technical debt, thus facilitating the transition to systems including the ARA API. Another advantage of eventification is that the same eventified logs can be used for many use cases, including the top one (or two) main reports and bidding use cases. For reporting, eventified logs can be appropriately aggregated, for example, at the campaign level, according to the slices used by the report. For bidding, eventification facilitates using conversions (or conversion values) as labels and training machine learning models characterized by interactive event features. The training of machine learning models uses a training dataset in units of a combination of interactive events and attributed conversions (or conversion values) as input, matching the structure of eventified logs. Building both reports and bidding based on the same log can reduce processing complexity and automatically provide cross-use case consistency.

[0080] Figure 2B is a reference Figure 2A A block diagram of an example data flow within the secure distribution system 202 described with reference to. According to some implementations of the present disclosure, the secure distribution system 202 is configured as an API denoising layer architecture.

[0081] As referenced Figure 2A As described, metadata 226, event-level reports 228, summary reports 230, and interaction data 232 can be processed in parallel by computing components to generate eventified debiased data 244 that can be used to determine interaction use cases 246.

[0082] Metadata 226 can be formatted as a table, where entries include context data corresponding to user interactions recorded by the API. Metadata 226 can be processed to generate mapped metadata 234.

[0083] The event-level reports 228 generated by the API include truncated raw event-level reports that can be assigned to one or more buckets. The event-level reports 228 can be processed to distinguish between events corresponding to false branches and true branches and to generate event API preprocessed data 236.

[0084] The summary report 230 transmitted by the API includes an aggregated summary report 238 generated based on the original summary report and added noise (false reports). The aggregated summary report 238 can be processed to estimate interaction parameters 242.

[0085] Interaction data 232 can be extracted and collected from different channels configured to generate interaction logs. The interaction data 232 can be processed to generate an aggregated interaction 240. The aggregated interaction 240 can also be used to determine the estimated interaction parameters 242.

[0086] The estimated interaction parameters 242 can be processed to generate debiased event data 244. The debiased event data 244 includes interaction data extracted from the aggregated summary report 238 and the event-level report 228 based on underlying 3PC transformation applied to protect user data privacy. The debiased event data 244 has the same structural type as the original interaction data, thus lacking private user information or having private user information replaced with generic user information. Without violating the privacy and security measures imposed by the system privacy settings, the debiased event data 244 can be processed to generate interaction use case data 246.

[0087] Figure 3 is a flowchart of an example process 300 according to some implementations of the present disclosure. The example process 300 can be performed using any components of the example system 100 described, for example, with reference to Figure 1 or the example system 200 described with reference to FIG. 2. The operations of the process 300 are described below for illustrative purposes only. The operations of the process 300 can be performed by any suitable device or system, such as any suitable data processing device. The operations of the process 300 can also be implemented as instructions stored on a non-transitory computer-readable medium. Executing the instructions causes one or more data processing devices to perform the operations of the process 300.

[0088] By one or more processors of a computing device (e.g., the server system 102 described with reference to Figure 1 or the secure distribution system 202 described with reference to FIG. 2) from a client device (e.g., the Figure 1Receives an aggregated summary report (302) via the API of the client devices 104, 204 described in and FIG. 2. The aggregated summary report includes data aggregation items. The data aggregation items include groupings of event attribution explicit action data aggregated to the slice level. The aggregated summary report is configured based on predefined slices through which the interaction provider system plans to understand conversion activities. For example, the aggregated summary report can be configured to provide information focused on answering specific questions such as "How many conversions are there in a particular region?" or "What is the total value of purchases yesterday?" The type of information used for aggregation is particularly useful for reporting use cases through which service (e.g., testing or advertising) provider systems can gain insights into the interaction ad campaigns provided, rather than based on individual clicks (or individual views). Although the aggregated summary report is structured to provide aggregated data associated with a specific topic, the underlying raw aggregated summary report can provide additional information beyond the initial aggregation scope. To extract the additional information, the aggregated summary report can be formatted relative to the applied structure configuration. The configuration of the aggregated summary report includes adjustment of the limits of the aggregated summary report. The limits in the aggregated summary report can be configured by spreading the per-interaction sensitivity parameter across a number of conversion distributions that can potentially be achieved by the API to protect the privacy of user data related to interaction events (e.g., interactions with the user interface that displays the interaction). The aggregated summary report provides the type of flexibility that can be used to collect attribution configurations.

[0089] Identify and remove false positives (304). False positives can be determined by analyzing whether event-level reports are available. The "hierarchical" structure includes the final arrangement of aggregation items in a tree structure, where each additional key splits a parent "leaf" into child "leaves". The hierarchical structure of the aggregated summary report includes slices of aggregation items corresponding to parent event nodes and child event nodes, where an irrelevant branch including one or more nodes can correspond to a false positive report. A false positive report is a slice of aggregation items that has been determined to have been artificially added to the aggregated summary report and is unrelated to any conversion of real user interactions. The results in a slice of aggregation items that erroneously appear to include a conversion define the false positives that can be identified based on the availability of event-level reports. The structural pattern (aggregation key) used for the aggregation items can be stored in a database (e.g., the digital component repository 212 described in reference to FIG. 2). For example, the ARA used to generate the aggregation structure can be provided by a content provider system (e.g., reference Figure 1The content provider systems 106, 206 described in FIGS. 1 and 2 pre-register before receiving the aggregated summary report. In some implementations, the content provider system may register multiple aggregated item keys that can be used for conversion. False positive identification may include processing the aggregated summary report using each of the aggregated keys being stored. In response to identifying false positives in the configured aggregated summary report that represent aggregated items full of errors, these false positives are removed from the data. Multiple techniques can be implemented to filter out false positives. In some implementations, event-level reports can be used to reduce false positives. For example, if considering click-through conversion (CTC) settings, each click can record three conversions within three time windows, each conversion corresponding to three metadata bits. If the aggregated summary report indicates that a click led to a conversion, but the event-level report reports that there was no conversion, the corresponding click may be a false positive on the noisy branch.

[0090] Determine and reduce noise to generate the original aggregated summary report (306). The noise identification procedure may be based on the hierarchical structure of the configured aggregated summary report. (For example, via an API) Add noise to the aggregated item data to protect user privacy, but the nature of the noise is related to the reference Figure 5Differences in the event-level reports depicted in FIGS. 0 to 7. The event-level reports use a local-differential privacy (DP) model to add noise, while the differential privacy mechanism in the aggregated summary reports is a centralized-DP model. Local-DP means that each interaction is processed with noise; centralized-DP means that noise is added only after some aggregation has occurred. Different from the noise addition mechanism in the event-level reports (randomized response), the aggregated summary reports use a Laplace noise addition mechanism. The noise added using the Laplace noise addition mechanism includes a random variable from the Laplace distribution that is added to the aggregated terms. For example, the event-level attribution conversion after an interaction event is truncated by a contribution bound; the truncated conversions are aggregated to the slice level; and Laplace noise is added to the aggregated term slices and reported. The aggregated summary reports do not use a specific structure nor a predefined key set for the aggregated terms; this is resolved by the content provider system. The magnitude of the noise added to the aggregated terms is technically fixed and is drawn from a Laplace distribution with a mean of 0 and a standard deviation of 2L1 / ∈, where L1 and ∈ are the sensitivity of the data and the desired privacy level, respectively. Noise can be reduced from all slices of the aggregated summary reports. The statistical distribution of the data can be determined to identify potential noise. For example, an aggregated term of a specific entry type (e.g., campaign level) can be compared to the sum of conversions across all interaction groups in the corresponding entry type (campaign). If the comparison indicates that the measurements correspond to approximately the same quantity, the aggregated term of the entry type in “A” can be combined with the aggregated terms in the sum of conversions across all interaction groups in the corresponding entry type “B” + “C” to improve the estimation of the quantity. In some implementations, the Laplace noise centered at 0 and additionally added to the aggregated summary reports can be reduced by taking a linear weighted average of “A” and “B” + “C”, which would represent an unbiased estimate of the conversion quantity. Another example method for reducing noise is to apply a skewed weighted average, which minimizes the final variance of the parent node estimate. The result is an improved estimate of the conversion (configuration) count included in the configured aggregated summary reports. Note that the skewed weighted average method can include a second “top-down” pass to pass information down the tree, ultimately benefiting the finest-grained aggregated terms. The top-down method looks at the difference between a slice and its children and spreads that difference across the child node distribution. The described noise reduction process provides “consistency” because the conversion count of each slice is exactly equal to the sum of its children; where the original API output does not share this property.

[0091] Determine interaction use cases (308) for the original aggregated summary report. For example, data content mapping can be used to identify digital components associated with statistics derived from the original aggregated summary report and obtained from a database. The mapped digital components can be stored electronically in a physical memory device as a single file or as a collection of files (such as video files, audio files, multimedia files, image files, or text files), and include advertisement information, such that the interaction is of the digital component type. For example, the digital component can be content that is intended to supplement the content of a web page or other resource presented by an application executed by a client device. More specifically, the digital component can include digital content related to the resource content (e.g., the digital component can relate to the same topic as the web page content or be related to a relevant topic). The configuration of the digital component can supplement and generally enhance the web page or application content provided to access the asset providing system.

[0092] Generate a trigger (310) to activate the operation of the asset providing system using the interaction use case. The trigger can automatically activate the execution of one or more operations corresponding to the determined interaction use case. The operations can include establishing a communication channel with the client device, transferring the digital component from the database to the client device, and / or transferring an offer corresponding to the digital component from the asset providing system to the client device. The operations can include automatically modifying the display of the client device to increase the visibility of the automatically triggered digital component display.

[0093] Figures 4A to 4C is a block diagram showing an example of an aggregated summary report according to the described implementation. Figure 4A Shows an example of an aggregated summary report 400A that can be generated by an API. Figure 4B Shows example aggregated summary report data 400B that can be included in an example aggregated summary report 400B. Figure 4C Shows an example of a denoised aggregated summary report 400C.

[0094] As Figures 4A to 4CAs shown, the aggregated summary report is structured hierarchically to include one or more parent nodes 402, 404 and one or more child nodes 406, 408, 410. Each parent node 402 or 404 can respectively have one or more child nodes 406, 408 or 410. Each node of a particular type can include a set of data types. For example, the parent nodes 402, 404 can include keys, conversion (configuration) types, and aggregated conversion (configuration) counts. Each child node of a particular type can include a set of data types. For example, the child nodes 406, 408, 410 can include the data type sets of the corresponding parent nodes and one or more additional data. In the example shown, the child nodes 406, 408, 410 include keys, conversion types, ad groups, and aggregated conversion counts.

[0095] In the example shown, the aggregated summary report 400A is aggregated by "keys" (e.g., campaign ID, ad group, conversion type), specific instances of these keys are "key - value" (campaign ID == 1, ad group == 1, conversion type == "sale"), and the aggregated item data corresponding to a specific key - value is an "aggregated item". The data reported in each box in the aggregated summary report 400A is an "aggregated item" representing a specific combination of event and conversion characteristics. The aggregated summary report 400A provides a data set whose basic granularity is at the aggregated item level. The aggregated items reported by the aggregated item - API are not completely accurate because the aggregated summary report 400A includes statistical noise added to the counts for differential privacy reasons. The data in the aggregated summary report 400A has been aggregated in a specific way: first, by campaign, then by conversion type; and then by interaction group. The details of the aggregated items and how the aggregated items can be organized are controlled by the content provider system. The content provider system can choose to aggregate the data in a way that can be mapped to specific actions of the interaction use case. Due to the introduction of the noise - adding mechanism to protect user data privacy, the aggregated summary report can report results even in the absence of actual conversions. In the example shown, the two boxes on the right report some conversions for campaign 2, where in fact there are no conversions.

[0096] The aggregated summary report 400A can be configured to comply with the following set limits: the number of contributions that can be registered for representation in the aggregated summary report for a given interaction. Event - level conversions after an interaction event can be truncated by the set limits, and the aggregated items can be based on the truncated conversions. If the conversion activity exceeds the limits, these limits will manifest as overly small aggregated items. The content provider system can configure the limits in the aggregated summary report by potentially distributing the per - interaction L1 sensitivity parameter across a number of conversions. The aggregated summary report can provide a type of flexibility that can be used to better capture the attribution picture.

[0097] In the example shown, the aggregated summary report data 400B includes two keys: a campaign ID and an interaction group ID. The hierarchical structure allows some flexibility in the content provider system. When the number of conversions is low, the noise added by the aggregated summary report can overwhelm the true conversion count in its aggregates, where collecting aggregates at a coarser level at the top of the hierarchical structure by the content provider system can make data collection and processing more efficient in terms of system resources. On the other hand, when the number of conversions is high, the impact of the noise can be relatively low compared to the true conversion count in the aggregates, and it makes more sense to collect aggregates at a finer-grained level in the lower part of the hierarchy. Using the hierarchical structure provides flexibility to the content provider system to fine-tune the trade-off between the information and accuracy in its aggregates depending on the specific situation. The hierarchical structure can help efficiently combine multiple levels of information to improve the quality of aggregates throughout the tree. The hierarchical structure can be used to identify true positive data 412 and false positive data 414 and separate them, and the false positive data can be reduced to generate the denoised aggregated summary report data 400C in which the false positive data 414 has been reduced.

[0098] In Figures 4A to 4C the example shown, the "slices" 402, 404, 406, 408, 410 correspond to data aggregated within a specific node or key-value combination on the aggregate tree. Although the aggregated summary report 400A shown includes a hierarchical structure, other aggregation techniques can be used when considering other types of aggregations such as total conversion value.

[0099] Figure 5 is a flowchart of an example debiasing process 500 according to some implementations of the present disclosure. Any component of the example system 100 described with reference to Figure 1 or the example system 200 described with reference to FIG. 2 can be used to perform the example process 500. The operations of the process 500 are described below for illustrative purposes only. The operations of the process 500 can be performed by any suitable device or system, such as any suitable data processing device. The operations of the process 500 can also be implemented as instructions stored on a non-transitory computer-readable medium. Executing the instructions causes one or more data processing devices to perform the operations of the process 500.

[0100] Event-level reporting (e.g., with reference to Figure 6BThe example event-level data 600B) described is received (502) by a computing device from an API of a client device. The term "event" in the event-level report corresponds to interaction events generated according to various conversion (configuration) types. The granularity of the event-level report is defined by interactions such as clicks or views. As the interaction progresses, the API returns information about any conversions that may occur (or not occur) within a predefined time after the interaction. Data from the event-level report can be paired with metadata corresponding to the respective conversions. To protect user privacy, the API does not return event-level data with full fidelity (no deviation). A small portion of the ad interactions are randomly selected by the API to be assigned random conversion (configuration) metadata. In some implementations, a limit can be set on how much metadata can be extracted from a conversion.

[0101] Process the event-level report to identify invalid metadata to identify events that are determined to be on a false branch (504). The event-level report can be processed using metadata filters applied by a metadata mapping engine (e.g., the metadata mapping engine 214 described in reference Figure 2A ). The event can be filtered based on metadata entries that can be mapped to the event in the event-level report. If the metadata for an identified log entry is not registered in the metadata mapping table, it is determined that the conversion (configuration) of the log entry is on a false branch, and the corresponding metadata is identified as invalid for each conversion type identifier (CTID), facilitating the identification of interaction events on the false branch. Interaction events on the false branch can be reduced to improve the signal-to-noise ratio and obtain a more accurate estimate. Invalid events are characterized by 3-bit conversion metadata not registered by a given CTID (or 1-bit for EVC / VTC) (see Identifying Clicks on the False Branch for more information). Invalid interaction events can be reduced from the event-level report.

[0102] For events not deterministically identified as being on a false branch, estimate the probability P(false|y i ) of each event being on a false branch (506). The term y i is a vector of the conversion counts reported by the API for event i. Note that the probability that event i is on the true branch is:

[0103] P(true|y i ) = 1 - P(false|y i ).

[0104] The probability estimation includes determining an estimate μ kw of the expected conversion count for each conversion type k and window w, where k = 1,..., K and w = 1,..., W (assuming there are a total of K types and W windows). Considering that it can be truncated at n conversions, determine the expected total conversion count, μ kw (Nkw μ(N≥n) is defined as μ(N≥n)=E(N|y=n,y=3).

[0105] If the conversion type k in window w is "before truncation", then

[0106]

[0107] If the conversion type k in window w is being truncated, then

[0108]

[0109] If the conversion type k in window w is "after truncation", then

[0110]

[0111] The parameters can be empirically estimated by combining event-level API outputs, aggregated item-level API outputs, and ad event log data without relying on any other data sources (such as 1PC data).

[0112] The probability estimate is as follows:

[0113]

[0114] The term ω is a specific configuration in the set of all possible conversion configurations for events that the event-level API can report (Ω), Q is the total number of ad events with CTID, p is the probability that an event noise eventually appears on the false branch of the API, p(ω) is the probability of drawing ω given that the event is on the false branch, and C(ω) is the total number of events with CTID reported by the API with configuration ω. Information can be retrieved from the API configuration. The numerator is the expected number of interaction events that end up on the false branch and have the conversion configuration ω, and the denominator is the total number of events reported by the API with configuration ω, which includes events on both the false branch and the true branch. The ratio is defined as the probability that an event is on the false branch given that ω is observed. Note that no specific distribution (e.g., Poisson distribution) for conversions on the true branch needs to be assumed, which may be inaccurate. Instead, the conditional probability can be directly estimated by leveraging knowledge of the noise mechanism of the event-level API.

[0115] The debiased total conversion count Conv for each conversion type and window can be obtained from the debiased aggregated item-level API output for each conversion type and window kw,agg , to estimate the unconditional mean. Dividing the debiased total conversion count by the total number Q of events from the interaction event log gives the conversion rate μ kw .

[0116] μ kw (Nkw ≥ n), where n = 0, 1, 2, 3

[0117] Use unconditional means to determine the truncated mean (508). The truncation estimation process ensures that the debiased event-level API counts match the Conv kw,agg matches, which significantly simplifies the design and implementation of the Newton Common Data Layer by naturally incorporating the described merging steps. The assumption is

[0118]

[0119] where

[0120] The assumption states that the ratio of the truncated means at different truncation points for each conversion type k and window w is the same as the corresponding ratio for the overall conversion count. By making this assumption, there is no longer a need to make assumptions about the distribution of the underlying true conversions. Instead of using 1PC data to determine the truncated events, the event API output is used, thus bypassing the dependence on 1PC data that may not be used. The probability P(N kw = n), n = 0, 1, 2, 3 cannot be directly estimated from the event API output, but the probability P(N = n), n = 0, 1, 2, 3 can be directly determined without assuming the true conversion distribution.

[0121] Truncation occurs due to the application of data collection limits to the generation of event API reports. Depending on different priority policies, the way of detecting the truncation window can vary. For reports without conversion priority, data collection (MPC) limits are applied based on the conversion times across all reporting windows. To detect the truncation window for a case, the last reporting window with a conversion report will be the truncation window. For reports with conversion bid-ability priority, MPC limits are applied based on the conversion bid-ability added to the conversion times across each reporting window. The truncation window can be split into 6 buckets for the case of 3 reporting windows and 2 buckets for the case of 1 reporting window. Each window can have bid-able and non-bid-able buckets. Determine whether the truncation occurs in the bid-able bucket or the non-bid-able bucket. If the current truncation window has non-bid-able conversions, the truncation occurs at the non-bid-able bucket. (Non-bid-able > 0). If the current truncation window has only bid-able conversions, the truncation occurs at the bid-able bucket. (Non-bid-able = 0)

[0122] Determine parameters for each event (510) based on a truncated window. The truncated window is used to estimate parameters for each event and apply a debiasing formula to event API data. For each event i, each window w, and conversion type k, the expected conversion count on the true branch is determined by applying a set of selection criteria:

[0123] If y i = ω ∈ Φ kwn,untrunc , then

[0124] If y i = ω ∈ Φ kw0,ptrunc , then

[0125] If y i = ω ∈ Φ kwn,trunc , then

[0126] The expected conversion count for each conversion type k and window is:

[0127]

[0128] The input data for the event API debiasing layer can be event API data, event API configuration, interaction data from the combined logs, aggregated API data. The input data can be used to estimate parameters to debias the event data.

[0129] Example process 500 presents several advantages. One advantage is that mild assumptions are applied to the distribution of true conversions (e.g., no need to assume a Poisson distribution). Example process 500 utilizes information from the event-level API, aggregated-item-level API, and interaction event logs without relying on any other data sources (e.g., 1PC data). Example process 500 can be applied when the coverage of 1PC tracking is low, non-existent, or even phased out in the future. Robust to configuration changes: Example process 500 can be applied to any configuration regardless of the specific choice of parameters (b, c, w, p). Even if the noise applied in the future changes the parameters, or different parameters can be provided for different service providers (e.g., advertisers), example process 500 remains valid. Since P is uniformly applied to all interaction events noise, as long as they have a reasonably large number of interaction events, P(false|y) can be estimated at a finer-grained level (such as the campaign or ad-group level). Therefore, an increase in p can have a relatively small impact on the estimate. noise. Example process 500 no longer relies on aggregated-item API data to estimate p(false|y), and increasing the noise level of the aggregated-item API data has no impact on the estimate. Thus, high noise becomes acceptable, and the browser can achieve better privacy guarantees while maintaining reasonable utility of the data. Concise and coherent: Example process 500 greatly reduces the complexity of computing p(false|y) and naturally matches the debiased event-level conversion counts with the debiased aggregated-item-level counts, which can simplify the original design of the Newton common data layer by incorporating steps to combine the debiased data from the two APIs. Example process 500 can be enhanced with a metadata filter: Identifying invalid metadata for each CTID allows identifying interaction events on the false branch to improve the signal-to-noise ratio and obtain even more accurate estimates.

[0130] Figure 6A is a block diagram of example event-level API input data 600A according to some implementations of the present disclosure. The data from the event-level report is "event-level" in the sense that it receives a record for each interaction and pairs that record with metadata corresponding to subsequent conversions. Example event-level API input data 600A can be 3PC explicit action data that forms the input into the API. In Figure 6A it, the series of boxes on the left represents interactions (e.g., interaction views), while the series of boxes on the right represents conversions that may or may not occur after the interaction presented (displayed) by the client device. Example event-level API input data 600A includes the depicted 15 interaction events (view IDs: 1 to 15 602A to 602E). The interaction events correspond to several interaction event characteristics, such as which campaign and interaction group they correspond to.

[0131] AsFigure 6A As shown in Figure 6A , the series of boxes on the left can include multiple campaigns 604A - 604E and interaction groups 606A - 606E (2 campaigns and 3 interaction groups). Conversions have their own metadata, such as conversion types 608A - 608G (such as, sale or purchase) and conversion values 612A - 612G (in US dollars). If a conversion occurs, it is attributed to the corresponding interaction view. Each of the views and conversions has complete metadata, as all information about each is known. For example, the second view comes from a specific campaign 610A - 610G (e.g., Campaign 1, Interaction Group 2, and generates 2 conversions), each conversion is a sale (purchase) event, and the total value is $10. The example event - level API input data 600A includes multiple views that do not correspond to any conversions, as indicated by the lack of conversions for views 4 through 15. The event - level report is generated by an API that uses the information described as input, but only reports a transformed version of the data back to the content provider system, as referenced Figure 6B as described.

[0132] Figure 6B is a block diagram of an example event - level report 600B according to some implementations of the present disclosure. The example event - level report 600B can be generated by an API that truncates the event - level API input data (e.g., the example event - level API input data 600A as referenced Figure 6A as described) and the corresponding metadata.

[0133] For example, view IDs: 1 602A - 602C do have 4 attributed conversions, but the event - level report only includes 1. Also, the series of boxes on the right representing conversions do not have all the information about the conversions, but only have information about the conversions that fall into "bucket 0" (the definition of "bucket" is controlled by the content provider system). On the other hand, different from the conversions on which the metadata is reduced, the complete metadata can be retained and mapped to the example event - level report 600B. In the Figure 6B example shown, the API correctly presents where no conversions occurred - namely, on views 4 through 15.

[0134] For example event-level report 600B, up to 3 conversions can be registered for a single interaction click; and up to 1 conversion can be registered for a single interaction view. In example event-level report 600B, "View ID: 1" is shown as including only 1 conversion, while in fact the event data includes 4 conversions. Example event-level report 600B is not susceptible to truncation, indicating that fusing these two reports can be a promising method for "filling" such truncation gaps in a post-processing step. The data from the event-level report is at the granularity of useful attributes for accurate conversion attribution. To protect privacy, the granular attribution is noise-added. Statistically, the noise addition is achieved via a randomized response mechanism. The randomized response mechanism implemented in example event-level report 600B involves a binary selection process with two different possibilities: (1) reporting the true conversion activity, or (2) reporting random conversion activity.

[0135] Figure 7A is a block diagram showing an example of a true-branch and noise-branch discrimination mechanism 700A according to some implementations of the present disclosure. The branch discrimination mechanism 700A uses the example event-level report 600B described Figure 6B to present.

[0136] For each event 702 (interaction), it can randomly select which of the two branches to use with a certain probability. The true-branch and noise-branch discrimination mechanism 700A includes a classification between a "true branch" 704A and a "noisy branch" 704B. The probability 706A that any given interaction falls on the noisy branch is small, while the probability 7068 that any given interaction falls on the true branch is high.

[0137] The event-level report uses a randomized response model. Given a fixed value p706A, the event-level report implements randomized response by randomly selecting with probability p706A, 706B whether each registered interaction event 702 will be on the true branch 704A or the noisy branch 704B (note: p is known to the content provider system, but the content provider system does not know whether the interaction event in the event-level report is on the true branch).

[0138] If the interaction event 702 is on the true branch 704A, the number of conversions reported by the event-level report is truncated 708A to the first c conversions, and the metadata of such conversions is also restricted - only some bits of the metadata can be encoded and revealed to the content provider system. Also, the accurate conversion time is not shown, but is grouped into time windows corresponding to multiple clicks 710A and views 710B. The API can send a conversion report containing such information to the content provider system.

[0139] If the interaction event 702 is on the noisy branch 704B, any information about true conversions or lack thereof will be suppressed by the noise 708B, and instead the API can randomly draw from a pool of possible configurations (where a configuration is a particular implementation of a conversion result) and generate conversion reports similar to those on the true branch 704A. Thus, the content provider system cannot tell whether the conversion reports sent from the API are on the true branch 704A or the noisy branch 704B.

[0140] Figure 7B Is a diagram of the event debiasing mechanism 700B according to some implementations of the present disclosure.

[0141] The event debiasing mechanism 700B uses a debiasing transformation (essentially a function) that utilizes information from the API and produces an unbiased estimate E[n i of the true conversion count. The estimation parameters are for use in obtaining data-based "summary statistics" that form the input to the debiasing transformation. The debiasing transformation constructs an estimator for the conversion corresponding to each advertising event i in the event-level report as follows:

[0142]

[0143] The terms "noisy branch" and "true branch" refer to whether the event to which the API attributes a conversion is on the noisy branch or the true branch. P(noisy branch) refers to the probability that an event is on the noisy branch 704B, and is equivalent for the true branch 704A. Each event is assigned to the noisy branch 704B or the true branch 704A with a certain probability.

[0144] When event denoising, the y registered conversions 714 corresponding to event i 712 are included in the event-level report. To estimate the true conversion for event i 712, the realization of the y registered conversions 714 is determined by the following two conditions: the probability that event i is on the noisy branch 704B or the true branch 704A of the API and what its true conversion could be. In the binary randomized response of the event-level report, each interaction event 712 is either on the noisy branch 704B or the true branch 704A, but the content provider system does not know the exact state. The probabilities that event 702 is on the true branch and on the noisy branch are Pr(true|y) and Pr(noisy|y). The event i 712 can be determined using the y registered conversions 714 (based on the conditional probabilities on the branches after the i orange circle). Simply put, estimating one conditional probability gives the other because Pr(true|y i ) + Pr(noisy|y i ) = 1.

[0145] Determine the estimate of the conversion on each branch. If the interaction event is on a noisy branch, there is no relevant information from the API, and given the event on the noisy branch 704B, the best estimate of the number of conversions is only E[n i , i.e., the average conversion rate (to be estimated again). If the interaction event is on the true branch, given that the event-level report reports at most 1 conversion after one view, the conversion from event i is considered to be truncated. Assume the event is on the true branch,

[0146] E[n i |y i = E[n i |n i ≥1].

[0147] The denoising transformation formula for interaction views can be used to generate denoised event data. Except for the term E[n i |y i , the same processing can be done for the case of interaction clicks (described in reference Figure 6A and Figure 6B ). For a given example, at most 3 conversions can be reported for a click: if y i <3 from the event-level report, it can be determined that the conversion for event i has not been truncated, and E[n i |y i = y can be used in the true branch formula. If y≥3 in the event-level report, truncation of the conversion for event i may have occurred, and the true branch formula E[n i |y i = E[n i |n i ≥3] can be used.

[0148] Figure 8 is a flowchart showing an example process 800 according to some implementations of the present disclosure. Any components of the example system 100 described in reference Figure 1 or the example system 200 of FIG. 2 can be used to execute the example process 800. The operations of the process 800 are described below for illustrative purposes only. The operations of the process 800 can be performed by any suitable device or system, such as any suitable data processing device. The operations of the process 800 can also be implemented as instructions stored on a non-transitory computer-readable medium. Executing the instructions causes one or more data processing devices to perform the operations of the process 800.

[0149] The event-level report (e.g., reference Figure 6BThe example event-level data 600B) described is received and processed (802) by a computing device from the API of a client device. The event-level report includes filtered interaction event data corresponding to interaction events generated according to multiple conversion types, as referenced Figure 6A and Figure 6B described, and multiple conversion types can be reported in a truncated format. The data included in the event-level report can be paired with metadata corresponding to the respective conversions. To protect user privacy, the API does not return event-level data with full fidelity. A small portion of the ad interactions are randomly selected by the API to be assigned random conversion metadata. In some implementations, a limit can be set on how much metadata can be extracted from a conversion.

[0150] The event-level report is processed to identify invalid metadata to remove events that are determined to be on a false branch (803). The event-level report can be processed using metadata filters applied by a metadata mapping engine (e.g., the metadata mapping engine 214 described in reference Figure 2A ). The event can be filtered based on the metadata entries that can be mapped to the event in the event-level report. If the metadata for an identified log entry is not registered in the metadata mapping table, then the conversion (configuration) of that log entry is determined to be on a false branch, and the corresponding metadata is identified as invalid for each conversion type identifier (CTID), facilitating the identification of interaction events on false branches. Interaction events on false branches can be reduced to improve the signal-to-noise ratio and obtain a more accurate estimate. Invalid events are characterized by 3-bit conversion metadata not registered by a given CTID (or 1-bit for EVC / VTC) (see Identifying Clicks on False Branches for more information). Invalid interaction events can be reduced from the event-level report. For events not deterministically identified as being on a false branch, the probability P(false|y i ) of each event being on a false branch is estimated. The processing can also include debiasing, as described in reference Figure 5 .

[0151] One or more processors of a computing device (e.g., the server system 102 described in reference Figure 1 or the secure distribution system 202 described in reference to FIG. 2) receive and process an aggregated summary report (804) from the API of a client device (e.g., the client devices 104, 204 described in reference Figure 1 and FIG. 2). The aggregated summary report includes data aggregation items. The data aggregation items include groupings of event attribution explicit action data aggregated to the slice level. The aggregated summary report is configured based on a predefined set of slices through which the interaction provider system plans to understand conversion activity. The processing can include debiasing, as described in reference Figure 3 .

[0152] The first step in combining the event-level reports and the aggregated summary reports is to estimate better aggregates based on the noisy data that is sent from the aggregated summary reports to the content provider system. The step is essentially about reducing the impact of the noise and is achieved by "post-processing" the aggregated summary report data. The post-processing is motivated by the fact that if the aggregates are used as a calibrated form of the event-level report data, it is important to ensure that these aggregates are as accurate as possible. To obtain improved aggregates from the post-processing, two important components need to be highlighted: handling "false positives" and leveraging the hierarchical nature of the aggregates. False positives in the aggregated summary reports. One of the main ways in which the statistical noise in the aggregated summary reports affects the data quality is via "false positives". A false positive is a slice of an aggregate that has not actually converted, but for which the aggregated summary report adds Laplace noise to the 0 in its result. The result in the slice appears to have converted, but in reality, it has not.

[0153] Identify and reduce false positives (806). False positives can be determined by leveraging the hierarchical nature of the aggregated summary reports. The hierarchical structure of the aggregated summary reports facilitates the efficient and flexible processing of the aggregated summary reports to identify use cases. The "hierarchical" structure includes the final arrangement of the aggregates in a tree structure, where each additional key splits a parent node "leaf" into child node "leaves". The hierarchical structure of the aggregated summary reports includes slices of aggregates corresponding to parent event nodes and child event nodes, where an irrelevant branch that includes one or more nodes can correspond to a false positive event. A false positive is a result of the structure of the aggregated summary reports, where the keys based on which the ARA has completed the aggregation must be pre-registered by the content provider system before the data has reached the content provider system. In this setting, the content provider system can register aggregate keys that may or may not have conversions, resulting in a large number of key-value pairs that include false positives. A false positive event is a slice of an aggregate that has been determined to have been artificially added to the aggregated summary reports and is not related to any conversions of real user interactions. The incorrectly appearing result of a conversion in the aggregate slice defines a false positive that can be identified based on the structure of the aggregated summary reports. The structural pattern (aggregation key) used for the aggregates can be stored in a database (e.g., the digital component repository 212 described with reference to FIG. 2). For example, the ARA used to generate the aggregation structure can be provided by the content provider system (e.g., with reference Figure 1The content provider systems 106, 206) described in FIGS. 1 and 2 pre-register before receiving the aggregated summary report. In some implementations, the content provider system may register multiple aggregated item keys that can be used for conversion. False positive identification may include processing the aggregated summary report using each of the aggregated keys being stored. In response to identifying false positives in the configured aggregated summary report that represent aggregated items full of errors, these false positives are reduced from the data. Multiple techniques can be implemented to filter out false positives. In some implementations, event-level reports can be used to reduce false positives. For example, if click-through conversion (CTC) settings are considered, each click can record three conversions within three time windows, with each conversion corresponding to three metadata bits. If the aggregated summary report indicates that a click led to a conversion, but the event-level report reports no conversion, the corresponding click may be a false positive on a noisy branch. Identifying an entry in the configured aggregated summary report as being on a noisy branch, the entry can be randomly attributed to a single bucket with no conversion. If the probability of an entry determined to be a false positive exceeds a set threshold (e.g., the chance that the entry will correspond to some real conversion is below an acceptable threshold), the aggregated item slice including the false positive is reduced, thereby helping to improve data quality.

[0154] Determine and reduce noise to generate the original aggregated summary report (808). The noise identification procedure can be based on the hierarchical structure of the configured aggregated summary report. Noise is added to the aggregated item data, but the nature of the noise is related to the reference Figure 5Differences in the event-level reports described up to Figure 7. The event-level reports use a local-differential privacy (DP) model to add noise, while the differential privacy mechanism in the aggregated summary reports is a centralized-DP model. Local-DP means that each interaction is processed with noise; centralized-DP means that noise is added only after some aggregation has occurred. Different from the noise-adding mechanism in the event-level reports (randomized response), the aggregated summary reports use the Laplace noise-adding mechanism. The noise added using the Laplace noise-adding mechanism includes a random variable from the Laplace distribution that is added to the aggregated terms. For example, the event-level attribution conversion after an interaction event is truncated by the contribution boundary; the truncated conversions are aggregated to the slice level; and Laplace noise is added to the aggregated term slice and reported. The aggregated summary reports do not use a specific structure nor a predefined key set through which to aggregate; these decisions are handled by the content provider system. The magnitude of the noise added to the aggregated terms is technically fixed, which is drawn from a Laplace distribution with a mean of 0 and a standard deviation of 2L1 / ∈, where L1 and ∈ are the sensitivity of the data and the desired privacy level, respectively. The noise can be reduced from all slices of the aggregated summary reports. The statistical distribution of the data can be determined to identify potential noise. For example, an aggregated term of a specific entry type (e.g., campaign level) can be compared with the sum of conversions across all interaction groups in the corresponding entry type (campaign). If the comparison indicates that the measurements correspond to roughly the same amount, the aggregated terms of the entry type in "A" can be combined with the aggregated terms in the sum of conversions across all interaction groups in the corresponding entry type "B" + "C" to improve the estimation of the amount. In some implementations, the Laplace noise centered at 0 and additionally added to the aggregated summary reports can be reduced by taking a linear weighted average of "A" and "B" + "C", which would represent an unbiased estimate of the conversion quantity. Another example method for reducing noise is to apply a skewed weighted average, which minimizes the final variance of the parent node estimate. The result is an improved estimate of the conversion count included in the configured aggregated summary reports. Note that the skewed weighted average method can include a second "top-down" pass to pass information down the tree, ultimately benefiting the finest-grained aggregated terms. The top-down method looks at the differences between a slice and its children and spreads that difference across the child node distribution. The described noise removal process provides "consistency" because the conversion count of each slice is exactly equal to the sum of its children; where the original API output does not share this property. The probability determined to be no conversion according to the event-level reports for each event (e.g., click) can be used to remove the aggregated term fragments for which the event-level reports indicate no conversion (configuration change). The use of the event-level reports in combination with the aggregated summary reports helps improve data quality.

[0155] Statistics are generated (810) by combining a debiased event-level report and a denoised aggregated summary report. In response to determining that noise is reduced from the aggregated items by post-processing the aggregated summary report data, the denoised aggregated items can be used to post-process the event-level report data and create a unified, more accurate eventized log for various use cases. Eventization can include bidding on training a machine learning model based on using a conversion (or conversion value) as a label and characterized by interactive event features. The training of the machine learning model uses a training data set in units of a combination of interactive events and attributed conversions (or conversion values) as input, which matches the structure of the eventized log. Building both reports and bids based on the same log can reduce processing complexity and automatically provide consistency across use cases. For example, a training system can be used to train a machine learning model on the denoised aggregated items and the event-level report data to predict values for the eventized log, resulting in a trained model. The conversion (or conversion value) can be provided as an input feature to the training system, and the interactive event features contain values that can be provided as the target output to the training system. The training system can pick the type of machine learning model to train, e.g., pick a predefined or default machine learning model type, or analyze the input features and the target output according to the event scenario based on the report to identify the type of machine learning model. For example, the type of machine learning model can include a gradient boosting tree model, a generalized linear model, a support vector machine, a decision tree model, or a neural network model, such as a multi-layer perceptron (MLP). Machine learning training algorithms such as minimizing error, calculating gradients, or performing backpropagation can be used to train the machine learning model. In some implementations, the training system can use metadata corresponding to the denoised aggregated items and the event-level report data to preprocess the values of the interactive event features to provide to the training system. For example, by using metadata that identifies the data type for the values, the system can preprocess the values so that the training system can interpret the values more accurately. In other words, the training system can map the conversion values in the cells to an encoded representation, which can be used as an input feature for the training of the machine learning model. For example, the system can convert each pair of denoised aggregated items and event-level report data into a predicted value for the eventized log (such as for a use case), which is in a format that can be interpreted by a content provider system and / or an asset provider system. The event-level data provided by the event-level report data can be processed for event denoising, which includes identifying the noise addition and conversion truncation on these events. Using the conversions reported by the event-level report for each event in the event-level report data, each event is transformed by implementing a denoising transformation. The denoising transformation corrects both the randomized response noise addition implemented by the event-level report and the truncation it imposes on the attributed conversion. The transformation is data-driven and takes as input a set of summary statistic values that index key aspects of the underlying data generation process.Summary statistics are determined based on refined aggregates obtained from an aggregated summary report after post - processing.

[0156] Generating a report is attributed to the integration event - level log (812) of the conversion of each event for which noise and truncation in the event - level report have been corrected. The correction utilizes information from the aggregated summary report by combining information from two APIs that leverage the information content of both the event - level report and the aggregated summary report. When aggregated, the integrated event - level log can be consistent with the post - processed aggregates from the aggregated summary report (this does not guarantee consistency with the original event - level report and the aggregated summary report processed separately as described in references Figure 3 and Figure 5 ).

[0157] Determine interaction use cases (814) for the integrated event - level log. For example, data content mapping can be used to identify digital components associated with the integrated event - level log and obtain the digital components from a database. The integrated event - level log is derived from an original aggregated summary report merged with de - biased event data. The mapped digital components can be stored electronically in a physical memory device as a single file or a collection of files (such as video files, audio files, multimedia files, image files, or text files) and include service (e.g., test or advertisement) information such that the interaction is of the digital component type. For example, the digital component can be content that is intended to supplement the content of a web page or other resource presented by an application executed by a client device. More specifically, the digital component can include digital content related to the resource content (e.g., the digital component can relate to the same topic as the web page content or be related to a relevant topic). The configuration of the digital components can supplement and generally enhance the web page or application content provided for accessing the asset - providing system.

[0158] Generate a trigger (816) to activate the operation of the asset - providing system using the interaction use case. The trigger can automatically activate the execution of one or more operations corresponding to the determined interaction use case. The operations can include establishing a communication channel with the client device, transferring the digital component from the database to the client device, and / or transferring an offer corresponding to the digital component from the asset - providing system to the client device. The operations can include automatically modifying the display of the client device to increase the visibility of the automatically triggered digital component display.

[0159] Figure 9 is a diagram of an example mechanism 900 for parameter learning and data fusion with aggregated summary report data according to some implementations of the present disclosure. For example, the example mechanism 900 can be included in the example process 800 described in reference Figure 8 ).

[0160] Example mechanism 900 indicates an example implementation of a debiasing formula that uses estimated summary statistics from data. The debiasing formula defines key summary statistics as Pr(noisy|y i )、E[n i and E[n i |n i ≥c] for different values of c. The summary statistics can be learned from a combination of the original event-level reporting logs and the post-processed aggregates. Figure 9 The execution order is shown.

[0161] Pr(noisy|y i )902A: is the probability that event 904A will be on the noisy branch given y i transformations reported by the event-level reporting for the corresponding event. By leveraging our knowledge of the API mechanism, the object can be determined directly from the event-level reporting. By both the way of adding noise probability and extracting the noisy branch configuration, the probability of each configuration appearing on the noisy branch can be estimated. The probability can be used to determine the likelihood that each configuration actually appears in the event-level reporting data of the ad technology. If the configuration appears more frequently in the event-level reporting data than expected, then from the perspective of the noisy branch, Pr(noisy|y i ) is small.

[0162] E[n i 902B: is the expected conversion rate 904B, which can be estimated by dividing the total conversion count obtained from the aggregated summary report after post-processing the aggregates by the count of ad events registered in the event-level reporting.

[0163] E[n i |n i ≥c]902C: uses the combination of two API reports. For example, if the event-level reporting hypothetically has only the true branch and implements conversion truncation. In the context example, the conversion count provided by the aggregated summary report is accurate, such that the count of conversions truncated by the event-level reporting can be determined by comparing the difference in the conversion counts reported between the two APIs. The average value 904C of how many conversions are truncated for each truncated ad event can be determined, thereby estimating the object as E[n i |n i ≥c]902C. Overall, the aggregated summary report can be appropriately configured such that it provides a reasonable overall count; and the noisy branch in the event-level reporting can be interpreted to account for the fact that the conversion counts it reports are not entirely accurate.

[0164] Figure 10A is a diagram of an example data flow for generating a false probability estimate according to some implementations of the present disclosure. For example, example mechanism 1000A can be included in the referenceFigure 8 in the described example process 800.

[0165] The filtered event transformation 1002 can be used to generate event information 1004, metadata count 1006, interaction count 1008, Ω interaction count 1010, and interaction data 1012. The metadata count 1006, interaction count 1008, Ω interaction count 1010, and interaction data 1012 can be combined by using the combined mapping data 1014 of the configuration map 1016 to generate mapping information 1018.

[0166] The combined mapping data 1014 can combine the aggregated data information and the interaction data information to determine the transformation using the following formula.

[0167] Cvr = aggregated_conversion_count <display_date X CTID X conversion_metadata X latency_window> / aggregated_interaction_count <display_date X CTID>

[0168] Conversion metadata: The mapping with respect to the conversion type and bidability can be used to generate conversion metadata using the metadata mapping table.

[0169] The latency window including the conversion date and the impression date can be used as a combined key to determine edge cases. For Ctid not in the aggregated data, such as unassigned interactions, set its Cvr to 0. If the aggregated conversion count is negative, set its CVR to 0. For aggregated conversions, there are no interactions from Java Advanced Imaging (JAI), and the click count from the event API data input can be processed to determine the false probability 1020 used to generate the false probability information 1020.

[0170] Figure 10B is a diagram of an example combined truncated mean data mechanism 1000B according to some implementations of the present disclosure. For example, the example mechanism 1000B can be included in the reference Figure 8 described in the example process 800.

[0171] The conversion information 1032, the false probability information 1020, and the aggregated data information 1024 can be used to derive the combined data 1026, which can be used to determine the truncated mean information 1038. Calculating the truncated mean over each slice of conversion metadata X latency window includes determining the overall truncated mean and determining the sliced truncated mean. The following includes example pseudocode for determining the overall truncated mean and determining the sliced truncated mean.

[0172] Calculate the overall truncated mean based on the following formula.

[0173]

[0174] Figure 10C is a diagram of an example combined truncated data mechanism 1000C according to some implementations of the present disclosure. For example, the example mechanism 1000C can be included in the example process 800 described in reference Figure 8 described above.

[0175] API data 1040 and false probability information 1042 are combined for determining event conversion 1044. The total truncated average information 1046 is used to generate an event conversion 1050 with a truncated average. The event conversion 1050 with a truncated average can be used together with conversion information 1048 and aggregated data information 1052 to generate combined information 1054, and the combined information can be processed to determine the truncated average information 1056 using the following formula for conversion:

[0176] Conv kw,agg <Conv kw,untrunc (true) + Conv kw,ptrunc (true) + Conv kw (false). If the aggregated item count on the sliced level is too small, the truncated average can remain as it is.

[0177] The following includes example pseudocode for combined truncated data.

[0178]

[0179] Figure 11A is an example of 3PC explicit action data 1100A according to some implementations of the present disclosure.

[0180] Figure 11B is an example of event-level reporting data 1100B according to some implementations of the present disclosure.

[0181] Figure 11C is an example of an evented log 1100C after event denoising data according to some implementations of the present disclosure.

[0182] Figure 11D is an example of event-level application programming interface output data 1100D according to some implementations of the present disclosure.

[0183] Figure 11E is an example of debiased event-level application programming interface result data 1100E according to some implementations of the present disclosure.

[0184] Figure 11FAn example of the probability estimate 1100E on the false branch according to some implementations of the present disclosure.

[0185] In some implementations, the components of the environments and systems described above can be any computer or processing device, such as, for example, a blade server, a general-purpose personal computer (PC), a workstation, a -based workstation, or any other suitable device. In other words, the present disclosure contemplates computers other than general-purpose computers and computers without a conventional operating system. Further, the components can be adapted to execute any operating system including the following: Mac Java TM 、Android TM 、 or any other suitable operating system. According to some implementations, the components can also include an email server, a web server, a cache server, a streaming data server, and / or other suitable servers or be communicatively coupled thereto.

[0186] The processors used in the environments and the systems described above can be a central processing unit (CPU), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or another suitable component. Generally, each processor can execute instructions and manipulate data to perform the operations of the various components. Specifically, such as in the communication between an external device, an intermediate device, and a target device, each processor can perform functions that are used to send requests and / or data to the components of the environment and to receive data from the components of the environment.

[0187] The components, environments, and systems described above can include one or more memories. The memory can include any type of memory or database engine and can take the form of volatile and / or non-volatile memory including, but not limited to, the following: magnetic media, optical media, random access memory (RAM), read-only memory (ROM), removable media, or any other suitable local or remote memory component. The memory can store various objects or data, including caches, classes, frameworks, applications, backup data, objects, jobs, web pages, web page templates, database tables, repositories of stored entity information and / or dynamic information, and any other appropriate information, including any parameters, variables, algorithms, instructions, rules, constraints, for reference corresponding to the purposes of the target device, the intermediate device, and the external device. Other components within the memory are possible.

[0188] Regardless of the specific implementation, "software" can include computer-readable instructions, firmware, wired and / or programmed hardware, or any combination thereof (transient or non-transient, as appropriate) on a tangible medium that, when executed, is operable to perform at least the processes and operations described herein. In fact, each software component can be written or described in whole or in part in any suitable computer language including: C, C++, Java TM , Visual Basic, assembly language, any suitable version of 4GL, and other languages. The software can instead include many sub-engines, third-party services, components, libraries, etc., as appropriate. Conversely, the features and functions of various components can be combined into a single component as appropriate.

[0189] The device can include any computing device such as: a smartphone, a tablet computing device, a PDA, a desktop computer, a laptop / notebook computer, a wireless data port, one or more processors within these devices, or any other suitable processing device. For example, the device can include a computer that includes: an input device, such as a keyboard, a touch screen, or other device that can accept user information; and an output device or a graphical user interface (GUI) that conveys information corresponding to the components of the environment and system described above, the information including digital data, visual information. The GUI interfaces with at least a portion of the environment and system described above for any suitable purpose, including generating a visual representation of a web browser.

[0190] The foregoing figures and the accompanying description illustrate example processes and computer-implementable techniques. The environment and system described above (or its software or other components) can contemplate using, implementing, or executing any suitable techniques to perform these and other tasks. It will be understood that these processes are for illustrative purposes only, and the described or similar techniques can be performed at any appropriate time, including simultaneously, separately, in parallel, and / or in combination. Additionally, many of the operations in these processes can be performed synchronously, simultaneously, in parallel, and / or in a different order than shown. Moreover, as long as the method remains appropriate, the process can have additional operations, fewer operations, and / or different operations.

[0191] On the other hand, while the present disclosure has been described in terms of specific implementations and generally associated methods, changes and permutations of these implementations and methods will be apparent to those skilled in the art. Accordingly, the above description of example implementations does not limit or constrain the present disclosure. Other changes, permutations, and alterations are possible without departing from the spirit and scope of the present disclosure.

[0192] Various implementations of the present disclosure have been described. However, it will be understood that various modifications can be made without departing from the spirit and scope of the present disclosure. Accordingly, other implementations are within the scope of the following claims.

[0193] In view of the implementations of the subject matter described above, the present application discloses the following list of examples, where a single feature of an example or more than one feature of the examples in combination and optionally in combination with one or more features of one or more other examples are further examples that also fall within the disclosure of the present application. In some implementations, any of the following examples can be performed simultaneously, individually, in parallel, and / or in combination with any other of the listed examples.

[0194] Example 1. A computer-implemented method includes: receiving, from an application programming interface, an aggregated summary report including aggregated event data collected by the application programming interfaces during events corresponding to interactions with a user interface, the aggregated event data being aggregated using a hierarchical structure corresponding to a data type; removing, from the aggregated summary reports, a first portion of the aggregated event data identified as false positives to maintain true positive aggregated event data; at each level of the hierarchical structure, reducing noise from the true positive aggregated event data to generate a denoised aggregated summary report; determining actions using the denoised aggregated summary reports; and providing instructions to an asset provider system to activate at least one of the actions using the denoised aggregated summary reports. As described above, due to anonymization, aggregation, information truncation, and addition of noise, data reported by the API to the content provider system can deviate from explicit action data measured by third-party cookies to protect user privacy. Technologies applied to protect user privacy are configured by API settings. Information derived from API settings is used to generate the denoised aggregated summary reports. In contrast to the described example, data analysis based on third-party cookies generally cannot facilitate such actions in a noisy / privacy-protected state. The processing of the aggregated event data described in this example includes removing private information or replacing it with generic information, facilitating statistical analysis based on interaction data derived from the aggregated event data, and in doing so, protecting user privacy. The denoised aggregated summary reports can be processed to generate interaction use case data without violating privacy and security measures imposed by system privacy settings. In contrast to systems that include third-party Cookies, in cases where specific actions can be blocked due to potential risks to system security, the system privacy settings of the described example promote transparency in actions and data processing, enabling the execution of one or more actions to be automatically activated. Actions can include establishing a communication channel with a client device, transferring digital components from a database to the client device, and / or transferring quotes corresponding to the digital components from an asset providing system to the client device. Actions can include automatically modifying the display of the client device to increase the visibility of automatically triggered digital component displays.

[0195] Example 2. The computer-implemented method as described in the foregoing example, wherein the aggregated summary report includes hierarchically structured event attribution configuration data as nodes distributed across multiple levels.

[0196] Example 3. The computer-implemented method as described in any of the foregoing examples, wherein the event attribution configuration data includes truncated configurations aggregated as data slices at one or more levels.

[0197] Example 4. The computer-implemented method according to any one of the preceding examples, wherein the noise includes Laplace noise added to each of these data slices.

[0198] Example 5. The computer-implemented method according to any one of the preceding examples, further comprising: generating a weighted average of a parent node and a child node to minimize the variance of the estimate of the parent node.

[0199] Example 6. The computer-implemented method according to any one of the preceding examples, wherein processing these aggregated summary reports to reduce the noise includes: applying a denoising transform to correct both the truncated data and the randomized response noise implemented by the Laplace noise.

[0200] Example 7. The computer-implemented method according to any one of the preceding examples, wherein applying the denoising transform includes using a set of summary statistical values of aspects of an index configuration type.

[0201] Example 8. The computer-implemented method according to any one of the preceding examples, wherein these aggregated summary reports include metadata corresponding to the configuration type applied to the structured event attribution configuration data.

[0202] Example 9. A computer-implemented system, comprising: a memory that stores application programming interface (API) information; and a server that performs operations including: receiving, from an application programming interface, an aggregated summary report including aggregated event data collected by these application programming interfaces during events corresponding to interactions with a user interface, the aggregated event data being aggregated using a hierarchical structure corresponding to a data type; removing a first portion of the aggregated event data identified as false positives from these aggregated summary reports to maintain true positive aggregated event data; reducing noise from the true positive aggregated event data at each level of the hierarchical structure to generate a denoised aggregated summary report; determining an operation using these denoised aggregated summary reports; and providing instructions to an asset provider system to activate at least one of these operations using these denoised aggregated summary reports.

[0203] Example 10. The computer-implemented system according to the preceding example, wherein these aggregated summary reports include hierarchically structured event attribution configuration data as nodes distributed across multiple levels.

[0204] Example 11. The computer-implemented system according to any one of the preceding examples, wherein the event attribution configuration data includes truncated configurations aggregated into data slices at one or more levels.

[0205] Example 12. The computer-implemented system according to any one of the preceding examples, wherein the noise includes Laplace noise added to each of these data slices.

[0206] Example 13. The computer-implemented system according to any one of the preceding examples, wherein these operations further include: generating a weighted average of the parent node and the child nodes to minimize the variance of the estimated value of the parent node.

[0207] Example 14. The computer-implemented system according to any one of the preceding examples, wherein processing these aggregated summary reports to reduce the noise includes: applying a denoising transform to correct both the truncated data and the randomized response noise implemented by the Laplace noise.

[0208] Example 15. The computer-implemented system according to any one of the preceding examples, wherein applying the denoising transform includes using a set of summary statistical values of aspects of the index configuration type.

[0209] Example 16. The computer-implemented system according to any one of the preceding examples, wherein these aggregated summary reports include metadata corresponding to the configuration type applied to the structured event attribution configuration data.

[0210] Example 17. A non-transitory computer-readable medium encoded with a computer program, the computer program including instructions that, when executed by one or more computers, cause the one or more computers to perform operations including: receiving, from an application programming interface, an aggregated summary report including aggregated event data, the aggregated event data being collected by the application programming interfaces during events corresponding to interactions with a user interface, the aggregated event data being aggregated using a hierarchical structure corresponding to a data type; removing a first portion of the aggregated event data identified as false positives from these aggregated summary reports to maintain true positive aggregated event data; reducing noise from the true positive aggregated event data at each level of the hierarchical structure to generate a denoised aggregated summary report; determining operations using these denoised aggregated summary reports; and providing instructions to an asset provider system to activate at least one of these operations using these denoised aggregated summary reports.

[0211] Example 18. The non-transitory computer-readable medium according to the previous example, wherein these aggregated summary reports include hierarchically structured event attribution configuration data as nodes distributed in multiple levels, wherein the event attribution configuration data includes truncated configurations aggregated into data slices at one or more levels, and wherein the noise includes Laplace noise added to each of these data slices.

[0212] Example 19. The non-transitory computer-readable medium according to any one of the foregoing examples, wherein the operations further include: generating a weighted average of the parent node and the child nodes to minimize the variance of the estimated value of the parent node.

[0213] Example 20. The non-transitory computer-readable medium according to any one of the foregoing examples, wherein processing the aggregated summary reports to reduce the noise includes: applying a denoising transform to correct both the truncated data and the randomized response noise implemented by the Laplace noise, wherein applying the denoising transform includes using a set of summary statistical values of aspects of an index configuration type, and wherein the aggregated summary reports include metadata corresponding to the configuration type applied to the structured event attribution configuration data.

[0214] Example 21. A computer-implemented method includes: receiving, from an application programming interface, an event-level report including branches that define records of events including interactions with the application programming interfaces, the event-level reports being paired with metadata corresponding to the configuration of the application programming interfaces; performing an identification of portions of the metadata that include invalid metadata; estimating, for each of the branches, a probability of being true or noisy by using the identification of the portions of the metadata that include the invalid metadata, classifying portions of the events as being on true branches; determining configuration parameters for the true branches, the configuration parameters including a number of average configurations for each event; generating raw application programming interface data by applying a debiasing model that uses the configuration parameters by combining multiple reporting windows and configuration types; determining an operation by using the raw application programming interface data; and providing instructions to an asset provider system to activate at least one of the operations by using the raw application programming interface data. To address limitations of event-level deviations generated by third-party cookies, the processing of event-level data described in the present example facilitates statistical analysis based on interaction data derived from event-level data while protecting user privacy. Specifically, the present example describes a process for generating debiased data that includes replacing truncated data with generic data, facilitating the use of the debiased data for statistical analysis. Another advantage of the present example is that the noise level of the aggregated API data (added as Laplace noise to the data aggregation terms) has no effect on the estimation of interaction data based on event-level data. The noise added by the Laplace noise addition mechanism includes random variables from a Laplace distribution added to the aggregation terms. Event-level attribution transformation after interaction events is truncated by a contribution boundary, the truncated transformations are aggregated to a slice level; and Laplace noise is added to the aggregation term slices and reported. Thus, the high noise in the aggregated summary report becomes acceptable after being removed by using a noise identification procedure based on the hierarchical structure of the configured aggregated summary report. The described example provides a concise and coherent process for generating event-level reports by applying a denoising strategy that matches API settings. Another advantage of the present example is that the process for generating debiased data can be enhanced by a metadata filter that is used to identify invalid metadata, which allows the identification of interaction events on false branches to improve the signal-to-noise ratio and obtain even more accurate estimates. In contrast to the described example, event-level reports generated by third-party cookies typically have no mapping to metadata indicating the presence of events available for false event identification.

[0215] Example 22. The computer-implemented method as described in the foregoing example, wherein the configuration parameters include a configuration window and a configuration count for each configuration type.

[0216] Example 23. The computer-implemented method according to any one of the foregoing examples, further comprising: when truncated for each configuration type, determining the truncated average to be the expected total configuration count for each such configuration type and window.

[0217] Example 24. The computer-implemented method according to any one of the foregoing examples, wherein determining the truncated average comprises: determining a total truncated average by aligning the aggregated item configuration counts with the application programming interface data using the display dates; and determining a sliced truncated average by aligning the total aggregated item configuration counts with the application programming interface data using the display dates and the latency window levels.

[0218] Example 25. The computer-implemented method according to any one of the foregoing examples, wherein the total truncated averages comprise total configuration counts truncated at the configured quantity.

[0219] Example 26. The computer-implemented method according to any one of the foregoing examples, wherein the total truncated averages comprise edge cases that do not exist in the event simulation but occur in the total aggregated item configuration counts.

[0220] Example 27. The computer-implemented method according to any one of the foregoing examples, wherein the sliced truncated averages comprise per-data-slice truncation ratios.

[0221] Example 28. The computer-implemented method according to any one of the foregoing examples, wherein estimating the probability that each of these branches is true or noisy comprises generating a ratio of the total interaction count over the display dates obtained from the interaction logs to the total number of conditional interactions.

[0222] Example 29. A computer-implemented system, comprising: a memory that stores application programming interface (API) information; and a server that performs operations comprising: receiving, from an application programming interface, an event-level report comprising branches that define records of events that include interactions with the application programming interfaces, the event-level reports being paired with metadata corresponding to the configurations of the application programming interfaces; performing an identification of portions of the metadata that comprise invalid metadata; estimating the probability that each of these branches is true or noisy by using the identification of these portions of the metadata that comprise the invalid metadata, classifying portions of the events as being on true branches; determining configuration parameters for the true branches, the configuration parameters comprising the average number of configurations for each event; generating raw application programming interface data by applying a debiasing model that uses the configuration parameters across multiple reporting windows and configuration types; determining operations using the raw application programming interface data; and providing instructions to an asset provider system to activate at least one of the operations using the raw application programming interface data.

[0223] Example 30. A computer-implemented system as described in the foregoing examples, wherein the configuration parameters include a configuration window and a configuration count for each configuration type.

[0224] Example 31. A computer-implemented system as described in any of the foregoing examples, the operations further comprising: when truncated for each configuration type, determining the truncated average as the expected total configuration count for each configuration type and window.

[0225] Example 32. A computer-implemented system as described in any of the foregoing examples, determining the truncated average includes: determining the total truncated average by aligning the aggregated item configuration count with the application programming interface data using the display date; and determining the sliced truncated average by aligning the total aggregated item configuration count with the application programming interface data using the display dates and the latency window levels.

[0226] Example 33. A computer-implemented system as described in any of the foregoing examples, wherein the total truncated averages include the total configuration count truncated at the configured number.

[0227] Example 34. A computer-implemented system as described in any of the foregoing examples, wherein the total truncated averages include edge cases that do not exist in the event simulation but occur in the total aggregated item configuration count.

[0228] Example 35. A computer-implemented system as described in any of the foregoing examples, wherein the sliced truncated averages include a truncation ratio per data slice.

[0229] Example 36. A computer-implemented system as described in any of the foregoing examples, wherein estimating the probability that each of these branches is true or noisy includes generating a ratio of the total interaction count on the display date obtained from the interaction log to the total number of conditional interactions.

[0230] Example 37. A non - transitory computer - readable medium encoded with a computer program, the computer program including instructions that, when executed by one or more computers, cause the one or more computers to perform operations including: receiving, from an application programming interface, an event - level report including branches that define records of events including interactions with the application programming interface, the event - level reports being paired with metadata corresponding to the configuration of the application programming interface; performing an identification of portions of the metadata that include invalid metadata; estimating, for each of the branches, the probability of being true or noisy by using the identification of the portions of the metadata that include the invalid metadata, classifying portions of the events as being on true branches; determining configuration parameters for the true branches, the configuration parameters including the number of average configurations for each event; generating raw application programming interface data by applying a de - biasing model that uses the configuration parameters across multiple reporting windows and configuration types; using the raw application programming interface data to determine operations; and providing instructions to an asset provider system to activate at least one of the operations using the raw application programming interface data.

[0231] Example 38. The non - transitory computer - readable medium as described in the preceding example, wherein the configuration parameters include a configuration window and a configuration count for each configuration type, and the operations further include: when truncated for each configuration type, determining a truncated average as the expected total configuration count for each configuration type and window.

[0232] Example 39. The non - transitory computer - readable medium as described in the preceding example, determining the truncated average includes: determining a total truncated average by aligning aggregate item configuration counts with the application programming interface data using a display date; and determining a sliced truncated average by aligning the total aggregate item configuration counts with the application programming interface data using the display dates and a latency window level, wherein the total truncated average includes a total configuration count truncated at a number of configurations, and wherein the total truncated average includes edge cases that do not exist in an event simulation but occur in the total aggregate item configuration count.

[0233] Example 40. The non - transitory computer - readable medium as described in any of the preceding examples, wherein the sliced truncated average includes a truncation ratio per data slice, and wherein estimating the probability of each of the branches being true or noisy includes generating a ratio of the total interaction count on a display date obtained from an interaction log to the total number of conditional interactions.

[0234] Example 41. A computer-implemented method includes: receiving, from an application programming interface, an event-level report including a biased record of events, the events including interactions with the application programming interfaces, the event-level reports being paired with metadata corresponding to a configuration applied by the application programming interfaces; generating an original event-level report from the event-level reports by applying a debiasing model using configuration parameters of the application programming interfaces to remove spurious events from the event-level reports; receiving, from the application programming interfaces, an aggregated summary report, the aggregated summary reports including an aggregated record of the events included in the event-level reports corresponding to the interactions with the application programming interfaces; generating an original aggregated summary report from the aggregated summary reports by using a metadata mapping to remove false positive events from the event-level reports; generating statistical data by matching the original event-level reports with the original aggregated summary reports according to event scenarios; determining an operation using the statistical data; and providing instructions to an asset provider system to activate at least one of the operations using the statistical data. Compared with an overall single type of interaction data report generated by third-party cookies, the techniques described in this example include an integration protocol for two types of interaction data reports: event-level reports and aggregated summary reports. The integration protocol for the two types of interaction data reports provides the advantage of facilitating interaction measurement with high utility without violating privacy and security measures imposed by system privacy settings. Specifically, the described example utilizes information from event-level reports and aggregated summary reports. When aggregated, the integrated event-level logs can be consistent with the post-processed aggregated items from the aggregated summary reports, which does not guarantee the consistency of the original event-level reports and the aggregated summary reports processed separately as described in Reference Examples 1 and 21. By leveraging the combined use of API data (which provides better measurement fidelity compared to using either report type alone), generating statistical data by merging event-level data and aggregated summary reports improves the accuracy of interaction measurement.

[0235] Example 42. The computer-implemented method as described in the foregoing example, wherein the biased record of the events includes anonymizing, aggregating, truncating information, and injecting noise into the interaction data to protect the privacy of users performing the interactions with the application programming interfaces.

[0236] Example 43. The computer-implemented method as described in any of the foregoing examples, wherein the aggregated summary reports include hierarchically structured event attribution configuration data as nodes distributed at multiple levels, and the event attribution configuration data includes truncated configurations aggregated as data slices at one or more levels.

[0237] Example 44. The computer-implemented method according to any one of the foregoing examples, wherein the noise includes Laplace noise added to each of these data slices.

[0238] Example 45. The computer-implemented method according to any one of the foregoing examples, comprising: generating a weighted average of a parent and a child to minimize the variance of an estimate of the parent node.

[0239] Example 46. The computer-implemented method according to any one of the foregoing examples, wherein processing these raw aggregated summary reports to reduce the noise includes: applying a denoising transform to correct both these truncated configurations and the randomized response noise implemented by the Laplace noise.

[0240] Example 47. The computer-implemented method according to any one of the foregoing examples, wherein the denoising transform uses a set of summary statistics indexing aspects of the configuration type.

[0241] Example 48. The computer-implemented method according to any one of the foregoing examples, wherein these aggregated summary reports include metadata corresponding to the configuration types applied to the structured event attribution configuration data.

[0242] Example 49. The computer-implemented method according to any one of the foregoing examples, wherein these configuration parameters include a configuration window and a configuration count for each configuration type.

[0243] Example 50. The computer-implemented method according to any one of the foregoing examples, further comprising: when truncated for a configuration, determining the truncated average to be the expected total configuration count for each configuration type and window.

[0244] Example 51. The computer-implemented method according to any one of the foregoing examples, determining the truncated average includes: determining a total truncated average by aligning the aggregated item configuration count with the application programming interface data; and determining a sliced truncated average by aligning the total aggregated item configuration count with the application programming interface data.

[0245] Example 52. The computer-implemented method according to any one of the foregoing examples, wherein these total truncated averages include the total configuration count truncated at the number of configurations.

[0246] Example 53. The computer-implemented method according to any one of the foregoing examples, wherein these total truncated averages include edge cases.

[0247] Example 54. The computer-implemented method according to any one of the foregoing examples, wherein these sliced truncated averages include a per-data-slice truncation ratio.

[0248] Example 55. The computer-implemented method according to any one of the foregoing examples, wherein estimating the probability that each of these branches is true or noisy includes generating a ratio of the total interaction count on the display date obtained from the interaction log to the total number of conditional interactions.

[0249] Example 56. A computer-implemented system, comprising: a memory that stores application programming interface (API) information; and a server that performs operations including: receiving, from the application programming interface, an event-level report including a biased record of events, the events including interactions with the application programming interfaces, the event-level reports being paired with metadata corresponding to the configurations applied by the application programming interfaces; generating, from the event-level reports, raw event-level reports by applying a debiasing model using the configuration parameters of the application programming interfaces to remove false events from the event-level reports; receiving, from the application programming interfaces, aggregated summary reports, the aggregated summary reports including aggregated records of the events corresponding to the interactions with the application programming interfaces included in the event-level reports; generating, from the aggregated summary reports, raw aggregated summary reports by using metadata mapping to remove false positive events from the event-level reports; generating statistical data by matching the raw event-level reports with the raw aggregated summary reports according to event scenarios; determining operations using the statistical data; and providing instructions to an asset provider system to activate at least one of the operations using the statistical data.

[0250] Example 57. The computer-implemented system according to the foregoing example, wherein the biased record of the events includes anonymizing, aggregating, truncating information, and injecting noise into the interaction data to protect the privacy of users performing the interactions with the application programming interfaces, wherein the aggregated summary reports include hierarchically structured event attribution configuration data as nodes distributed in multiple levels, and the event attribution configuration data includes truncated configurations aggregated into data slices at one or more levels, and wherein the noise includes Laplace noise added to each of the data slices.

[0251] Example 58. The computer-implemented system according to any one of the foregoing examples, wherein the aggregated summary reports include metadata corresponding to the configuration types applied to the structured event attribution configuration data, and wherein the configuration parameters include a configuration window and a configuration count for each configuration type.

[0252] Example 59. A non-transitory computer-readable medium encoded with a computer program, the computer program comprising instructions that, when executed by one or more computers, cause the one or more computers to perform operations including: receiving, from an application programming interface, an event-level report including a biased record of events, the events including interactions with the application programming interface, the event-level reports being paired with metadata corresponding to a configuration applied by the application programming interface; generating an original event-level report from the event-level reports by applying a debiasing model using configuration parameters of the application programming interface to remove spurious events from the event-level reports; receiving, from the application programming interface, an aggregated summary report, the aggregated summary reports including an aggregated record of the events corresponding to the interactions with the application programming interface included in the event-level reports; generating an original aggregated summary report from the aggregated summary reports by using a metadata mapping to remove false positive events from the event-level reports; generating statistical data by matching the original event-level reports with the original aggregated summary reports according to event scenarios; using the statistical data to determine an operation; and providing instructions to an asset provider system to activate at least one of the operations using the statistical data.

[0253] Example 60. The non-transitory computer-readable medium as described in the foregoing example, wherein the biased record of events includes anonymizing, aggregating, information truncating, and noise injection of interaction data to protect the privacy of users performing the interactions with the application programming interface, wherein the aggregated summary reports include hierarchically structured event attribution configuration data as nodes distributed in multiple levels, and the event attribution configuration data includes truncated configurations aggregated as data slices at one or more levels, wherein the noise includes Laplace noise added to each of the data slices.

Claims

1. A computer-implemented method comprising: receiving, from an application programming interface, an aggregated summary report including aggregated event data collected by the application programming interface during events corresponding to interactions with a user interface, the aggregated event data being aggregated using a hierarchical structure corresponding to a data type; removing a first portion of the aggregated event data identified as false positives from the aggregated summary report to maintain true positive aggregated event data; At each level of the hierarchical structure, reducing noise from the true positive aggregated event data to generate a denoised aggregated summary report; determining an action using the denoised aggregated summary report; as well as Instructions are provided to an asset provider system to activate at least one of the operations using the de-noised aggregated summary report.

2. The computer-implemented method of claim 1, wherein: The aggregated summary report includes hierarchically structured event attribution configuration data as nodes distributed in a plurality of levels.

3. The computer-implemented method of claim 2, wherein: The event attribution configuration data includes truncated configurations aggregated into data slices at one or more levels.

4. The computer-implemented method of claim 3, wherein: The noise includes Laplace noise added to each of the data slices.

5. The computer-implemented method of claim 4, further comprising: Generate a weighted average of the parent and child nodes to minimize the variance of the parent's estimate.

6. The computer-implemented method of claim 5, wherein: Processing the aggregated summary report to reduce the noise includes: A denoising transform is applied to correct both the truncated data and the randomized response noise achieved by the Laplace noise.

7. The computer-implemented method of claim 6, wherein: Applying the denoising transform includes using a set of summary statistics for aspects of an index configuration type.

8. The computer-implemented method of claim 7, wherein: The aggregated summary report includes metadata corresponding to the configuration type applied to the structured event attribution configuration data.

9. A computer-implemented system comprising: a memory storing application programming interface (API) information; as well as A server performs the following operations: receiving, from an application programming interface, an aggregated summary report including aggregated event data collected by the application programming interface during events corresponding to interactions with a user interface, the aggregated event data being aggregated using a hierarchical structure corresponding to a data type; removing a first portion of the aggregated event data identified as false positives from the aggregated summary report to maintain true positive aggregated event data; At each level of the hierarchical structure, reducing noise from the true positive aggregated event data to generate a denoised aggregated summary report; determining an action using the denoised aggregated summary report; as well as Instructions are provided to an asset provider system to activate at least one of the operations using the de-noised aggregated summary report.

10. The computer-implemented system of claim 9, wherein: The aggregated summary report includes hierarchically structured event attribution configuration data as nodes distributed in a plurality of levels.

11. The computer-implemented system of claim 10, wherein: The event attribution configuration data includes truncated configurations aggregated into data slices at one or more levels.

12. The computer-implemented system of claim 11, wherein: The noise includes Laplace noise added to each of the data slices.

13. The computer-implemented system of claim 12, wherein: The operations further include: Generate a weighted average of the parent and child nodes to minimize the variance of the parent's estimate.

14. The computer-implemented system of claim 13, wherein: Processing the aggregated summary report to reduce the noise includes: A denoising transform is applied to correct both the truncated data and the randomized response noise achieved by the Laplace noise.

15. The computer-implemented system of claim 14, wherein: Applying the denoising transform includes using a set of summary statistics for aspects of an index configuration type.

16. The computer-implemented system of claim 15, wherein: The aggregated summary report includes metadata corresponding to the configuration type applied to the structured event attribution configuration data.

17. A non-transitory computer readable medium encoded with a computer program, the computer program comprising instructions that, when executed by one or more computers, cause the one or more computers to perform operations comprising: receiving, from an application programming interface, an aggregated summary report including aggregated event data collected by the application programming interface during events corresponding to interactions with a user interface, the aggregated event data being aggregated using a hierarchical structure corresponding to a data type; removing a first portion of the aggregated event data identified as false positives from the aggregated summary report to maintain true positive aggregated event data; At each level of the hierarchical structure, reducing noise from the true positive aggregated event data to generate a denoised aggregated summary report; determining an action using the denoised aggregated summary report; as well as Instructions are provided to an asset provider system to activate at least one of the operations using the de-noised aggregated summary report.

18. The non-transitory computer readable medium of claim 17, wherein: The aggregated summary report includes hierarchically structured event attribution configuration data as nodes distributed in multiple levels, wherein the event attribution configuration data includes truncated configurations aggregated into data slices at one or more levels, wherein the noise includes Laplace noise added to each of the data slices.

19. The non-transitory computer readable medium of claim 18, wherein: The operations further include: Generate a weighted average of the parent and child nodes to minimize the variance of the parent's estimate.

20. The non-transitory computer readable medium of claim 19, wherein: Processing the aggregated summary report to reduce the noise includes: A denoising transformation is applied to correct both the truncated data and the randomized response noise implemented by the Laplace noise, wherein applying the denoising transformation includes using a set of summary statistics of aspects of an indexed configuration type, and wherein the aggregated summary report includes metadata corresponding to the configuration type applied to the structured event attribution configuration data.

21. A computer-implemented method comprising: receiving, from an application programming interface, an event level report including branches, the branch definition including a record of events of interactions with the application programming interface, the event level report being paired with metadata corresponding to a configuration of the application programming interface; performing identification of portions of the metadata including invalid metadata; classifying a portion of the event as being on a true branch by estimating a probability of each of the branches being true or noisy using the identification of the portion of the metadata including the invalid metadata; determining a configuration parameter for the true branch, the configuration parameter comprising an average number of configurations for each event; generating raw application programming interface data by applying a debiasing model using the configuration parameters and by associating a plurality of report windows and configuration types; determining an operation using the raw application programming interface data; as well as Instructions are provided to an asset provider system to activate at least one of the operations using the raw application programming interface data.

22. The computer-implemented method of claim 21, wherein: The configuration parameters include a configuration window and a configuration count for each configuration type.

23. The computer-implemented method of claim 22, further comprising: When truncated for each configuration type, a truncated average is determined to be the expected total configuration count for that configuration type and window.

24. The computer-implemented method of claim 21, determining the truncated mean comprises: determining an overall truncated average by aligning the aggregated item configuration counts with the application programming interface data using presentation dates; as well as A sliced ​​truncated average is determined by aligning a total aggregate item configuration count with the application programming interface data using the presentation date and latency window level.

25. The computer-implemented method of claim 24, wherein: The total truncated average includes the total configuration count truncated at the number of configurations.

26. The computer-implemented method of claim 24, wherein: The overall truncated average includes edge cases that were not present in the event simulation but were present in the overall aggregate item configuration count.

27. The computer-implemented method of claim 21, wherein: The sliced ​​truncated means include a truncation ratio for each data slice.

28. The computer-implemented method of claim 21, wherein: Estimating the probability that each of the branches is true or noisy includes generating a ratio of a total interaction count on a presentation date obtained from an interaction log relative to a total number of conditional interactions.

29. A computer-implemented system comprising: a memory storing application programming interface (API) information; as well as A server performs the following operations: receiving, from an application programming interface, an event level report including a branch, the branch definition including a record of events of an interaction with the application programming interface, the event level report being paired with metadata corresponding to a configuration of the application programming interface; performing identification of portions of the metadata including invalid metadata; classifying a portion of the event as being on a true branch by estimating a probability of each of the branches being true or noisy using the identification of the portion of the metadata including the invalid metadata; determining a configuration parameter for the true branch, the configuration parameter comprising an average number of configurations for each event; generating raw application programming interface data by applying a debiasing model using the configuration parameters and by associating a plurality of report windows and configuration types; determining an operation using the raw application programming interface data; as well as Instructions are provided to an asset provider system to activate at least one of the operations using the raw application programming interface data.

30. The computer implemented system of claim 29, wherein: The configuration parameters include a configuration window and a configuration count for each configuration type.

31. The computer implemented system of claim 30, the operations further comprising: When truncated for each configuration type, a truncated average is determined to be the expected total configuration count for that configuration type and window.

32. The computer implemented system of claim 29, determining the truncated mean comprises: determining an overall truncated average by aligning the aggregated item configuration counts with the application programming interface data using presentation dates; as well as A sliced ​​truncated average is determined by aligning a total aggregate item configuration count with the application programming interface data using the presentation date and latency window level.

33. The computer implemented system of claim 32, wherein: The total truncated average includes the total configuration count truncated at the number of configurations.

34. The computer implemented system of claim 32, wherein: The overall truncated average includes edge cases that were not present in the event simulation but were present in the overall aggregate item configuration count.

35. The computer implemented system of claim 29, wherein: The sliced ​​truncated means include a truncation ratio for each data slice.

36. The computer implemented system of claim 29, wherein: Estimating the probability that each of the branches is true or noisy includes generating a ratio of a total interaction count on a presentation date obtained from an interaction log relative to a total number of conditional interactions.

37. A non-transitory computer readable medium encoded with a computer program, the computer program comprising instructions that, when executed by one or more computers, cause the one or more computers to perform operations comprising: receiving, from an application programming interface, an event level report including a branch, the branch definition including a record of events of an interaction with the application programming interface, the event level report being paired with metadata corresponding to a configuration of the application programming interface; performing identification of portions of the metadata including invalid metadata; classifying a portion of the event as being on a true branch by estimating a probability of each of the branches being true or noisy using the identification of the portion of the metadata including the invalid metadata; determining a configuration parameter for the true branch, the configuration parameter comprising an average number of configurations for each event; generating raw application programming interface data by applying a debiasing model using the configuration parameters and by associating a plurality of report windows and configuration types; determining an operation using the raw application programming interface data; as well as Instructions are provided to an asset provider system to activate at least one of the operations using the raw application programming interface data.

38. The non-transitory computer readable medium of claim 37, wherein: The configuration parameters include a configuration window and a configuration count for each configuration type, and the operation further includes: When truncated for each configuration type, a truncated average is determined to be the expected total configuration count for that configuration type and window.

39. The non-transitory computer readable medium of claim 37, determining the truncated mean comprises: determining an overall truncated average by aligning the aggregated item configuration counts with the application programming interface data using presentation dates; as well as A sliced ​​truncated average is determined by aligning a total aggregate item configuration count with the application programming interface data using the presentation date and latency window level, wherein the total truncated average includes the total configuration count truncated at the number of configurations, wherein the total truncated average includes edge cases that are not present in an event simulation but appear in the total aggregate item configuration count.

40. The non-transitory computer readable medium of claim 37, wherein: The sliced ​​truncated means include a truncation ratio per data slice, and wherein estimating the probability that each of the branches is true or noisy includes generating a ratio of a total interaction count on a presentation date obtained from an interaction log relative to a total number of conditional interactions.

41. A computer-implemented method comprising: receiving, from an application programming interface, an event-level report comprising a biased record of events, the events comprising interactions with the application programming interface, the event-level report paired with metadata corresponding to a configuration applied by the application programming interface; generating a raw event level report from the event level report by applying a debiasing model using configuration parameters of the application programming interface to remove false events from the event level report; receiving an aggregated summary report from the application programming interface, the aggregated summary report comprising an aggregated record of events included in the event-level report corresponding to the interactions with the application programming interface; generating an original aggregated summary report from the aggregated summary report by using metadata mapping to remove false positive events from the event level report; generating statistical data by matching the original event-level reports with the original aggregated summary reports according to event scenarios; determining an action using said statistical data; as well as Instructions are provided to an asset provider system to activate at least one of the operations using the statistical data.

42. The computer-implemented method of claim 41, wherein: The biased recording of events includes anonymization, aggregation, information truncation, and noise injection of interaction data to protect the privacy of a user performing the interaction with the application programming interface.

43. The computer-implemented method of claim 42, wherein: The aggregated summary report includes hierarchically structured event attribution configuration data as nodes distributed in a plurality of levels, and the event attribution configuration data includes truncated configurations aggregated into data slices at one or more levels.

44. The computer-implemented method of claim 43, wherein: The noise includes Laplace noise added to each of the data slices.

45. The computer-implemented method of claim 44, further comprising: Generate a weighted average of the parent and child to minimize the variance of the parent's estimate.

46. ​​The computer-implemented method of claim 45, wherein: Processing the raw aggregated summary report to reduce the noise includes: A denoising transform is applied to correct both the truncated configuration and the randomized response noise achieved by the Laplace noise.

47. The computer-implemented method of claim 46, wherein: The denoising transform uses a set of summary statistics indexing aspects of the configuration type.

48. The computer-implemented method of claim 41, wherein: The aggregated summary report includes metadata corresponding to a configuration type applied to the structured event attribution configuration data.

49. The computer-implemented method of claim 41, wherein: The configuration parameters include a configuration window and a configuration count for each configuration type.

50. The computer-implemented method of claim 49, further comprising: When truncated for configurations, a truncated average is determined as the expected total configuration count for each configuration type and window.

51. The computer-implemented method of claim 41, determining the truncated mean comprises: determining an overall truncated average by aligning aggregate item configuration counts with the application programming interface data; as well as The sliced ​​truncated average is determined by aligning the total aggregate item configuration count with the application programming interface data.

52. The computer-implemented method of claim 51 , wherein: The total truncated average includes the total configuration count truncated at the number of configurations.

53. The computer-implemented method of claim 51 , wherein: The overall truncated mean includes edge cases.

54. The computer-implemented method of claim 41, wherein: The sliced ​​truncated means include a truncation ratio for each data slice.

55. The computer-implemented method of claim 41, wherein: Estimating the probability that each of the branches is true or noisy includes generating a ratio of a total interaction count on a presentation date obtained from an interaction log relative to a total number of conditional interactions.

56. A computer-implemented system comprising: a memory storing application programming interface (API) information; as well as A server performs the following operations: receiving, from an application programming interface, an event-level report comprising a biased record of events, the events comprising interactions with the application programming interface, the event-level report paired with metadata corresponding to a configuration applied by the application programming interface; generating a raw event level report from the event level report by applying a debiasing model using configuration parameters of the application programming interface to remove false events from the event level report; receiving an aggregated summary report from the application programming interface, the aggregated summary report comprising an aggregated record of events included in the event-level report corresponding to the interactions with the application programming interface; generating an original aggregated summary report from the aggregated summary report by using metadata mapping to remove false positive events from the event level report; generating statistical data by matching the original event-level reports with the original aggregated summary reports according to event scenarios; determining an action using said statistical data; as well as Instructions are provided to an asset provider system to activate at least one of the operations using the statistical data.

57. The computer implemented system of claim 56, wherein: The biased record of events includes anonymization, aggregation, information truncation and noise injection of interaction data to protect the privacy of users performing the interaction with the application programming interface, wherein the aggregated summary report includes hierarchically structured event attribution configuration data as nodes distributed in multiple levels, and the event attribution configuration data includes truncated configurations aggregated into data slices at one or more levels, wherein the noise includes Laplace noise added to each of the data slices.

58. The computer implemented system of claim 56, wherein: The aggregated summary report includes metadata corresponding to configuration types applied to the structured event attribution configuration data, and wherein the configuration parameters include a configuration window and a configuration count for each configuration type.

59. A non-transitory computer readable medium encoded with a computer program, the computer program comprising instructions that, when executed by one or more computers, cause the one or more computers to perform operations comprising: receiving, from an application programming interface, an event-level report comprising a biased record of events, the events comprising interactions with the application programming interface, the event-level report paired with metadata corresponding to a configuration applied by the application programming interface; generating a raw event level report from the event level report by applying a debiasing model using configuration parameters of the application programming interface to remove false events from the event level report; receiving an aggregated summary report from the application programming interface, the aggregated summary report comprising an aggregated record of events included in the event-level report corresponding to the interactions with the application programming interface; generating an original aggregated summary report from the aggregated summary report by using metadata mapping to remove false positive events from the event level report; generating statistical data by matching the original event-level reports with the original aggregated summary reports according to event scenarios; determining an action using said statistical data; as well as Instructions are provided to an asset provider system to activate at least one of the operations using the statistical data.

60. The non-transitory computer readable medium of claim 59, wherein: The biased record of events includes anonymization, aggregation, information truncation and noise injection of interaction data to protect the privacy of users performing the interaction with the application programming interface, wherein the aggregated summary report includes hierarchically structured event attribution configuration data as nodes distributed in multiple levels, and the event attribution configuration data includes truncated configurations aggregated into data slices at one or more levels, wherein the noise includes Laplace noise added to each of the data slices.

Citation Information

Cited By

  • Big data advertisement label classification system based on AI analysis

    CN120822133A

  • Big data advertisement tag classification system based on ai analysis

    CN120822133B