Advertisement data consanguinity analysis method based on interpretable AI and computer system

By constructing a heterogeneous graph for advertising data lineage analysis and combining it with an anomaly detection model with dynamic thresholds, the misjudgment problem caused by static thresholds in advertising data anomaly detection is solved, and accurate detection of abnormal data and rapid location of responsible nodes are achieved.

CN120671048AInactive Publication Date: 2025-09-19GUANGZHOU TAIDONG TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510777969.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-11
Publication Date
2025-09-19
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing technology for advertising data anomaly detection has the problem of misjudgment caused by unreasonable static threshold settings, is difficult to adapt to the dynamic changes of data, and lacks an effective root cause location mechanism.

Method used

By adopting an advertising data lineage analysis method based on explainable AI, constructing a heterogeneous graph and quantifying the node influence weights, combined with an anomaly detection model with dynamic thresholds, we can achieve anomaly detection of advertising billing data and locate the responsible nodes.

Benefits of technology

It improves the accuracy of anomaly detection, can quickly locate the root cause of the anomaly, reduce misjudgments, and improve the efficiency and accuracy of anomaly handling.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120671048A_ABST
    Figure CN120671048A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data processing, in particular to an advertisement data consanguinity analysis method based on interpretable AI and a computer system. The method comprises the following steps: acquiring advertisement data and an advertisement charging rule, wherein the advertisement data comprises advertisement consumption data; according to the advertisement data and the advertisement charging rule, calculating a cost index to obtain advertisement charging data; constructing a heterogeneous graph comprising data nodes, rule nodes and service nodes, and quantifying the influence weight of each node on downstream data to obtain a blood relationship graph; inputting the advertisement charging data into a pre-constructed anomaly detection model, and judging that the advertisement charging data is abnormal data in response to the situation that a loss value of the anomaly detection model is greater than a target threshold value; wherein the target threshold value is determined by the statistical magnitude of the historical advertisement charging data. According to the method, the interpretability of the advertisement data and the accuracy of anomaly detection can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing technology. More specifically, the present invention relates to an advertising data lineage analysis method and computer system based on explainable AI. Background Art

[0002] In the large-scale media advertising billing business, data lineage analysis is a key technology for ensuring billing accuracy and traceability. This analysis often relies on lineage maps. However, traditional data lineage maps only present the data flow path and fail to explain the causal relationship between data changes and business and billing rules. This makes it difficult for business personnel to understand whether fluctuations in advertising consumption data are caused by billing rule changes, data transmission anomalies, or other factors, leading to a lack of basis for decision-making.

[0003] Secondly, when detecting anomalies in advertising data, static thresholds are often relied upon. However, this approach cannot adapt to dynamic changes in data. For example, normal fluctuations such as a surge in ad consumption (e.g., clickthroughs) during a promotional event can easily be misidentified as anomalies. However, truly abnormal data may not be detected in time due to improper threshold settings, resulting in misjudgment. Furthermore, even if an anomaly is detected, manual investigation is still required, making it difficult to quickly locate the root cause of the anomaly.

[0004] Finally, advertising billing rules are typically described in natural language and include complex combinations of conditions (for example, "When the ad runs between 6 PM and 10 PM on weekdays, targets men aged 25-35, and receives more than 500 clicks, the CPC will increase from 2 yuan to 3 yuan"). Traditional methods often rely on simple regular expression matching or basic natural language processing techniques, which are unable to accurately parse semantics. Manually maintaining the mapping between billing rules and data fields is not only time-consuming and labor-intensive, but also prone to errors.

[0005] Therefore, how to avoid misjudgment of static thresholds during anomaly detection and improve the accuracy of anomaly detection is a technical problem that needs to be solved urgently. Summary of the Invention

[0006] In order to solve the technical problem that the static threshold leads to low accuracy in detecting anomalies in advertising data, the present invention provides solutions in the following aspects.

[0007] In a first aspect, the present invention provides an advertising data lineage analysis method based on explainable AI, comprising: obtaining advertising data and advertising billing rules, the advertising data including advertising consumption data; calculating cost indicators based on the advertising data and the advertising billing rules to obtain advertising billing data; utilizing the advertising data and the advertising billing data to form data nodes; constructing a heterogeneous graph including data nodes, rule nodes, and business nodes, and quantifying the influence weight of each node in the heterogeneous graph on downstream data to obtain a lineage map; inputting the advertising billing data into a pre-built anomaly detection model, and in response to the loss value of the anomaly detection model being greater than a target threshold, determining that the advertising billing data is anomaly data; wherein the target threshold is determined by the statistics of historical advertising billing data.

[0008] Furthermore, the method for determining the target threshold includes: determining the target threshold according to the mean and standard deviation of the historical advertising billing data, and the target threshold is positively correlated with the mean and the standard deviation.

[0009] Furthermore, the target threshold value is calculated as follows:

[0010] threshold=μ+k·σ(1+α·event_weight);

[0011] Wherein, threshold is the target threshold, μ is the mean, σ is the standard deviation, α is the preset business event impact coefficient, k is a constant, and event_weight is the impact weight of the business event on data fluctuation.

[0012] Furthermore, after detecting abnormal data, the method further includes: searching for a responsible node causing the abnormality in the bloodline map according to a back propagation algorithm based on the abnormal node corresponding to the abnormal data.

[0013] Furthermore, the responsible node causing the anomaly is searched in the bloodline map according to the back propagation algorithm, including: determining the upstream node corresponding to the abnormal node, the input edge of the upstream node and the parent node of the upstream node; performing weighted summation on the influence weight of the input edge and the responsibility score of the parent node to obtain the responsibility score of the upstream node; and locating the responsible node based on the responsibility score of the upstream node.

[0014] Furthermore, quantifying the influence weight of each node in the heterogeneous graph on downstream data includes: quantifying based on SHAP value.

[0015] Furthermore, it also includes: inputting the advertising billing rule text into a pre-built mapping model for parsing to obtain a structured logical expression; the mapping model is obtained by fusing the BERT model, the BiLSTM model, and the CRF model, wherein the BERT model is used to extract semantic features, the BiLSTM model is used to further learn the semantics of the semantic features, and the CRF model is used to output the probability of the label sequence.

[0016] Furthermore, it also includes: monitoring the rule knowledge base, which includes multiple advertising billing rules; in response to the update of the advertising billing rules in the rule knowledge base, judging through a logical reasoning algorithm whether the updated advertising billing rules conflict with the original advertising billing rules, and if so, issuing a warning and feeding back the conflict path; if not, generating a rule difference report, and updating the objects associated with the advertising billing rules according to the updated advertising billing rules, the objects including the bloodline map.

[0017] Furthermore, it also includes: visualizing the bloodline map, abnormal analysis results, and output results of the mapping model, so that relevant personnel can make decisions based on the visualized content, and the decision content includes adjusting advertising billing rules and advertising delivery strategies.

[0018] In a second aspect, the present invention provides a computer system for advertising data lineage analysis based on explainable AI, comprising a processor and a memory, wherein the memory stores computer program instructions, and when the computer program instructions are executed by the processor, an advertising data lineage analysis method based on explainable AI according to any one of the first aspects is implemented.

[0019] The beneficial effects of the present invention are as follows: the present invention quantifies the influence weight of nodes on downstream data through SHAP values, and combines with the visualization of heterogeneous graphs to achieve deep semantic interpretation of data lineage relationships, providing intuitive and accurate basis for subsequent decision-making; by performing anomaly detection based on dynamic thresholds, it avoids misjudgments caused by unreasonable static threshold settings, and improves the accuracy of anomaly detection; based on abnormal nodes, through back-propagation attribution reasoning of heterogeneous graphs, it achieves rapid positioning of the root causes of data anomalies, thereby improving the efficiency and accuracy of anomaly handling; through mapping models, it achieves high-precision parsing of advertising billing rule texts described in complex natural languages, combines lineage graphs to achieve dynamic mapping of rules and data, and introduces rule version management and rule conflict detection functions, reducing manual maintenance costs and error rates, thereby improving the efficiency and accuracy of rule management and maintenance. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] The above and other objects, features and advantages of the exemplary embodiments of the present invention will become readily understood by reading the following detailed description with reference to the accompanying drawings. In the accompanying drawings, several embodiments of the present invention are shown in an illustrative and non-limiting manner, and the same or corresponding reference numerals represent the same or corresponding parts, wherein:

[0021] Figure 1 is a flowchart schematically illustrating a method for analyzing advertising data lineage based on explainable AI according to a first embodiment of the present invention;

[0022] Figure 2 is a flowchart schematically illustrating a method for analyzing advertising data lineage based on explainable AI according to a second embodiment of the present invention;

[0023] Figure 3 Schematically illustrates a structural block diagram of an advertising data lineage analysis system based on explainable AI according to a third embodiment of the present invention;

[0024] Figure 4 It is a block diagram schematically showing the structure of a computer system for analyzing advertising data lineage based on explainable AI according to the fourth embodiment of the present invention. DETAILED DESCRIPTION

[0025] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work shall fall within the scope of protection of the present invention.

[0026] The specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0027] Example 1

[0028] Figure 1 It is a flowchart schematically illustrating an advertising data lineage analysis method based on explainable AI according to an embodiment of the present invention.

[0029] In this embodiment, the present invention proposes an advertising data lineage analysis method based on explainable AI, such as Figure 1 As shown, the method of the present invention includes:

[0030] S101. Acquire advertisement data and advertisement billing rules.

[0031] Specifically, advertising data of multiple dimensions can be collected from multiple advertising data sources (e.g., advertising delivery platforms). In this embodiment, the advertising data includes at least three dimensions, specifically, advertising consumption data (e.g., exposure rate, click-through rate, etc.), advertising value estimation data, advertising tag data, etc.

[0032] Furthermore, advertising billing rules can be obtained from a preset rule knowledge base, wherein advertising billing rules include advertiser billing rules and large media billing rules, such as CPC (Cost Per Click) rules and CPA (Cost Per Action) rules.

[0033] S102: Calculate a cost indicator according to the advertisement data and the advertisement billing rule to obtain advertisement billing data.

[0034] Specifically, based on the advertisement consumption data and the corresponding advertisement billing rules, the cost indicators (such as cost, profit, etc.) of each advertisement plan are calculated, thereby obtaining advertisement billing data with business significance.

[0035] By calculating the advertising consumption data, the original data is converted into result data with business significance, providing an important basis for the subsequent construction and analysis of the bloodline map.

[0036] S103. Utilize the advertising data and the advertising billing data to form data nodes; construct a heterogeneous graph including data nodes, rule nodes, and business nodes, and quantify the influence weight of each node in the heterogeneous graph on downstream data to obtain a lineage graph.

[0037] Specifically, based on advertising data, advertising billing rules, and advertising billing data, a heterogeneous graph is constructed, including data nodes (e.g., advertising consumption data, advertising billing data), rule nodes (e.g., CPA rules, CPC rules), and business nodes (e.g., advertiser ID, advertising campaign ID). The nodes in the heterogeneous graph are connected through edges. For example, "data → rule" indicates that a data field is referenced by a rule, "rule → business" indicates that a rule is executed in a specific business, and "business → data" indicates that data is generated or modified.

[0038] Furthermore, the influence weight of each node in the heterogeneous graph on the downstream data is quantified to obtain a lineage map. In this embodiment, the influence weight of each node on the downstream data is calculated by the SHAP value. The SHAP (Shapley Additive Explanation) value is a method for interpreting the output results of machine learning models. It can be used to measure the contribution of each feature to the model's prediction results, thereby making complex machine learning models more interpretable.

[0039] Specifically, the calculation expression of the influence weight is:

[0040]

[0041] Where, SHAP i is the influence weight of feature i (corresponding to the node in the heterogeneous graph), F is the set of all features (i.e., the set of nodes in the heterogeneous graph), f(S) represents the predicted value of the model when the feature subset is S, f(S∪{i}) is the predicted value of the model after adding feature i to the feature subset S, S is the feature subset that does not include feature i, |S| is the size of the feature subset S, and |F| is the total number of features.

[0042] The SHAP value can be used to quantify the contribution of each node to data changes and highlight key lineage links. At the same time, semantic associations between data nodes, rule nodes, and business nodes are established. That is, the influence weight calculated by the SHAP value can clarify the causal relationship between data changes and business and rules, thereby improving the interpretability of the lineage map and providing a scientific basis for subsequent business personnel in decision-making.

[0043] In one embodiment, the method of the present invention also includes visualizing the bloodline map. Specifically, dynamic bloodline map visualization can be achieved through D3.js (a JavaScript library). More specifically, D3.js can be used to display detailed information (such as the original text of the rule, data statistical characteristics), highlight abnormal nodes and track paths, and the scope of influence of rule changes when a node is hovered, thereby facilitating business personnel to intuitively understand the causal relationship between data changes and business and rules, and provide an intuitive and reliable basis for decision-making.

[0044] To sum up, the bloodline map can intuitively display the flow path of data in the entire business process and the causal relationship between each node, providing a basic framework for data interpretation and analysis, and providing a scientific basis for business personnel's decision-making.

[0045] S104: Input the advertising billing data into a pre-built anomaly detection model to perform anomaly detection and obtain an anomaly detection result.

[0046] Specifically, the time series advertising billing data is input into a trained anomaly detection model, and in response to the loss value of the anomaly detection model being greater than a target threshold, the advertising billing data is determined to be anomaly data; wherein the target threshold is determined by the statistics of historical advertising billing data.

[0047] In this embodiment, the anomaly detection model used is trained using an LSTM-Autoencoder model (a deep learning model). When training the anomaly detection model, the model is optimized by minimizing the reconstruction error. In this embodiment, the reconstruction error uses the mean squared error (MSE), which is used as the loss value of the anomaly detection model. When the loss value of the anomaly detection model exceeds a target threshold, the input data is determined to be anomaly data.

[0048] Specifically, the calculation expression of the target threshold is:

[0049] threshold=μ+k·σ(1+α·event_weight);

[0050] Where threshold is the target threshold, μ is the mean, σ is the standard deviation, α is the preset business event impact coefficient, k is a constant, and event_weight is the weight of the business event's impact on data fluctuations. The business event impact coefficient is manually set based on the importance of the business. In this embodiment, it can be divided into 6 levels, ranging from 0 to 5, with a larger number indicating a greater weight. Furthermore, the weight of the business event's impact on data fluctuations is the same as the weight calculated in S103 based on the SHAP value.

[0051] By dynamically setting target thresholds based on historical data, it can adapt to dynamic changes in data and accurately distinguish normal fluctuations from abnormal fluctuations. This solves the problem that the static threshold anomaly detection method cannot adapt to dynamic changes in data, reduces misjudgments and missed judgments, and thus improves the accuracy of anomaly detection.

[0052] Example 2

[0053] In this embodiment, the present invention provides an advertising data lineage analysis method based on explainable AI, such as Figure 2 As shown, the method of the present invention includes:

[0054] S201: Acquire advertisement data and advertisement billing rules.

[0055] S202: Calculate a cost indicator according to the advertisement data and the advertisement billing rule to obtain advertisement billing data.

[0056] S203. Utilize the advertising data and the advertising billing data to form a data node; construct a heterogeneous graph including data nodes, rule nodes, and business nodes, and quantify the influence weight of each node on downstream data to obtain a lineage map.

[0057] S204: Input the advertisement consumption data into a pre-built anomaly detection model to perform anomaly detection and obtain an anomaly detection result.

[0058] It should be noted that the difference between Example 1 and Example 2 is that in addition to S101 (i.e., S201), S102 (i.e., S202), S103 (i.e., S203), and S104 (i.e., S204) in Example 1, Example 2 also includes S205, S206, S207, and S208.

[0059] S205 , based on the abnormal nodes corresponding to the abnormal data obtained by detection, searching for the responsible nodes causing the abnormality in the bloodline graph according to a back propagation algorithm.

[0060] Specifically, the upstream node corresponding to the abnormal node, the input edge of the upstream node, and the parent node of the upstream node are determined; the influence weight of the input edge and the responsibility score of the parent node are weighted and summed to obtain the responsibility score of the upstream node; based on the responsibility score, the responsible node that caused the abnormality is located, specifically, the upstream node with the largest responsibility score is regarded as the responsible node that caused the data abnormality.

[0061] In this embodiment, the calculation expression of the responsibility score is:

[0062] Resp(node)=∑ edge∈incoming_deges w(edge)·Resp(parent);

[0063] Where Resp(node) is the responsibility score of the upstream node of the abnormal node, node is the upstream node, incoming_deges is the set of input edges of the upstream node, w(edge) is the influence weight of the current input edge (i.e., the influence weight calculated according to the SHAP value in S203), edge is the current input edge, Resp(parent) is the responsibility score of the parent node corresponding to the current input edge, and parent is the parent node corresponding to the current input edge of the upstream node.

[0064] The responsibility score of the parent node can be a preset initial value or obtained through model training.

[0065] By quantifying the responsibility scores of upstream nodes associated with an anomaly, we can quickly and accurately locate the node responsible for the anomaly, improving the efficiency and accuracy of exception handling, thereby ensuring data accuracy and the normal operation of the business. It should be understood that a responsible node is the root cause of the data anomaly, such as a node with misconfigured rules or a node with incorrect data collection.

[0066] The present invention performs anomaly detection through an anomaly detection model combined with a dynamic threshold, and traces back to locate the root cause through the bloodline map. Compared with traditional static threshold detection, it is more accurate and efficient.

[0067] S206: Input the advertisement billing rule text into a pre-built mapping model for parsing to obtain a structured logical expression.

[0068] Since advertising billing rules are usually described in natural language and contain complex combinations of conditions, it is necessary to parse the advertising billing rule text described in natural language and convert it into structured logical expressions that computers can understand and use, so as to improve the efficiency and accuracy of maintaining advertising billing rules.

[0069] Specifically, a mapping model is constructed. In this embodiment, the mapping model is obtained by integrating the BERT (Bidirectional Encoder Representations from Transformers) model, the BiLSTM (Bidirectional Long Short-Term Memory Network) model, and the CRF (Conditional Random Field) model. Among them, the BERT model extracts the semantic features of the text, the BiLSTM model further learns the contextual information of the text, and the CRF model calculates the probability of the label sequence based on the features output by the BiLSTM model, thereby realizing the conversion of the advertising billing rules described in natural language into structured logical expressions. Among them, the probability formula P(y|x) for calculating the label sequence by the CRF model is:

[0070]

[0071] Where Z(x) is the normalization factor, x is the input advertising billing rule text, y is the predicted label sequence, ω yt-1,yt is the label transfer weight, v t,yt is the emission weight, and T is the length of the advertising billing rule text sequence.

[0072] Furthermore, the mapping model is fine-tuned using the labeled advertising billing rule samples to obtain the final mapping model. The loss function used in fine-tuning is:

[0073]

[0074] Where n is the number of advertising billing rule samples, y i is the labeled i-th logical expression, x i The input advertising billing rule text for the i-th advertisement.

[0075] The mapping model of the present invention can perform high-precision parsing of advertisement billing rule texts described in natural language, ensuring that computers accurately understand and execute complex advertisement billing rules, thereby improving the accuracy of executing advertisement billing rules.

[0076] S207: Monitor the rule knowledge base, and in response to receiving a rule knowledge base update, perform conflict detection and update records.

[0077] It is understood that the rule knowledge base stores multiple advertising billing rules. In one embodiment, rule document version control can be implemented based on Git-LFS (Git Large File Storage, a Git extension), and information such as the content, time, and personnel of each rule change can be recorded, and a rule difference report can be generated.

[0078] Specifically, when it is detected that a new advertising billing rule has been added to the rule knowledge base (or the original advertising billing rule has been modified), a logical reasoning algorithm is used to detect whether the new advertising billing rule (or the modified advertising billing rule) conflicts with the original advertising billing rule. If there is a conflict (for example, the original rule is that CPA is greater than 10, and the new rule is that CPA is less than 8), a warning is issued and the conflict path is fed back; if there is no conflict, the content, time, personnel and other information of the rule change are recorded, and a rule difference report is generated.

[0079] By performing conflict detection and version management on the updates of knowledge base rules, the accuracy and consistency of the rules can be ensured, which effectively reduces the manual maintenance costs and the risk of billing rule execution errors, and improves the efficiency and reliability of rule maintenance.

[0080] S208: Feedback on rule updates and / or anomaly detection results.

[0081] In one embodiment, it also includes: visual bloodline maps, anomaly detection results and rule analysis results, and feedback to business personnel so that business personnel can make decisions based on the visual data / content (such as adjusting advertising billing rules, adjusting advertising delivery strategies, etc.). The decision results will be fed back to the rule knowledge base to trigger the rule update process.

[0082] By visualizing bloodline maps, anomaly detection results, and rule analysis results, it can provide business personnel with intuitive and accurate scientific basis for decision-making, thereby improving the accuracy and efficiency of decision-making.

[0083] Specifically, if there is a rule update and no conflict exists, the updated billing rules are fed back to the steps (objects) related to advertising billing rules, including steps S202, S203, S206, etc. Similarly, if abnormal data is detected, the abnormality detection results are fed back to the steps related to advertising data usage, including steps S202, S203, S204, etc.

[0084] Specifically, for rule updates, it includes: recalculating advertising data according to the updated advertising billing rules, adjusting the node relationships and influence weight calculations of the lineage graph according to the updated advertising billing rules, and updating the rule mapping and conflict logic according to the updated advertising billing rules to achieve continuous optimization. Similarly, for abnormal results, it includes: checking whether there are problems in the process of calculating cost indicators based on abnormal results; the lineage graph optimizes the calculation of node relationships and influence weights based on abnormal results, and the mapping model analyzes whether the advertising billing rules are reasonable based on abnormal results, and then adjusts the parsing algorithm (mapping model) and rule management strategy. By feeding back rule updates and anomaly detection results to related processes, closed-loop optimization is achieved, reducing the cost of manual rule parsing and anomaly troubleshooting, and improving efficiency.

[0085] Example 3

[0086] In this embodiment, the present invention also provides an advertising data lineage analysis system based on explainable AI, which is used to implement the advertising data lineage analysis method based on explainable AI described in the first aspect.

[0087] Specifically, if Figure 3 As shown, the system includes: a data acquisition module, an advertising data calculation module, a bloodline map construction module, an anomaly detection and positioning module, a rule parsing and management module, a visualization and decision-making module, and a rule knowledge base module. The data acquisition module is connected to the advertising data calculation module; the advertising data calculation module is also connected to the bloodline map construction module, the anomaly detection and positioning module, and the rule knowledge base module; the bloodline map construction module is also connected to the anomaly detection and positioning module and the rule knowledge base module; the anomaly detection and positioning module is also connected to the rule parsing and management module, which is also connected to the visualization and decision-making module and the rule knowledge base module, and the visualization and decision-making module is also connected to the rule knowledge base module.

[0088] In one embodiment, the data collection module is used to collect advertising data and advertising billing rules; the advertising data calculation module is used to calculate cost indicators based on the advertising data and advertising billing rules to obtain advertising billing data; the lineage map construction module is used to construct a lineage map (including constructing a heterogeneous graph and calculating SHAP values); the anomaly detection and positioning module is used to perform anomaly detection and anomaly root location on the advertising data; the rule parsing and management module is used to parse and map the advertising billing rule text described in natural language, and perform rule conflict detection and rule version management; the visualization and decision module is used to visualize the lineage map, anomaly detection results and rule parsing results, and perform feedback display, as well as receive the decision results of business personnel and feed them back to the rule knowledge base module; the rule knowledge base module includes a preset rule knowledge base for receiving decision results.

[0089] In this system, a feedback mechanism is set up. Specifically, rule update feedback: the decision results of business personnel are fed back to the rule knowledge base module, triggering the rule update. The updated rule information is also fed back to the advertising data calculation module, the bloodline map construction module and the rule parsing and management module. The advertising data calculation module recalculates the advertising data according to the new rules (i.e., the new advertising billing rules); the bloodline map construction module adjusts the node relationship and SHAP value calculation according to the new rules; the rule parsing and management module updates the rule mapping and conflict detection logic to achieve continuous optimization of the system.

[0090] Anomaly Detection Feedback: Anomaly detection and location module anomaly results are fed back to the advertising data calculation module, the bloodline map construction module, and the rule analysis and management module. The advertising data calculation module checks the calculation process for problems based on anomalies. The bloodline map construction module uses anomaly information to optimize node relationships and influence weight calculations. The rule analysis and management module analyzes the rationality of the rules based on the anomaly cause and adjusts the analysis algorithm and rule management strategy, forming a closed-loop optimization.

[0091] Example 4

[0092] Figure 4 It is a block diagram schematically showing the structure of a computer system for analyzing advertising data lineage based on explainable AI according to this embodiment.

[0093] In a second aspect, the present invention also provides a computer system for analyzing advertising data lineage based on explainable AI. Figure 4 As shown, the computer system includes a processor and a memory, and the memory stores computer program instructions. When the computer program instructions are executed by the processor, an advertising data lineage analysis method based on explainable AI according to the first aspect of the present invention is implemented.

[0094] The computer system also includes other components well known to those skilled in the art, such as a communication interface. The configuration and functions of these components are known in the art and will not be described in detail here.

[0095] In the present invention, the aforementioned memory can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. For example, the computer-readable storage medium can be any suitable magnetic storage medium or magneto-optical storage medium, such as resistive random access memory (RRAM), dynamic random access memory (DRAM), static random access memory (SRAM), enhanced dynamic random access memory (EDRAM), high-bandwidth memory (HBM), hybrid memory cube (HMC), etc., or any other medium that can be used to store the required information and can be accessed by an application, module, or both. Any such computer storage medium can be part of a device or accessible or connectable to a device. Any application or module described in the present invention can be implemented using computer-readable / executable instructions that can be stored or otherwise retained by such a computer-readable medium.

[0096] In the description of this specification, “a plurality of” means at least two, for example, two, three or more, etc., unless otherwise clearly defined.

[0097] While several embodiments of the present invention have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Numerous modifications, variations, and alternatives will occur to those skilled in the art without departing from the concept and spirit of the present invention. It should be understood that various alternatives to the embodiments of the present invention described herein may be employed in practicing the present invention.

Claims

1. A method for analyzing advertising data lineage based on explainable AI, characterized by: include: Acquiring advertising data and advertising billing rules, wherein the advertising data includes advertising consumption data; Calculating a cost indicator according to the advertisement data and the advertisement billing rules to obtain advertisement billing data; Using the advertisement data and the advertisement billing data to form a data node; Construct a heterogeneous graph including data nodes, rule nodes, and business nodes, and quantify the influence weight of each node in the heterogeneous graph on downstream data to obtain a lineage map; The advertising billing data is input into a pre-built anomaly detection model, and in response to a loss value of the anomaly detection model being greater than a target threshold, the advertising billing data is determined to be anomaly data; wherein the target threshold is determined by statistics of historical advertising billing data.

2. The advertising data lineage analysis method based on explainable AI according to claim 1 is characterized in that: The method for determining the target threshold includes: determining the target threshold according to the mean value and the standard deviation of the historical advertising billing data, wherein the target threshold is positively correlated with the mean value and the standard deviation.

3. The advertising data lineage analysis method based on explainable AI according to claim 2 is characterized in that: The calculation expression of the target threshold is: threshold=μ+k·σ(1+α·event_weight); Wherein, threshold is the target threshold, μ is the mean, σ is the standard deviation, α is the business event impact coefficient, k is a constant, and event_weight is the impact weight of the business event on data fluctuation.

4. The advertising data lineage analysis method based on explainable AI according to claim 1 is characterized in that: After the abnormal data is detected, the method further includes: searching for a responsible node causing the abnormality in the bloodline map according to a back propagation algorithm based on the abnormal node corresponding to the abnormal data.

5. The advertising data lineage analysis method based on explainable AI according to claim 4 is characterized in that: Searching for the responsible node causing the abnormality in the bloodline graph according to a back-propagation algorithm, including: determining the upstream node corresponding to the abnormal node, the input edge of the upstream node, and the parent node of the upstream node; Performing a weighted summation on the influence weight of the input edge and the responsibility score of the parent node to obtain the responsibility score of the upstream node; The responsible node is located based on the responsibility score of the upstream node.

6. The advertising data lineage analysis method based on explainable AI according to claim 1 is characterized in that: Quantifying the influence weight of each node in the heterogeneous graph on downstream data includes: quantifying based on SHAP value.

7. The advertising data lineage analysis method based on explainable AI according to claim 1 is characterized in that: Also includes: The advertising billing rule text is input into a pre-built mapping model for parsing to obtain a structured logical expression; the mapping model is obtained by fusing the BERT model, the BiLSTM model, and the CRF model, wherein the BERT model is used to extract semantic features, the BiLSTM model is used to further learn the semantics of the semantic features, and the CRF model is used to output the probability of the label sequence.

8. The advertising data lineage analysis method based on explainable AI according to claim 7 is characterized in that: Also includes: A monitoring rule knowledge base, wherein the rule knowledge base includes a plurality of advertising billing rules; In response to the update of the advertising billing rules in the rule knowledge base, a logical reasoning algorithm is used to determine whether the updated advertising billing rules conflict with the original advertising billing rules. If so, a warning is issued and the conflict path is fed back; if not, a rule difference report is generated, and the objects associated with the advertising billing rules are updated according to the updated advertising billing rules, and the objects include the bloodline map.

9. The advertising data lineage analysis method based on explainable AI according to claim 1 is characterized in that: Also includes: The bloodline map, abnormality analysis results, and output results of the mapping model are visualized so that relevant personnel can make decisions based on the visualized content, including adjusting advertising billing rules and advertising delivery strategies.

10. A computer system for analyzing advertising data lineage based on explainable AI, characterized in that: It includes a processor and a memory, the memory stores computer program instructions, and when the computer program instructions are executed by the processor, an advertising data lineage analysis method based on explainable AI according to any one of claims 1 to 9 is implemented.

Citation Information

Cited By

  • Method for rapidly completing competitive product advertisement analysis based on AI model

    CN121707606A