A method and system for discovering causal relationships based on data mining
Patent Information
- Application Number
- CN202611079510.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-21
- Publication Date
- 2026-09-15
- Estimated Expiration
- 2046-07-21
AI Technical Summary
当该变量丧失自然波动后,因果发现算法中的倾向性得分模型无法有效估计不同用户被干预的真实概率,导致模型估计结果失效
[0019] Beneficial Effects: This application proposes a data mining-based causal relationship discovery method and system. By establishing a policy association mapping mechanism based on user identifiers and timestamps at the data access end, it generates intervention impact identifiers for each behavior log record and introduces multi-level intervention status labeling and differentiated sample hierarchical management. This effectively distinguishes between naturally observed samples and samples contaminated by policy feedback, eliminating the problem of spurious causal backflow caused by policy intervention at the source. Simultaneously, by calculating the difference in association metrics between natural samples and intervention samples, it achieves precise localization of the relationship between intervention-sensitive variables. Based on this, it applies controlled constraints to propensity score estimation and conditional independence testing, ensuring that the causal framework reconstruction is entirely based on pure and semantically consistent natural observation samples. This solution significantly improves the robustness of incremental causal discovery in complex internet environments, effectively avoiding misjudging spurious correlations generated by human intervention as genuine causality, and providing a reliable decision-making basis for iterative optimization of business strategies.
Smart Images

Figure CN122596217B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data mining technology, and in particular to a method and system for discovering causal relationships based on data mining. Background Technology
[0002] In internet platform operations, the platform periodically performs causal discovery tasks. This involves retrieving user behavior logs from the data lake within a preset time window and running causal discovery algorithms to continuously update the understanding of causal relationships in the user behavior conversion chain. The causal graph output by this task is used to guide subsequent product strategy adjustments. For example, based on the positive causal paths identified in the causal graph, strategic interventions can be implemented for specific user groups.
[0003] In scenarios where business is constantly iterating, policy interventions will generate new user behavior data and accumulate it in the platform's log stream. Then, during the data retrieval in the next cycle, this intervened data will be mixed with uninterrupted natural observation data and enter the processing scope of the causal discovery task.
[0004] When the input data of a causal discovery algorithm mixes data influenced by policy intervention with unaffected natural observation data, two types of problems arise: First, strategic intervention causes a certain behavioral variable (such as whether or not to claim a platform coupon) to have a highly consistent value among the intervened user group (all have claimed it), while this variable should have normal individual differences under natural conditions (i.e. some users claim it and some do not). When this variable loses its natural fluctuation, the propensity score model in the causal discovery algorithm cannot effectively estimate the true probability of different users being intervened, causing the model estimation results to fail.
[0005] Secondly, strategic intervention can significantly improve certain user behavior metrics in the short term. This short-term improvement may be mistaken by the causal discovery algorithm as a natural correlation between the metric and other variables, leading the algorithm to incorrectly determine that there is a causal relationship between these variables and include them in the causal graph.
[0006] Because the existing causal discovery process lacks traceable identification of whether each data record has been affected by policy intervention, the algorithm indiscriminately uses all data (including intervened and uninterrupted data) to reconstruct the causal graph during runtime. The direct consequence is that the data changes artificially created by the previous round of policy intervention are mixed into the current causal structure, causing the newly generated causal graph to deviate from the user's true behavioral mechanism under natural conditions, i.e., forming a pseudo-causal backflow.
[0007] To address the aforementioned issues, existing technologies urgently need improvement. Summary of the Invention
[0008] In view of the shortcomings of the prior art, this application provides a data mining-based causal relationship discovery method and system, which aims to solve the problem of how to identify and isolate data records affected by policy intervention in periodic causal discovery tasks, so as to suppress the interference of policy feedback pollution on the accuracy of causal graph reconstruction.
[0009] Firstly, a causal relationship discovery method based on data mining, the method comprising the following steps: S1: Obtain user behavior log data and historical policy intervention records; S2: Using the user identifier and event timestamp as matching keys, associate the historical policy intervention records with the user behavior log data to generate an intervention impact identifier for each piece of user behavior log data; S3: Based on the intervention impact identifier, identify the intervention status label of each user behavior log data, and divide the user behavior log data into different sample levels according to the intervention status label, and configure corresponding availability constraints for each sample level; S4: Calculate the correlation metrics between variables on the user behavior log data at different sample levels, and identify the relationships between sensitive variables affected by the intervention based on the differences in the correlation metrics between different sample levels. S5: Based on the availability constraints corresponding to each sample level and the relationship between the sensitive variables, the execution process of the causal inference logic is constrained and modified to reconstruct the causal structure diagram that excludes the interference of policy intervention feedback pollution.
[0010] Furthermore, the user behavior log data includes user identifiers, event timestamps, user profile features, contextual features, and business conversion tags; The historical policy intervention records include policy identifier, policy type, policy issuance timestamp, policy effective time window, set of user group identifiers covered by the policy, and intervention variables applied by the policy.
[0011] Furthermore, step S2 includes: S21: Build a distributed association mapping engine, using user identifier and event timestamp as the joint matching key; S22: For each user behavior log data, iterate through the historical policy intervention records that are effective within the current time window, verify whether the user identifier corresponding to the user behavior log data falls into the user identifier set of the historical policy intervention records, and determine whether the event timestamp falls within the policy effective time window or the delayed impact period. S23: When the user identifier corresponding to the user behavior log data falls into the user identifier set of the historical policy intervention record, and the event timestamp falls within the policy effective time window or the lag effect period, an intervention impact identifier of the user behavior log data is generated. The intervention impact identifier includes the matched strategy identifier and the intervention intensity attenuation coefficient calculated based on the time difference between the event timestamp and the strategy effective time window; S24: When a user behavior log data matches multiple historical policy intervention records within the current time window, the multiple policy identifiers are appended to the same intervention impact identifier.
[0012] Furthermore, step S22 includes: S221: Expand the set of user identifiers in the historical policy intervention records into a Bloom filter; S222: For each user behavior log data, the user identifier corresponding to the user behavior log data is checked by querying the Bloom filter to see if it falls into the user identifier set; S223: Obtain a preset lag impact threshold, wherein the preset lag impact threshold is the maximum duration for which the historical policy intervention records still have a residual impact on user behavior, calculated from the end time of the policy effective time window; S224: Calculate the time difference between the event timestamp and the end time of the policy effective time window. If the time difference is less than or equal to zero, it is determined that the event timestamp falls within the policy effective time window. If the time difference is greater than zero and less than or equal to the preset lag effect threshold, it is determined that the event timestamp falls within the lag effect period.
[0013] Furthermore, step S23 includes: S231: If the user identifier corresponding to the user behavior log data falls into the user identifier set of the historical policy intervention record, and the event timestamp falls within the policy effective time window or the lag effect period, obtain the policy identifier of the matched historical policy intervention record. S232: Obtain the time difference between the event timestamp and the end time of the policy effective time window. If the time difference is less than or equal to zero, set the intervention intensity attenuation coefficient to 1. S233: If the time difference is greater than zero and less than or equal to the preset lag effect threshold, the intervention intensity attenuation coefficient is calculated according to the following formula: Intervention intensity attenuation coefficient = 1 - time difference / preset lag effect threshold; S234: Associate the strategy identifier and the intervention intensity attenuation coefficient with the corresponding user behavior log data to form the intervention impact identifier corresponding to the user behavior log data.
[0014] Furthermore, step S3 includes: S31: Read the intervention intensity attenuation coefficient from the intervention impact identifier of each user behavior log data; S32: Based on the comparison results between the intervention intensity attenuation coefficient and the preset direct effect threshold and the preset influence dissipation threshold, each piece of user behavior log data is marked as a no-intervention state label, a direct-intervention state label, a delayed feedback state label, or an undetermined state label. S33: Divide the user behavior log data marked with the non-intervention state label into the non-intervention sample level, and configure the highest sample weight and full algorithm access permissions for the non-intervention sample level; The user behavior log data marked with the direct intervention status label is divided into the direct intervention sample level, and an algorithm access permission mask is configured for the direct intervention sample level to disable the permissions for propensity score estimation and partial correlation calculation. The user behavior log data marked with the lag feedback status label is divided into the lag feedback sample level, the lag feedback sample level is configured with a sample weight that is smaller than that of the uninterrupted sample level, and an algorithm access permission mask that restricts participation in the confounding factor search is configured. The user behavior log data marked with the "undeterminable status" label is divided into the "undeterminable sample level", and the lowest sample weight is configured for the "undeterminable sample level".
[0015] Furthermore, step S4 includes: S41: Obtain a set of variable pairs, which includes candidate variable pairs formed by combining variables from user profile features, context features, and business conversion tags in pairs; S42: Extract the variable values corresponding to each candidate variable from the uninterrupted sample level and calculate the first correlation measure; extract the variable values corresponding to each candidate variable from the directly intervened sample level and calculate the second correlation measure; S43: Calculate the difference between the first correlation measure and the second correlation measure; S44: If the difference value is greater than the preset sensitivity threshold, then the candidate variable pair is marked as an intervention-sensitive relationship.
[0016] Furthermore, step S42 includes: S421: For each candidate variable pair, take any one variable in the candidate variable pair as the first variable and the other variable in the candidate variable pair as the second variable; S422: From all sample records of the uninterrupted sample level, extract the values of the first variable and the second variable in each sample record, calculate the first partial correlation coefficient between the first variable and the second variable, and use it as the first correlation measure. S423: Extract the values of the first variable and the second variable in each sample record from all sample records at the direct intervention sample level, and calculate the second partial correlation coefficient between the first variable and the second variable as the second correlation measure.
[0017] Furthermore, step S5 includes: S51: Based on the sample weights and algorithm access permission masks of each sample level, select a subset of samples that are allowed to participate in the bias score estimation from each sample level, run the bias score estimation model on the sample subset, and obtain the bias score of each sample. S52: Based on the sample weights and algorithm access permission masks of each sample level, select a subset of samples from each sample level that are allowed to participate in the conditional independence test; S53: When performing the conditional independence test, for candidate variable pairs marked as the intervention-sensitive relationship, the regular statistical test process is intercepted, and the candidate edges of the candidate variable pairs in the causal graph skeleton are disconnected or their confidence weights are reduced. S54: Based on the propensity score, the result of the conditional independence test, and the result of disconnecting or reducing the weight of the candidate edges, reconstruct the causal structure graph that excludes the interference of policy intervention feedback pollution.
[0018] Secondly, a causal relationship discovery system based on data mining, used to implement the steps in any of the above methods, the system comprising: Acquisition module: Acquires user behavior log data and historical policy intervention records; Generation module: Using user identifier and event timestamp as matching keys, associate the historical policy intervention records with the user behavior log data to generate an intervention impact identifier for each piece of user behavior log data; Segmentation Module: Based on the intervention impact identifier, identify the intervention status label of each user behavior log data, and divide the user behavior log data into different sample levels according to the intervention status label, and configure corresponding availability constraints for each sample level; Calculation module: Calculates correlation metrics between variables on user behavior log data at different sample levels, and identifies the relationships of sensitive variables affected by intervention based on the differences in correlation metrics between different sample levels; Reconstruction Module: Based on the availability constraints corresponding to each sample level and the relationship between the sensitive variables, the execution process of the causal inference logic is constrained and corrected to reconstruct the causal structure diagram that excludes the interference of policy intervention feedback pollution.
[0019] Beneficial Effects: This application proposes a data mining-based causal relationship discovery method and system. By establishing a policy association mapping mechanism based on user identifiers and timestamps at the data access end, it generates intervention impact identifiers for each behavior log record and introduces multi-level intervention status labeling and differentiated sample hierarchical management. This effectively distinguishes between naturally observed samples and samples contaminated by policy feedback, eliminating the problem of spurious causal backflow caused by policy intervention at the source. Simultaneously, by calculating the difference in association metrics between natural samples and intervention samples, it achieves precise localization of the relationship between intervention-sensitive variables. Based on this, it applies controlled constraints to propensity score estimation and conditional independence testing, ensuring that the causal framework reconstruction is entirely based on pure and semantically consistent natural observation samples. This solution significantly improves the robustness of incremental causal discovery in complex internet environments, effectively avoiding misjudging spurious correlations generated by human intervention as genuine causality, and providing a reliable decision-making basis for iterative optimization of business strategies. Attached Figure Description
[0020] Figure 1 This is a flowchart of a data mining-based causal relationship discovery method proposed in this application.
[0021] Figure 2 This is a structural diagram of a causal relationship discovery system based on data mining proposed in this application.
[0022] Labeling Explanation: 201, Acquisition Module; 202, Generation Module; 203, Partition Module; 204, Calculation Module; 205, Reconstruction Module. Detailed Implementation
[0023] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of the embodiments. The components of the embodiments of this application described and marked in the accompanying drawings can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely represents selected embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.
[0024] It should be noted that similar reference numerals and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. Furthermore, in the description of this application, terms such as "first," "second," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0025] Because the existing causal discovery process lacks traceable identification of whether each data record has been affected by policy intervention, the algorithm indiscriminately uses all data (including intervened and uninterrupted data) to reconstruct the causal graph during runtime. The direct consequence is that the data changes artificially created by the previous round of policy intervention are mixed into the current causal structure, causing the newly generated causal graph to deviate from the user's true behavioral mechanism under natural conditions, i.e., forming a pseudo-causal backflow.
[0026] Therefore, there is an urgent need for a solution that can identify and isolate data records affected by policy intervention in periodic causal discovery tasks, so as to suppress the interference of policy feedback contamination on the accuracy of causal graph reconstruction.
[0027] To solve the above problems, please refer to Figure 1 This application provides a data mining-based causal relationship discovery method, which includes the following steps: S1: Obtain user behavior log data and historical policy intervention records; S2: Using user ID and event timestamp as matching keys, associate historical policy intervention records with user behavior log data to generate an intervention impact identifier for each user behavior log data; S3: Based on the intervention impact identifier, identify the intervention status label of each user behavior log data, and divide the user behavior log data into different sample levels according to the intervention status label, and configure corresponding availability constraints for each sample level; S4: Calculate the correlation metrics between variables on user behavior log data at different sample levels, and identify the relationships of sensitive variables affected by the intervention based on the differences in the correlation metrics between different sample levels. S5: Based on the availability constraints and sensitive variable relationships corresponding to each sample level, the execution process of causal inference logic is constrained and modified to reconstruct the causal structure diagram that excludes the interference of strategy feedback pollution.
[0028] It's important to note that causal relationship discovery refers to uncovering causal dependencies between variables through statistical analysis and algorithmic deduction of observed data, typically presented as a directed acyclic graph (DAG). This causal dependency doesn't simply mean that two variables change simultaneously; rather, it means that a change in one variable has a directional impact on the change in the other. For example, in user behavior analysis, receiving a coupon and completing an order may be simultaneously related, but only after excluding common influencing factors such as user spending power, access frequency, and activity time period can we further determine whether the former has a causal effect on the latter.
[0029] In observational studies, propensity score is used to measure the conditional probability of a sample receiving a specific treatment or intervention, often used to eliminate confounding bias. More specifically, its role is to place treated and untreated samples, which are otherwise difficult to compare directly due to uneven sample distribution, into a more comparable analytical framework. For example, in the analysis of activation strategies, platforms often tend to issue incentives to inactive users. In this case, if the probability of the sample receiving the intervention is not estimated, it is easy to misjudge the user's inactivity as the reason for the change in conversion.
[0030] Conditional independence testing, used to determine whether a statistical correlation still exists between two variables given certain control variables, is a core step in constructing the skeleton of a causal graph. For example, if there is a correlation between page dwell time and whether an order is placed, and this correlation disappears after controlling for the source of the visit and the price range of the product, it can be concluded that the direct connection between the two may not hold, but rather is formed indirectly by the control variables. Because propensity score and conditional independence testing are highly dependent on the naturalness of the sample distribution, once feedback behavior actively generated by platform strategies is mixed into the data, it will directly affect the authenticity of the causal structure identification.
[0031] In the technical solution of this application, one feasible approach is to manually write scripts to periodically export recent user behavior logs from the business database in full and manually download historical policy intervention records from the operations management backend. While this approach can acquire basic data, it often lacks a dynamic time window alignment mechanism when dealing with massive and frequently updated internet platform data, and the extracted field dimensions may not be standardized enough, easily overlooking key contextual features. To more clearly and completely describe the feasible implementation of this step, user behavior log data can at least originate from a data tracking system, server-side behavior logs, order system event tables, or behavior details accumulated in message queues.
[0032] Historical policy intervention records can at least be derived from policy execution logs on the A / B testing platform. When actually acquiring these records, a unified data extraction standard can be agreed upon first. For example, user identifiers can be uniformly mapped to the same internal primary key format, event timestamps can be uniformly converted to time fields with the same time zone and precision, and the creation time, approval time, issuance time, and effective time in the policy records can be saved separately to avoid mistaking the policy configuration time as the actual start time of the policy's impact.
[0033] If a single user on the platform may be affected by multiple policies simultaneously, when extracting historical policy intervention records, context fields such as policy priority, policy channel, policy target type, and policy revocation status can also be retained to determine which type of policy affected a particular log entry. Through this information acquisition method, the data entering subsequent processes is no longer simply behavior details and policy lists, but standardized input with unified primary keys, unified time semantics, and unified field definitions.
[0034] More specifically, after acquiring the data, two data streams can be merged using table join operations in a relational database, with user ID and event timestamp as the association conditions. For example, if the user ID of a log entry exists in the coverage list of a certain policy, and the log's timestamp is after the policy execution, then a field as an intervention impact identifier is appended to the log entry, recording the name of the policy.
[0035] In practice, this association process can first expand the historical policy intervention records, that is, convert the set of user group identifiers covered into a detailed structure that can be matched on a user-by-user basis. If the policy records store audience group identifiers rather than user-by-user lists, the audience members at the corresponding time point can be parsed from the audience group snapshot table and then associated with the log data.
[0036] Furthermore, the matching of event timestamps should not solely rely on the condition that the event is later than the policy issuance time. Instead, it should prioritize judgment logic that the event falls within the policy's effective time window or within the period of sustained influence after the policy takes effect. This is because although some policies are issued at a certain time, their actual impact on user behavior may occur later. For example, there may be a delivery delay in message delivery, a usage period for coupons after they are claimed, and a cache refresh period for recommendation policies.
[0037] Therefore, in a more complete implementation example, an impact range description field can be generated for each policy record, which includes at least the start and end points of the impact. Then, each user behavior log is judged one by one: if the user identifier matches, the log time falls within the impact range, and the behavior type corresponding to the log is related to the intervention variable applied by the policy, then the policy is written into the intervention impact identifier.
[0038] If the same log entry hits multiple policies simultaneously, the intervention impact identifier can be recorded as a multi-value structure, containing at least the policy identifier, policy type, and hit order. After this processing, the intervention impact identifier is no longer just a simple marker of whether or not an intervention occurred, but a traceable result reflecting which policy, when, and how it was affected, providing a basis for subsequent, more granular sample hierarchical segmentation.
[0039] In some implementations, logs with an additional policy name field can be marked as intervened, while those without the field can be marked as uninterrupted, thus dividing the samples into two basic levels.
[0040] For the intervention-affected level, a simple Boolean constraint can be configured, such as setting its availability in subsequent calculations to false, i.e., directly removing this part of the data. This simple binary partitioning and direct removal constraint method, while intuitive, may lead to the accidental deletion of a large amount of data containing some natural observational value, resulting in a sharp reduction in sample size, and it cannot provide fine-grained access control for different underlying algorithms of causal inference.
[0041] Therefore, in a more preferred implementation, the intervention status label may not be a single binary label, but rather divided into multiple sample levels based on the intervention intensity, the timing of the intervention, and the scope of the intervention. For example, it can be distinguished into a no-intervention level, a direct intervention level, a delayed-effect level, and a weakly correlated intervention level.
[0042] Among them, the no-intervention level indicates that the log records did not hit any policy impact in terms of object and time; the direct intervention level indicates that the log records occurred directly within the policy's effective period, and the behavior type has a direct correspondence with the intervention variable applied by the policy; the delayed impact level indicates that although the log records did not take effect immediately, they are still within the subsequent observation period where the policy may continue to have an impact; the weakly correlated intervention level indicates that the log records hit the policy-covered objects, but there is only an indirect correlation between the behavior type and the policy variable, and it cannot be directly identified as a strongly contaminated sample.
[0043] Correspondingly, availability constraints can be designed as access rules based on algorithms, stages, and purposes, rather than simply retaining or eliminating. For example, for propensity score estimation, only the uninvolved level and some weakly associated intervention levels can be allowed to participate; for conditional independence testing, the uninvolved level can be allowed as the main sample, while the delayed-effect level can participate in a reduced-weighted manner; for the final causal graph edge confidence correction, the difference information between all levels can be referenced.
[0044] Furthermore, the sample level can also include a sample weight field and an access permission mask field. The sample weight is used to express the credibility of the sample in the current algorithm stage, and the access permission mask is used to express whether the sample is allowed to be read by a certain algorithm module.
[0045] This hierarchical and constraint configuration approach avoids the waste of information caused by simple elimination, and allows different algorithmic steps to run within their respective suitable data boundaries, thereby more effectively isolating policy feedback pollution.
[0046] After dividing the dataset into tiers, the Pearson correlation coefficient between two variables (e.g., page dwell time and click-through rate) is calculated on both the intervened and uninterrupted datasets. If the correlation coefficient is only 0.1 in the uninterrupted dataset but spikes to 0.8 in the intervened dataset, it indicates that this strong association is artificially created by strategic intervention, and therefore this pair of variables is marked as a sensitive variable relationship.
[0047] In one specific embodiment, based on the aforementioned binary constraints and the identified relationships between sensitive variables, the causal inference algorithm skips the excluded data and forcibly deletes nodes containing sensitive variables when constructing the causal graph. While this coarse-grained constraint correction method can eliminate contamination to some extent, in the complex underlying logic of causal inference, simply deleting nodes would disrupt the topological integrity of the entire causal network and lacks targeted intervention mechanisms for core aspects such as propensity score estimation and conditional independence testing.
[0048] Therefore, step S5 further proposes to constrain and modify the execution process of the causal inference logic. Its core is not to filter or modify the result level after the causal graph is output, but to control the input of each calculation step in different algorithm stages of the causal inference pipeline according to the contaminated sample and the contaminated relationship, so as to block the contamination before it spreads.
[0049] Specifically, this constraint correction is controlled in the following three algorithmic stages: In the propensity score estimation stage, the highly uniform values of the processing variables in the intervention data prevent the propensity score model from effectively estimating the true intervention probability. In this stage, constraint correction primarily involves restricting samples from strong intervention states from entering the model training data to ensure that the processing variables maintain sufficient value fluctuation in the training set.
[0050] In the conditional independence test, strong associations between certain variables in the intervened data are artificially created, rather than natural causal dependencies. In this stage, constraint correction mainly involves identifying variable relationships sensitive to the intervention, preventing these identified variable pairs from entering the standard independence test process.
[0051] In the causal graph reconstruction stage, contamination information not completely eliminated in previous stages may still be passed into the final structure through candidate edges. In this stage, constraint correction mainly involves performing retention, weight reduction, or disconnection processing on candidate edges based on the judgment results of the first two stages.
[0052] The control of the above three stages is not an isolated operation, but a controlled execution process that proceeds step by step according to the algorithm's execution order. This mechanism does not simply delete the interfered data, but rather restricts contamination information to the early stages of algorithm execution, ensuring that subsequent stages operate on a controlled data basis, while preserving as much effective information from natural observation data as possible.
[0053] Through the above technical solution, the system establishes a traceable logical feedback chain on top of the physical data stream, making the correspondence between data records and policy interventions clear. By calculating the differences in correlation metrics between variables at different sample levels, the system can quantitatively identify which variable relationships are more sensitive to policy interventions, thereby avoiding misjudging intervention-distorted variables as confounding factors in their natural state.
[0054] Based on this, the system performs controlled execution of causal inference steps such as propensity score estimation and conditional independence testing, according to the availability constraints of the sample level and the identification results of intervention-sensitive relationships. This effectively suppresses the interference of artificial correlations introduced by strategy feedback on the reconstruction of the causal graph, so that the final output causal structure graph can more accurately reflect the behavioral transformation link of users in their natural state.
[0055] Furthermore, user behavior log data includes user identifiers, event timestamps, user profile features, contextual features, and business conversion tags; Historical policy intervention records include policy identifier, policy type, policy issuance timestamp, policy effective time window, set of user group identifiers covered by the policy, and intervention variables applied by the policy.
[0056] User behavior log data refers to the user's operational traces recorded by the system on the platform. This data includes user identifiers, used to uniquely identify individual users; event timestamps, used to record the time when the behavior occurred; user profile features, used to describe the user's static attributes; contextual features, used to describe the environmental information when the user's behavior occurred; and business conversion tags, used to mark whether the user has completed a specific business goal.
[0057] Various fields in user behavior log data need to be standardized and defined during the integration phase. For example, user profile features can include relatively stable attributes such as account status, membership level, and historical activity range; contextual features can include attributes related to the current behavior environment, such as access terminal type, page scenario, traffic source, and geographic information; and business conversion tags can be defined according to specific business requirements, such as whether a click was made, whether a purchase was made, whether an order was placed, and whether a user was retained.
[0058] User behavior log data can originate from platform event tracking systems, server-side log systems, or data lakes. For example, when a user browses products, clicks on ads, adds items to their cart, or completes payment on an e-commerce platform, the front-end application or back-end service generates a log record. This record includes: user identifier (e.g., user_id_A), event timestamp (e.g., 2023-10-26 14:30:05), user profile features (e.g., gender: female, age: 28, membership level: VIP3), contextual features (e.g., device type: mobile app, page scenario: product details page, traffic source: recommendation slot), and business conversion tags (e.g., whether clicked: 1, whether added to cart: 0, whether placed an order: 1).
[0059] After being cleaned, deduplicated, and standardized, this log data is stored in a distributed file system or data warehouse, waiting to be retrieved by causal discovery tasks.
[0060] Historical strategy intervention records refer to a collection of detailed information about strategies implemented by the platform in the past to influence user behavior. This information includes: strategy identifier, i.e., a unique identifier for each strategy; strategy type, referring to the classification of the strategy, such as coupon distribution, content push, or price adjustment; strategy issuance timestamp, indicating the time point when the strategy was issued or activated by the system; strategy effective time window, indicating the specific time period during which the strategy is expected to affect user behavior; set of user group identifiers covered by the strategy, referring to a list of unique identifiers for the user groups affected by the strategy; and intervention variables imposed by the strategy, referring to the specific user behavior or environmental factors that the strategy changed or introduced.
[0061] The intervention variables applied by the strategies in the historical strategy intervention records need to be clearly recorded as their objects, such as whether they affect the content exposure, pricing rights, message reach, or page ranking, so as to determine the correlation path between them and log behavior later.
[0062] In a specific embodiment, within a content recommendation platform, historical policy intervention records can be specifically represented as the following structured data: Policy identifier: For example, PUSH_CAMPAIGN_20230315_A (indicating a Class A push campaign launched on March 15, 2023).
[0063] Strategy type: For example, targeted content push (specifies a strategy to influence users by pushing specific content).
[0064] Policy issuance timestamp: For example, 2023-03-15 09:00:00 (the policy was issued by the system at 9:00 AM on that day).
[0065] Policy effective time window: For example, from 10:00:00 on March 15, 2023 to 23:59:59 on March 17, 2023 (the policy will take effect at 10:00 AM on March 15 and last until midnight on March 17).
[0066] The set of user group identifiers covered by the policy: for example, a URI pointing to a list of user IDs in distributed storage, or a reference to a Bloom filter instance that contains users who have been inactive for the past 30 days but have browsing history.
[0067] Intervention variables imposed by the strategy: For example, exposure of recommended content (indicating that the strategy influences user behavior by increasing the exposure of specific recommended content).
[0068] These records are generated and stored by the platform's online A / B testing system. During the data acquisition phase, the system extracts these records through API interfaces or database queries and integrates them with user behavior log data. For example, if a record in the user behavior log shows that user_X clicked on a recommended content at 14:30:00 on 2023-03-16, the system will use the aforementioned historical policy intervention records to determine whether user_X is within the coverage area of the PUSH_CAMPAIGN_20230315_A policy and whether the click time is within the policy's effective time window.
[0069] Through this specific implementation plan, the system achieves automated and standardized acquisition of multi-source input information.
[0070] Further, step S2 includes: S21: Build a distributed association mapping engine, using user identifier and event timestamp as the joint matching key; S22: For each user behavior log data, iterate through the historical policy intervention records that are effective within the current time window, verify whether the user identifier corresponding to the user behavior log data falls into the user identifier set of the historical policy intervention records, and determine whether the event timestamp falls within the policy effective time window or the delayed impact period. S23: When the user identifier corresponding to the user behavior log data falls into the set of user identifiers in the historical policy intervention records, and the event timestamp falls within the policy effective time window or the lag period, an intervention impact identifier is generated in the user behavior log data. The intervention impact identifier includes the matched strategy identifier and the intervention intensity attenuation coefficient calculated based on the time difference between the event timestamp and the strategy effective time window; S24: When a user behavior log data matches multiple historical policy intervention records within the current time window, the multiple policy identifiers are appended to the same intervention impact identifier.
[0071] Among them, the distributed association mapping engine refers to a large-scale data processing module built on a distributed computing framework (such as Spark or Flink), which can distribute massive log data and policy records to multiple computing nodes for parallel processing.
[0072] In a more specific implementation, user behavior log data can come from a data collection system, a server-side business log system, a message queue, or the results exported from an offline data warehouse. For example, when a user engages in browsing, clicking, adding to cart, placing an order, making a payment, or claiming a discount, the web application, mobile application, or server-side interface writes logs containing user identifiers, event timestamps, event types, session information, and device information to a Kafka message queue, which is then consumed in real time by a distributed association mapping engine. Historical policy intervention records can come from a policy platform, an operations and deployment platform, a recommendation system console, or a marketing system's policy release table. The records should include at least the policy identifier, policy name, policy effective time window, storage location of the user identifier set, policy type, and necessary lag effect configurations.
[0073] Using user identifiers and event timestamps as the combined matching key means that data alignment requires not only consistency of objects but also logical overlap between the time of the action and the time interval of the policy's effect. To ensure the stable use of the combined matching key, it is usually necessary to first unify the user identifiers in the logs, such as mapping device IDs, account IDs, and member IDs to a unified primary user identifier; at the same time, time zone alignment and millisecond or second-level standardization of event timestamps should be performed, and obviously abnormal future time or missing time records should be removed.
[0074] The multi-value appending mechanism refers to the use of arrays or lists in the data structure design of intervention impact identifiers. When a log hits multiple policy verification rules in succession, the later policy will not overwrite the earlier policy. Instead, all the policy identifiers that are hit will be recorded in sequence.
[0075] To facilitate data storage and subsequent retrieval, the intervention impact identifier can be stored using structured fields, such as a list of policy identifiers, a list of corresponding decay coefficients, the first hit time, and the hit source window. Alternatively, it can be mounted as a JSON object in the log extension fields.
[0076] Furthermore, during distributed execution, partitioning can be performed first according to user identifiers, and then sorted within each partition according to event timestamps. This can reduce cross-node data migration, prevent related logs and policy records of the same user from being distributed across too many nodes, improve matching efficiency, and reduce network transmission overhead.
[0077] Further, step S22 includes: S221: Expand the set of user identifiers from historical policy intervention records into a Bloom filter; S222: For each user behavior log data, check whether the user ID corresponding to the user behavior log data falls into the user ID set by querying the Bloom filter; S223: Obtain the preset lag impact threshold. The preset lag impact threshold is the maximum duration for which historical policy intervention records still have a residual impact on user behavior, calculated from the end time of the policy effective time window. S224: Calculate the time difference between the event timestamp and the end time of the policy effective time window. If the time difference is less than or equal to zero, the event timestamp is determined to fall within the policy effective time window. If the time difference is greater than zero and less than or equal to the preset lag effect threshold, the event timestamp is determined to fall within the lag effect period.
[0078] A Bloom filter is a probabilistic data structure used to test whether an element belongs to a set. It maps elements to a bit array using multiple hash functions. Although there is a very small false positive rate (i.e., elements that are not in the set may be incorrectly identified as being in the set), there will never be a false negative (i.e., elements that are in the set will never be incorrectly identified as not being in the set).
[0079] In one specific embodiment, the system first exports a set of user identifiers corresponding to a historical policy intervention record from the policy platform. This set can be a text file, a partition table, or a user list in a key-value database. Then, an offline task reads each user identifier one by one, performs multiple hash mappings on each identifier, writes the corresponding positions into a bit array, and finally generates a Bloom filter file. This file is then distributed to the local cache or shared memory area of each computing node in the distributed association mapping engine. In this way, when processing each user behavior log data, the node does not need to load the complete user list; it only needs to query the Bloom filter to quickly complete the initial screening.
[0080] If the business scenario is more sensitive to false positives, a secondary, more precise check can be added after the Bloom filter hits, such as reconfirming the data in the detailed user table or key-value cache, to reduce the impact of false positives while maintaining high throughput. The preset lag impact threshold is a buffer period set based on business experience or historical data statistics to quantify the lag impact after the strategy fails.
[0081] There are several ways to obtain the preset lag impact threshold. In an executable example, the system first counts the decline trajectory of key user behaviors after the end of historical activities according to strategy type. For example, it counts the changes in indicators such as click-through rate, order rate, and dwell time in the hours after the push ends. When the relevant indicators basically fall back to near the natural baseline after a certain time length, the time length is registered as the candidate lag impact threshold for the corresponding strategy type. Then, the operations personnel or algorithm configuration system combine historical stability, business tolerance, and sample size to confirm and write it into the strategy configuration table.
[0082] The calculation of the time difference involves converting the absolute time of the log occurrence into a relative time relative to the policy end time, thus providing a unified basis for subsequent state determination. To avoid time calculation errors, the same time zone and the same time precision are usually used, and standardized processing is performed for scenarios such as crossing days, weeks, and daylight saving time.
[0083] In one specific embodiment, a content recommendation platform issues a targeted push strategy for 10 million low-frequency active users, effective from 8 PM to midnight on Fridays. These 10 million users are selected from the platform's user profile system. First, an offline tagging task outputs a user list file. Then, a build task reads each user's identifier and generates a Bloom filter that occupies only a small amount of memory through multiple hash write operations. This Bloom filter is then distributed to each node in the real-time processing cluster.
[0084] Meanwhile, the system sets a preset lag threshold of 24 hours based on the half-life of historical push effects. This 24-hour period is not arbitrary, but is derived from the retrospective analysis of the effects of multiple targeted push campaigns on the platform: within one day after the end of the campaign, the click and dwell behaviors of users affected by the push are still significantly higher than the natural baseline, while after this period, the relevant behaviors gradually return to normal. Therefore, this is written into the strategy record as an executable threshold.
[0085] When processing a log generated at 10:00 AM on Saturday, the system first uses a Bloom filter with minimal memory overhead to confirm that the user belongs to the target audience. Then, it reads the policy end time and the log's event timestamp to calculate that the log occurred 10 hours after the policy ended. Since this time is after the policy ended but does not exceed the preset lag threshold of 24 hours, the system accurately determines that the log falls within the lag period.
[0086] If the log entry occurs after midnight on Sunday and exceeds the preset lag impact threshold, then even if the user belongs to the push notification group, it will no longer be marked according to the lag impact of this strategy, thereby avoiding the continued judgment of behavior that has returned to its natural state as polluted data.
[0087] The combination of these two aspects achieves the following results: the Bloom filter solves the computational and storage bottlenecks in matching large-scale user identifier sets, and the preset lag impact threshold compensates for the logical flaw of relying solely on the policy's effective time window in the time dimension. The combination of these two aspects enables the system to accurately identify implicitly contaminated data within the policy decay period in large-scale data scenarios, preventing it from being incorrectly classified into the untouched sample level. Because the purity of the untouched sample level is guaranteed, the data foundation upon which propensity score estimation and conditional independence tests based on this sample level rely is more reliable, and the accuracy of causal graph skeleton construction is also improved.
[0088] Furthermore, step S23 includes: S231: If the user identifier corresponding to the user behavior log data falls into the user identifier set of the historical policy intervention record, and the event timestamp falls within the policy effective time window or the delayed impact period, obtain the policy identifier of the matched historical policy intervention record. S232: Obtain the time difference between the event timestamp and the end time of the policy effective time window. If the time difference is less than or equal to zero, set the intervention intensity attenuation coefficient to 1. S233: If the time difference is greater than zero and less than or equal to the preset lag effect threshold, the intervention intensity attenuation coefficient shall be calculated according to the following formula: Intervention intensity attenuation coefficient = 1 - time difference / preset lag effect threshold; S234: Associate the strategy identifier and the intervention intensity attenuation coefficient with the corresponding user behavior log data to form the intervention impact identifier corresponding to the user behavior log data.
[0089] The intervention intensity decay coefficient is a continuous value between 0 and 1, used to characterize the degree of contamination of the behavior reflected in the current log by the policy. When the time difference is less than or equal to zero, it means that the behavior occurred directly within the policy's effective period, at which point the contamination level is highest, and the coefficient is set to the maximum value of 1. When the behavior occurs within the lag period, a linear decay model is used for calculation, meaning that as the time difference increases, the decay coefficient gradually decreases from 1 to 0. This calculation method transforms discrete time states into continuous intensity values, allowing each log entry to obtain an accurate contamination characterization indicator.
[0090] In one specific embodiment, the system generates a policy impact detail record for each user behavior log that hits the policy. This detail record includes information such as the policy identifier, the end time of the policy effective time window, the event timestamp, the preset lag impact threshold, and the intervention intensity attenuation coefficient.
[0091] When generating the intervention intensity attenuation coefficient, the system first determines whether the event timestamp is earlier than or equal to the end time of the policy's effective time window. If this condition is met, the intervention intensity attenuation coefficient for that log entry is directly set to 1. If the event timestamp is later than the end time of the policy's effective time window, the system further determines whether the difference between the event timestamp and the end time is within a preset lag effect threshold. If it is within the threshold, the corresponding attenuation coefficient is calculated according to the attenuation rule that the closer to the end time, the stronger the effect, and the closer to the end of the lag effect period, the weaker the effect (e.g., calculated using the formula attenuation coefficient = 1 - time difference / preset lag effect threshold). If the difference between the event timestamp and the end time exceeds the preset lag effect threshold, it indicates that the effect of the policy has completely dissipated, and the log entry is not considered a policy hit.
[0092] If the same log hits multiple policies at the same time, the system generates a corresponding policy impact detail record for each policy hit. Each record calculates the intervention intensity attenuation coefficient independently to avoid multiple policies sharing the same attenuation coefficient, which would mask the attenuation differences between different policies.
[0093] In some business scenarios, different types of strategies can be configured with different preset lag impact thresholds. For example, coupon strategies, content push strategies, and price discount strategies can each use their own configured threshold parameters, but the output format of all strategies still maintains a unified intervention impact identifier structure so that subsequent steps can be processed uniformly.
[0094] To facilitate subsequent reading and calculation in the causal inference process, the intervention intensity attenuation coefficient can be uniformly retained with a preset decimal precision and stored using the same rounding rules.
[0095] Furthermore, step S3 includes: S31: Read the intervention intensity attenuation coefficient in the intervention impact identifier of each user behavior log data; S32: Based on the comparison results between the intervention intensity attenuation coefficient and the preset direct effect threshold and preset impact dissipation threshold, each user behavior log data is marked as a no-intervention state label, a direct intervention state label, a delayed feedback state label, or an indeterminate state label. S33: Divide user behavior log data marked as non-intervention status into the non-intervention sample level, and configure the highest sample weight and full algorithm access permissions for the non-intervention sample level; User behavior log data marked with direct intervention status labels are divided into the direct intervention sample level, and an algorithm access permission mask is configured for the direct intervention sample level to disable the propensity score estimation and partial correlation calculation permissions. User behavior log data labeled as delayed feedback status are divided into delayed feedback sample level, sample weights of delayed feedback sample level are configured to be smaller than those of uninterrupted sample level, and algorithm access permission mask is configured to restrict participation in confounding factor search. User behavior log data marked with an "uncertain status" label are classified into the "uncertain sample" level, and the lowest sample weight is configured for the "uncertain sample" level.
[0096] The preset direct impact threshold and preset influence dissipation threshold are two key parameters used to segment continuous decay coefficient intervals. Typically, the direct impact threshold is set relatively high (e.g., 0.8) to identify heavily contaminated data, while the influence dissipation threshold is set relatively low (e.g., 0.2) to identify data where contamination has largely dissipated. Using these two thresholds, the system can discretize continuous decay coefficients into four state labels with clear business semantics, and thereby divide the data into four independent sample levels.
[0097] For each level, the system tailors sample weights and algorithm access permission masks, thus overlaying a robust logical access control network on top of the physical data. To illustrate the specific sources of the preset direct impact threshold and preset influence dissipation threshold, in one specific embodiment, the system first selects log samples from several historical policy activity periods, groups them according to the actual degree of contamination as verified by manual checks or backtracking, and then observes the impact of samples within different decay coefficient intervals on the stability of causal inference results. When samples in a certain interval significantly increase bias and disrupt the stability of propensity score estimation, the values near the upper bound of that interval are used as candidate values for the preset direct impact threshold. When samples in a certain interval have slight contamination but their impact on the results is already small, and deletion would significantly reduce the sample size, the values near the lower bound of that interval are used as candidate values for the preset influence dissipation threshold. After multiple rounds of offline verification, a set of thresholds that balances sample purity and sample size is selected and written to the configuration center.
[0098] Algorithm access permission masks can be stored using binary switches, enumerated labels, or permission lists. For example, the allowed and prohibited algorithm modules can be pre-registered for each sample level. When the algorithm task starts, the scheduler first reads the permission mask and then filters the input data.
[0099] Sample weights can also be written to the feature table as an additional field of the log.
[0100] In one specific embodiment, the system sets a preset direct effect threshold of 0.8 and a preset effect dissipation threshold of 0.2. These thresholds can be derived from offline verification results of historical policy activity samples: when the decay coefficient is higher than 0.8, the samples are usually still significantly dominated by the policy, and if they enter the core causal inference chain, they are prone to amplifying spurious correlations; when the decay coefficient is lower than 0.2, the residual effect of the samples is already weak, but there may still be a small amount of contamination that cannot be completely confirmed, so it is more suitable to place them in a conservative level of processing.
[0101] When a log entry is read and its decay coefficient is 1 (greater than 0.8), the system marks it as a direct intervention status label, classifies it into the direct intervention sample level, and explicitly writes "prohibited from participating in propensity score estimation" and "prohibited from participating in partial correlation calculation" into its permission mask.
[0102] When another log entry is read with a decay coefficient of 0.5 (between 0.2 and 0.8), the system marks it as a hysteresis feedback status label, classifies it into the hysteresis feedback sample level, sets its sample weight to 0.5 (lower than 1.0 for uninterrupted samples), and writes a restriction on its participation in confounding factor search in the permission mask.
[0103] When a log entry with a decay coefficient of 0.1 (less than 0.2) is read, the system classifies it into the "undetermined sample" level, assigns it the lowest weight of 0.1, and only allows it to participate in the final confidence smoothing calculation as marginal reference data. If a log entry does not match any historical policy intervention records, or if its event time is neither within the policy effective time window nor within the lag effect period, the system can directly mark it as an "uninterventional" state, classify it into the "uninterventional sample" level, assign it the highest sample weight, and grant it full algorithm access.
[0104] Through this layered processing, logs with different levels of contamination can enter different processing paths in subsequent processes, instead of being crudely retained or uniformly deleted.
[0105] Further, step S4 includes: S41: Obtain the set of variable pairs, which contains candidate variable pairs formed by combining variables from user profile features, context features, and business conversion tags in pairs; S42: Extract the values of each candidate variable from the uninterventional sample level and calculate the first correlation measure; extract the values of each candidate variable from the direct intervention sample level and calculate the second correlation measure. S43: Calculate the difference between the first correlation measure and the second correlation measure; S44: If the difference value is greater than the preset sensitivity threshold, then the candidate variable pair is marked as an intervention-sensitive relationship.
[0106] In the preceding steps, the system has divided user behavior log data into four sample levels. Based on this, the system constructs a set of candidate variable pairs from user profile features, contextual features, and business conversion tags through pairwise pairing to cover all possible combinations of variables that may be correlated.
[0107] Subsequently, the system utilizes the uninterrupted sample level and the directly intervened sample level, defined in the previous steps, as the reference sample under natural conditions and the experimental sample under policy intervention, respectively. For each pair of variables in the candidate variable pair set, the system calculates its association strength at both the uninterrupted sample level and the directly intervened sample level, obtaining a first correlation measure and a second correlation measure. By calculating the difference between the two, the system can quantify the degree to which policy intervention distorts the association relationship of the variable pair.
[0108] The system presets a sensitivity threshold. When the difference between a pair of variables exceeds this threshold, it indicates that the association strength between the pair of variables has changed significantly before and after the intervention, meaning that their association is greatly affected by the policy intervention. Therefore, the pair of variables is marked as intervention-sensitive.
[0109] The implementation principle is as follows: Causal dependencies that exist naturally typically maintain relatively stable strength across different sample levels; however, artificially created dependencies through policy intervention will exhibit significantly different strengths in the intervened samples compared to natural samples. Through cross-level comparisons, the system can identify which variable relationships are significantly distorted by policy intervention, thus pinpointing the contamination impact at the variable relationship level. In subsequent conditional independence tests, variable pairs marked as intervention-sensitive relationships will not enter the routine testing process, preventing distorted variable relationships from contaminating the causal graph while preserving naturally existing dependencies from being mistakenly deleted.
[0110] Further, step S42 includes: S421: For each candidate variable pair, take any one variable in the candidate variable pair as the first variable and the other variable in the candidate variable pair as the second variable; S422: From all sample records at the uninterrupted sample level, extract the values of the first and second variables in each sample record, calculate the first partial correlation coefficient between the first and second variables, and use it as the first correlation measure. S423: Extract the values of the first and second variables in each sample record from all sample records at the direct intervention sample level, and calculate the second partial correlation coefficient between the first and second variables as the second correlation measure.
[0111] In business environments with multiple variables, the correlation between two variables may be influenced by the transmission of other variables. Directly calculating the simple correlation coefficient between two variables can easily misjudge indirect correlations transmitted through other variables as direct correlations. Therefore, this step uses the partial correlation coefficient as a measure of the strength of the association between variables.
[0112] The partial correlation coefficient measures the net correlation between two variables, controlling for the influence of other variables. In actual calculations, for each candidate variable pair, the system treats that candidate variable pair as the variable to be analyzed and all other variables as control variables. The specific calculation method is as follows: First, a covariance matrix containing all variables is constructed. Then, by matrix inversion, the partial correlation coefficients corresponding to the candidate variable are solved at both the uninterventional sample level and the directly intervened sample level to obtain the first and second partial correlation coefficients.
[0113] By employing partial correlation coefficients, the system can more accurately assess the degree to which a strategic intervention distorts the direct relationship between two variables, after excluding the linear effects of other variables. When comparing the first and second partial correlation coefficients, the system is actually comparing the net effect of the strategic intervention on the direct association between the two variables after eliminating the influence of all other known factors.
[0114] Based on cross-level comparisons using partial correlation coefficients, the system can focus the identification of intervention-sensitive relationships on those directly distorted by policy intervention, reducing the probability of misjudgment due to the transmission effects of other variables. This ensures that the variable pairs ultimately labeled as intervention-sensitive relationships can more accurately reflect the artificial correlations introduced by policy intervention, providing a reliable basis for the controlled execution of subsequent causal inference.
[0115] In some preferred embodiments, in order to achieve more refined control in the causal inference process, the above-mentioned availability constraint based on the sample level only limits which algorithm steps samples of different states can participate in, which is a control method based on sample source.
[0116] Building upon this, deeper intervention measures are needed at the algorithm level: on the one hand, in the propensity score estimation stage, it is necessary to ensure that the processing variables retain sufficient value fluctuations in the training data so that the model can effectively estimate the intervention probability of different users; on the other hand, in the conditional independence test stage, it is necessary to prevent artificially created short-term strong associations from being misjudged as natural dependencies. To this end, this application, based on sample-level control, further introduces a deep controlled execution mechanism for the causal inference logic execution process, including intervention sensitivity determination for candidate variable pairs, detection, interception, or weight reduction of intervention-sensitive relationships, and implementation of differentiated retention strategies for candidate edges.
[0117] Specifically, step S5 includes: S51: Based on the sample weights and algorithm access permission masks of each sample level, select a subset of samples from each sample level that are allowed to participate in the bias score estimation, run the bias score estimation model on the sample subset, and obtain the bias score of each sample. S52: Based on the sample weights and algorithm access permission masks of each sample level, select the sample subsets that are allowed to participate in the conditional independence test from each sample level; S53: When performing conditional independence tests, for candidate variable pairs marked as having a sensitive relationship to intervention, intercept the routine statistical test process and disconnect or reduce the confidence weight of the candidate edge of the candidate variable pair in the skeleton of the causal graph. S54: Based on the propensity score, the results of the conditional independence test, and the results of disconnecting or reducing the weight of candidate edges, a causal structure graph is reconstructed to exclude the feedback pollution of the intervention strategy.
[0118] In real-world causal inference pipelines, propensity score estimation and conditional independence testing are two crucial underlying algorithms that are highly sensitive to data distribution. Indiscriminate use of mixed data can lead to a loss of heterogeneity in the processed variables within the sample, resulting in severe multicollinearity bias. Therefore, this specific implementation introduces the concepts of sample weights and algorithm access masking to achieve algorithm-level data access isolation.
[0119] The sample weight can be understood as an expression of the credibility of a sample in a certain stage of the algorithm. It is not necessarily fixed, but can vary with the sample level, the strength of the policy influence, and the purpose of the algorithm. For example, the same sample at the delay influence level may be given a lower weight in the propensity score estimation stage, but may be allowed to participate with a medium weight in the conditional independence test stage.
[0120] Algorithm access mask can be understood as a data visibility control marker for algorithm modules. It is used to answer questions such as whether a certain level of sample is allowed to be read by a certain algorithm, whether it is allowed to participate in training, and whether it is only allowed to participate in result verification.
[0121] For ease of implementation, access permission masks can be stored in a field-based configuration manner, such as attaching independent flags to each sample record to indicate whether access to propensity score estimation, conditional independence testing, or edge confidence correction is allowed; alternatively, they can be stored in a hierarchical unified configuration table manner, that is, first defining the access rules for each sample level to each algorithm module, and then filtering samples according to the rules at runtime.
[0122] By combining these two control variables, the system no longer sends all samples into the underlying algorithm at once, but instead allows different algorithms to only access the data range that they can handle and are suitable for processing, thereby reducing the risk of contamination transmission from the source.
[0123] For a specific example, in step S51, the system instantiates a logistic regression model as a propensity score estimation model. When preparing training data for this propensity score estimation model, the system filters data based on the algorithm access permission mask corresponding to each sample level. For example, a certain algorithm access permission mask indicates that samples at the directly intervened sample level are not visible to the propensity score estimation model.
[0124] At this point, when the system retrieves training data, it only allows samples from the uninterrupted sample level to enter the training set. Simultaneously, it substitutes the sample weights configured for the uninterrupted sample level into the logistic regression loss function for calculation. After the model completes its fit on this training set, it outputs a propensity score for each sample, which is the estimated probability that the sample will accept policy intervention.
[0125] Through the above data filtering and weighted training based on permission masks, the values of the processing variables in the training data on which the propensity score estimation model relies retain the normal fluctuations under natural conditions. The model can effectively estimate the intervention probability of different users and avoid model failure caused by highly consistent values of processing variables.
[0126] In steps S52 and S53, the controlled execution process of the conditional independence test is as follows: The system uses the PC algorithm as the causal graph construction algorithm. When constructing the causal graph, the PC algorithm first establishes a complete undirected graph containing all variables (i.e., there are candidate edges between every pair of all variables), and then uses the conditional independence test to determine whether each candidate edge should be retained.
[0127] When the PC algorithm reaches a candidate variable pair (such as coupon issuance and user activity), the system queries whether the variable pair has been marked as a sensitive relationship to intervention before initiating the independence test of the variable pair.
[0128] If the query results show that the variable pair has been marked as a sensitive relationship for intervention, the system triggers an interception mechanism: instead of performing partial correlation coefficient calculations and statistical significance tests on the variable pair, the system directly processes the candidate edge in the completely undirected graph. That is, the candidate edge is removed from the graph, or its confidence weight is reduced from its initial value (e.g., from 1.0 to 0.1), so that the confidence of the edge is significantly reduced in the subsequent causal graph reconstruction process, thereby avoiding artificially created strong associations from being mistakenly included in the final causal structure graph.
[0129] Finally, in step S54, the weighted scores calculated above and the intercepted and corrected test results are combined to generate the final directed acyclic graph.
[0130] To describe this process more clearly and completely, in a more detailed implementation example, the sample subset selection in step S51 can first read the sample hierarchy configuration table, and then check the access permission mask field on each sample record. Only samples that simultaneously meet the requirements of hierarchy permission and record permission are included in the propensity score estimation model. Before entering the propensity score estimation model, user profile features, context features, and historical behavior features can also be uniformly encoded and missing value handled to avoid affecting the estimation stability due to inconsistent field formats. The propensity score output by the model can then be written back to the sample records as auxiliary balancing information in the subsequent causal inference stage.
[0131] The sample subset selection in step S52 can adopt different rules than those in step S51. For example, it can allow some delayed-impact hierarchical samples to participate in order to retain more natural structural information, but still maintain strict restrictions on directly intervening hierarchical samples.
[0132] The disconnection or reduction of confidence weights in step S53 can be performed in stages based on the degree of sensitivity: if a candidate variable pair consistently exhibits a strong intervention-sensitive relationship across multiple time windows, the candidate edge is directly disconnected; if a candidate variable pair only exhibits moderate sensitivity in a local window, the candidate edge is retained but its priority for entering subsequent orientation stages is reduced. The advantage of this approach is that the system does not mechanically delete all sensitive relationships indiscriminately, but rather implements differentiated control based on the strength of contamination evidence, thereby suppressing spurious causal relationships while preserving as much true structural information as possible.
[0133] Through the above specific implementation methods, the system incorporates controlled execution logic into the execution process of the causal inference algorithm. In the propensity score estimation stage, the system filters training data based on sample weights and algorithm access permission masks, ensuring that the propensity score model is trained only on samples where the processing variables retain natural fluctuations, thus avoiding model estimation failure caused by highly consistent variable values in the intervened samples.
[0134] In the conditional independence test, the system intercepts or downweights the marked variable relationships based on the intervention sensitivity judgment results, avoiding the destruction of the causal graph topology caused by directly deleting variables, and preventing artificially created short-term associations from being incorrectly included in the causal graph.
[0135] Through the controlled execution of the above two steps, the artificial correlation introduced by the strategy intervention is blocked at different stages of the causal inference process, and the final causal structure diagram can more accurately reflect the real behavioral transformation path of users in their natural state.
[0136] Please refer to Figure 2 This application also proposes a data mining-based causal relationship discovery system for implementing the steps in any of the above methods. The system includes: Module 201: Acquire user behavior log data and historical policy intervention records; Generation module 202: Using user identifier and event timestamp as matching keys, it associates historical policy intervention records with user behavior log data to generate an intervention impact identifier for each user behavior log data; Segmentation Module 203: Based on the intervention impact identifier, identify the intervention status label of each user behavior log data, and divide the user behavior log data into different sample levels according to the intervention status label, and configure corresponding availability constraints for each sample level; Calculation module 204: Calculates the correlation metrics between variables on user behavior log data at different sample levels, and identifies the relationships of sensitive variables affected by intervention based on the differences in the correlation metrics between different sample levels; Reconstruction Module 205: Based on the availability constraints and sensitive variable relationships corresponding to each sample level, the execution process of causal inference logic is constrained and corrected to reconstruct a causal structure diagram that excludes policy intervention feedback pollution.
[0137] Among them, the acquisition module 201 refers to the module used to acquire user behavior log data and historical policy intervention records. Specifically, it can be implemented by means of data interface, file reading or database query. For example, user behavior log data can be pulled from the user behavior database through the API interface, and historical policy intervention records can be obtained from the policy management system.
[0138] The generation module 202 is used to associate historical policy intervention records with user behavior log data using user identifier and event timestamp as matching keys, and generate an intervention impact identifier for each user behavior log data. Specifically, it can be implemented using a distributed association mapping engine. For example, a distributed association mapping engine based on Spark or Hadoop can be built, using user identifier and event timestamp as joint matching keys to efficiently associate user behavior log data and historical policy intervention records.
[0139] The segmentation module 203 is used to identify the intervention status label of each user behavior log data based on the intervention impact identifier, and to segment the user behavior log data into different sample levels according to the intervention status label, configuring corresponding availability constraints for each sample level. For example, based on the comparison result of the intervention intensity attenuation coefficient and the preset threshold, the data is marked as no intervention, direct intervention, delayed feedback, or indeterminate status, and corresponding sample weights and algorithm access permissions are configured for each level.
[0140] The calculation module 204 is used to calculate the correlation metrics between variables on user behavior log data at different sample levels. Based on the differences in the correlation metrics between different sample levels, the module identifies the relationships of sensitive variables affected by the intervention. Specifically, it can be implemented using statistical analysis tools or machine learning algorithms. For example, it can calculate the partial correlation coefficients between variables at the uninterventional sample level and the directly intervened sample level, and identify the relationships of sensitive variables by comparing the differences in these coefficients.
[0141] The reconstruction module 205 is used to constrain and correct the execution process of causal inference logic based on the availability constraints and sensitive variable relationships corresponding to each sample level, and reconstruct a causal structure graph that excludes policy intervention feedback contamination. For example, when performing conditional independence tests, for variable pairs identified as having intervention-sensitive relationships, the conventional statistical test process is intercepted, and their candidate edges in the causal graph skeleton are disconnected or their confidence weights are reduced.
[0142] Specifically, this system aims to address the challenges faced by causal discovery algorithms when processing a mixture of policy-interventional data and natural observation data. By finely identifying and processing the effects of interventions, it reconstructs a more accurate causal structure diagram.
[0143] The above description is merely an embodiment of this application and is not intended to limit the scope of protection of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application.
Claims
1. A causal relationship discovery method based on data mining, characterized in that, The method includes the following steps: S1: Obtain user behavior log data and historical policy intervention records; S2: Using user identifier and event timestamp as matching keys, associate the historical policy intervention records with the user behavior log data to generate an intervention impact identifier for each piece of user behavior log data; the user behavior log data includes user identifier, event timestamp, user profile features, context features, and business conversion tags; the historical policy intervention records include policy identifier, policy type, policy issuance timestamp, policy effective time window, set of user group identifiers covered by the policy, and intervention variables applied by the policy; S3: Based on the intervention impact identifier, identify the intervention status label of each user behavior log data, and divide the user behavior log data into different sample levels according to the intervention status label, and configure corresponding availability constraints for each sample level; Step S3 includes: S31: Read the intervention intensity attenuation coefficient from the intervention impact identifier of each user behavior log data; S32: Based on the comparison results between the intervention intensity attenuation coefficient and the preset direct effect threshold and the preset influence dissipation threshold, each piece of user behavior log data is marked as a no-intervention state label, a direct-intervention state label, a delayed feedback state label, or an undetermined state label. S33: Divide the user behavior log data marked with the non-intervention state label into the non-intervention sample level, and configure the highest sample weight and full algorithm access permissions for the non-intervention sample level; The user behavior log data marked with the direct intervention status label is divided into the direct intervention sample level, and an algorithm access permission mask is configured for the direct intervention sample level to disable the permissions for propensity score estimation and partial correlation calculation. The user behavior log data marked with the lag feedback status label is divided into the lag feedback sample level, the lag feedback sample level is configured with a sample weight smaller than that of the uninterrupted sample level, and an algorithm access permission mask that restricts participation in the confounding factor search is configured. The user behavior log data marked with the "undeterminable status" label is divided into the "undeterminable sample level", and the lowest sample weight is configured for the "undeterminable sample level". S4: Calculate the correlation metrics between variables on the user behavior log data at different sample levels, and identify the relationships between sensitive variables affected by the intervention based on the differences in the correlation metrics between different sample levels. S5: Based on the availability constraints corresponding to each sample level and the relationship between the sensitive variables, the execution process of the causal inference logic is constrained and modified to reconstruct the causal structure diagram that excludes the interference of strategy feedback pollution.
2. The causal relationship discovery method based on data mining according to claim 1, characterized in that, Step S2 includes: S21: Build a distributed association mapping engine, using user identifier and event timestamp as the joint matching key; S22: For each user behavior log data, iterate through the historical policy intervention records that are effective within the current time window, verify whether the user identifier corresponding to the user behavior log data falls into the user identifier set of the historical policy intervention records, and determine whether the event timestamp falls within the policy effective time window or the delayed impact period. S23: When the user identifier corresponding to the user behavior log data falls into the user identifier set of the historical policy intervention record, and the event timestamp falls within the policy effective time window or the lag effect period, an intervention impact identifier of the user behavior log data is generated. The intervention impact identifier includes the matched strategy identifier and the intervention intensity attenuation coefficient calculated based on the time difference between the event timestamp and the strategy effective time window; S24: When a user behavior log data matches multiple historical policy intervention records within the current time window, the multiple policy identifiers are appended to the same intervention impact identifier.
3. The causal relationship discovery method based on data mining according to claim 2, characterized in that, Step S22 includes: S221: Expand the set of user identifiers in the historical policy intervention records into a Bloom filter; S222: For each user behavior log data, the user identifier corresponding to the user behavior log data is checked by querying the Bloom filter to see if it falls into the user identifier set; S223: Obtain a preset lag impact threshold, wherein the preset lag impact threshold is the maximum duration for which the historical policy intervention records still have a residual impact on user behavior, calculated from the end time of the policy effective time window; S224: Calculate the time difference between the event timestamp and the end time of the policy effective time window. If the time difference is less than or equal to zero, it is determined that the event timestamp falls within the policy effective time window. If the time difference is greater than zero and less than or equal to the preset lag effect threshold, it is determined that the event timestamp falls within the lag effect period.
4. The causal relationship discovery method based on data mining according to claim 3, characterized in that, Step S23 includes: S231: If the user identifier corresponding to the user behavior log data falls into the user identifier set of the historical policy intervention record, and the event timestamp falls within the policy effective time window or the lag effect period, obtain the policy identifier of the matched historical policy intervention record. S232: Obtain the time difference between the event timestamp and the end time of the policy effective time window. If the time difference is less than or equal to zero, set the intervention intensity attenuation coefficient to 1. S233: If the time difference is greater than zero and less than or equal to the preset lag effect threshold, the intervention intensity attenuation coefficient is calculated according to the following formula: Intervention intensity attenuation coefficient = 1 - time difference / preset lag effect threshold; S234: Associate the strategy identifier and the intervention intensity attenuation coefficient with the corresponding user behavior log data to form the intervention impact identifier corresponding to the user behavior log data.
5. The causal relationship discovery method based on data mining according to claim 1, characterized in that, Step S4 includes: S41: Obtain a set of variable pairs, which includes candidate variable pairs formed by combining variables from user profile features, context features, and business conversion tags in pairs; S42: Extract the variable values corresponding to each candidate variable from the uninterrupted sample level and calculate the first correlation measure; extract the variable values corresponding to each candidate variable from the directly intervened sample level and calculate the second correlation measure; S43: Calculate the difference between the first correlation measure and the second correlation measure; S44: If the difference value is greater than the preset sensitivity threshold, then the candidate variable pair is marked as an intervention-sensitive relationship.
6. The causal relationship discovery method based on data mining according to claim 5, characterized in that, Step S42 includes: S421: For each candidate variable pair, take any one variable in the candidate variable pair as the first variable and the other variable in the candidate variable pair as the second variable; S422: Extract the values of the first variable and the second variable in each sample record from all sample records of the uninterrupted sample level, and calculate the first partial correlation coefficient between the first variable and the second variable as the first correlation measure. S423: Extract the values of the first variable and the second variable in each sample record from all sample records at the direct intervention sample level, and calculate the second partial correlation coefficient between the first variable and the second variable as the second correlation measure.
7. The causal relationship discovery method based on data mining according to claim 6, characterized in that, Step S5 includes: S51: Based on the sample weights and algorithm access permission masks of each sample level, select a subset of samples that are allowed to participate in the bias score estimation from each sample level, run the bias score estimation model on the sample subset, and obtain the bias score of each sample. S52: Based on the sample weights and algorithm access permission masks of each sample level, select a subset of samples from each sample level that are allowed to participate in the conditional independence test; S53: When performing the conditional independence test, for candidate variable pairs marked as the intervention-sensitive relationship, the routine statistical test process is intercepted, and the candidate edges of the candidate variable pairs in the causal graph skeleton are disconnected or their confidence weights are reduced. S54: Based on the propensity score, the result of the conditional independence test, and the result of disconnecting or reducing the weight of the candidate edges, reconstruct the causal structure graph that excludes the interference of policy intervention feedback pollution.
8. A causal relationship discovery system based on data mining, characterized in that, The system, for implementing the steps of any one of the methods of claims 1-7, comprises: Acquisition module: Acquires user behavior log data and historical policy intervention records; The generation module associates the historical policy intervention records with the user behavior log data using the user identifier and event timestamp as matching keys, generating an intervention impact identifier for each piece of user behavior log data. The user behavior log data includes the user identifier, event timestamp, user profile features, context features, and business conversion tags. The historical policy intervention records include the policy identifier, policy type, policy issuance timestamp, policy effective time window, set of user group identifiers covered by the policy, and intervention variables applied by the policy. Segmentation Module: Based on the intervention impact identifier, identify the intervention status label of each user behavior log data, and divide the user behavior log data into different sample levels according to the intervention status label, and configure corresponding availability constraints for each sample level; The segmentation module is also used to read the intervention intensity attenuation coefficient in the intervention impact identifier of each user behavior log data; Based on the comparison results between the intervention intensity attenuation coefficient and the preset direct effect threshold and preset impact dissipation threshold, each piece of user behavior log data is marked as a no-intervention state label, a direct-intervention state label, a delayed feedback state label, or an undetermined state label. The user behavior log data marked with the non-intervention state label is divided into the non-intervention sample level, and the non-intervention sample level is configured with the highest sample weight and full algorithm access permissions. The user behavior log data marked with the direct intervention status label is divided into the direct intervention sample level, and an algorithm access permission mask is configured for the direct intervention sample level to disable the permissions for propensity score estimation and partial correlation calculation. The user behavior log data marked with the lag feedback status label is divided into the lag feedback sample level, the lag feedback sample level is configured with a sample weight smaller than that of the uninterrupted sample level, and an algorithm access permission mask that restricts participation in the confounding factor search is configured. The user behavior log data marked with the "undeterminable status" label is divided into the "undeterminable sample level", and the lowest sample weight is configured for the "undeterminable sample level". Calculation module: Calculates correlation metrics between variables on user behavior log data at different sample levels, and identifies sensitive variable relationships affected by intervention based on the differences in correlation metrics between different sample levels; Reconstruction Module: Based on the availability constraints corresponding to each sample level and the relationship between the sensitive variables, the execution process of the causal inference logic is constrained and corrected to reconstruct the causal structure diagram that excludes the interference of policy intervention feedback pollution.
Citation Information
Patent Citations
Complex task-based high-quality pseudo-annotation data set construction method
CN120297445A
Industrial equipment remote operation and maintenance management method and system based on digital twinning
CN121500847A