Causal inference method based on problem disassembly and algorithm matching

The method decomposes cause-and-effect inference into stages for accurate and adaptable causal relationship discovery, addressing flexibility and applicability issues in existing methods, ensuring reliable causal analysis across diverse scenarios.

CN120317367APending Publication Date: 2025-07-15EVALUATION & DEMONSTRATION RES CENT OF THE CHINESE PEOPLES LIBERATION ARMY ACAD OF MILITARY SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510386648.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

The existing causal inference methods lack flexibility and applicability, making it difficult to accurately mine and analyze the complex causal relationship between data under different application scenarios and usage conditions.

Method used

The causal inference problem is broken down into four major categories of phased sub-problems, and corresponding algorithm matching rules are designed, including causal relationship identification, causal effect estimation, intervention effect prediction, traceability analysis and other steps. Combined with data preprocessing and a variety of causal graph drawing methods, a structural causal model and potential result framework are used to estimate and test causal effect.

Benefits of technology

The full coverage of the causal inference process is achieved, ensuring the accuracy of causal relationship identification and effect calculation, and having good adaptability and reliability, and providing a variety of causal discovery means in complex systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120317367A_ABST
    Figure CN120317367A_ABST
Patent Text Reader

Abstract

The invention discloses a cause and effect inference method based on problem disassembly and algorithm matching, and relates to the technical field of cause and effect inference methods, and the method comprises the following steps: S1, defining a cause and effect inference problem; according to the cause-and-effect inference method based on problem disassembly and algorithm matching, a cause-and-effect inference problem is disassembled into four types of sub-problems, and a complete cause-and-effect inference algorithm matching rule is designed, so that a complex problem in a cause-and-effect inference process can be effectively solved; through six steps of causal inference problem definition, causal inference data set generation, causal relationship identification and inspection, causal effect estimation and inspection, anti-fact inference calculation and causal traceability analysis, full coverage of four types of sub-problems of causal inference is realized, and multiple causal relationship discovery means can be provided. And a suitable causal effect evaluation algorithm is automatically matched for each pair of causal variables, so that the accuracy of causal inference results such as causal relationship identification and causal effect calculation is ensured, and good adaptability is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of causal inference methods, and specifically to a causal inference method based on problem decomposition and algorithm matching. Background Art

[0002] In modern data science, understanding the causal relationships between variables is crucial for predictive analysis, decision support, and policy making. Traditional statistical methods often can only explain the correlations between variables, but cannot directly prove the existence of causal relationships. In recent years, with the improvement of computing power and the development of related algorithms, causal inference techniques have developed rapidly. However, most of the existing causal inference methods are targeted at specific scenarios or assumed conditions, lacking flexibility and applicability. Therefore, under the premise of facing different application scenarios and usage conditions, how to make good use of existing experimental data to accurately mine and analyze the complex causal relationships between data has become an important topic in causal inference. In order to effectively mine and verify the causal relationships between variables in various usage scenarios and problem backgrounds, a causal inference method based on problem decomposition and algorithm matching is proposed. Summary of the Invention

[0003] Aiming at the deficiencies of the prior art, the present invention provides a causal inference method based on problem decomposition and algorithm matching, which solves the problems raised in the above background art.

[0004] To achieve the above objectives, the present invention is realized through the following technical solutions: A causal inference method based on problem decomposition and algorithm matching, comprising the following steps:

[0005] S1: Definition of Causal Inference Problem

[0006] Decompose the causal inference problem into four major categories of phased sub-problems. The solution of different sub-problems corresponds to different steps of the method of the present invention, specifically as follows:

[0007] S11: Causal Relationship Identification Sub-problem

[0008] The causal relationship identification sub-problem is to solve the problem of who is the cause and who is the effect, and it is necessary to determine the independent variable, dependent variable, confounding variable, etc., and construct a causal graph, corresponding to S1 to S3 of the method of the present invention;

[0009] S12: Causal Effect Estimation Sub-problem

[0010] After determining the causal relationship, the causal effect estimation sub-problem needs to analyze the degree of influence of the independent variable on the dependent variable, that is, calculate the causal effect value of the independent variable on the dependent variable, corresponding to S1 to S4 of the method of the present invention;

[0011] S13: Intervention Effect Prediction Sub-problem

[0012] The intervention effect prediction sub - problem is a counterfactual reasoning problem, which is the consideration of the results of causal relationship identification and causal effect calculation, analyzing the possible impact of changes in the independent variable on the dependent variable, corresponding to S1 to S5 of the method of the present invention;

[0013] S14: Tracing analysis sub - problem

[0014] The tracing analysis sub - problem is to trace back and analyze the key indicators affecting the dependent variable after determining the causal relationship and calculating the causal effect, so as to form the analysis conclusion of the current experimental plan and the optimization suggestions for the subsequent experimental plan, corresponding to S4 and S6 of the method of the present invention;

[0015] S2: Causal inference dataset generation

[0016] Import and parse the original experimental data, provide data pre - processing operations such as outlier removal, normalization, data interpolation, etc., to generate the dataset used for the current causal inference analysis, so as to ensure data quality and consistency.

[0017] S3: Causal relationship discovery and testing

[0018] Based on the generated causal inference dataset, draw a causal graph to solve the causal relationship identification sub - problem in the causal inference process. The inspection and identification of the causal graph are as follows:

[0019] S31: Causal graph construction

[0020] Provide three methods for drawing causal graphs, which can be selected according to the actual usage scenarios:

[0021] S311: Manual drawing

[0022] Manual drawing is applicable to the situation where experts have a full understanding of the causal relationship between research objects in the current scenario and the number of variables is not large. Manually assign the causal relationship of each variable and draw the causal graph. If the relationship between variables in the scenario is clear and the number is small, the causal graph can be manually drawn according to expert experience.

[0023] S312: Semi - automatic drawing

[0024] Semi - automatic drawing is applicable to the situation where experts have a full understanding of the causal relationship between research objects in the current scenario and the number of variables is relatively large. It can automatically identify the variables with correlation relationships, and only the direction of the causal relationship needs to be determined manually. If the variable relationship in the scenario is clear, but the number is relatively large, manual drawing is time - consuming at this time. Extract the potential causal relationships between each variable through correlation calculation to obtain an undirected association relationship graph, and then the expert adds the causal relationship direction and modifies and improves the obtained causal graph.

[0025] S313: Fully automatic drawing

[0026] Fully automatic drawing is applicable to the situation where experts are not very familiar with the research object of this study. It can automatically discover causal relationships and draw causal structure diagrams. The fully automatic drawing provides the PC algorithm and the causal entropy algorithm. These two algorithms automatically complete the discovery of causal relationships and the determination of directions, and automatically draw them into causal diagrams.

[0027] S32: Causal relationship test

[0028] For the drawn causal diagram and the causal relationships it contains, it is necessary to test the correctness of the drawn causal diagram structure and causal relationships, as follows:

[0029] S321: Test for the correctness of the causal diagram structure

[0030] Perform a directed acyclic graph detection (DAG completeness test) on the drawn causal diagram. If a cycle is detected, the cycle will be marked in red on the causal diagram drawing and editing interface. With the support of experts, change the direction of the causal pairs in the cycle or remove the causal relationships in the cycle.

[0031] S322: Test for the correctness of causal relationships

[0032] For any pair of causal relationships x - y in the causal diagram, calculate the absolute value of the difference in probability distributions of this pair of causal relationships based on the intervention calculus (Do calculus), and use it as the result of the Do calculus and judge whether it is greater than the threshold. If it is greater than the threshold, retain this causal pair; if it is less than the threshold, remove this causal pair and automatically update and adjust the causal diagram.

[0033] S4: Causal effect estimation and test

[0034] Combining the causal structure identification based on the structural causal model and the causal effect estimation based on the potential outcome framework can solve the sub - problem of causal effect estimation in the causal inference process and achieve the quantitative calculation of causal relationships.

[0035] S41: Causal structure identification

[0036] The definitions of the applicable criteria for causal variables using the structural causal model include three types: the backdoor criterion, the frontdoor criterion, and instrumental variables. These three criteria cover the three major characteristics of causal structures, and thus define three causal variable structures: causal structures that conform to the backdoor criterion, causal structures that conform to the frontdoor criterion, and causal structures that conform to instrumental variables. Combining means such as d - separation and intervention calculus to identify the causal structure to which each pair of causal variables belongs can effectively understand the specific composition of causal relationships between causal variables, control the bias brought by confounding variables, and ensure the accuracy and reliability of the subsequent causal effect calculation results.

[0037] S42: Causal effect calculation

[0038] The causal effect value is calculated using the causal effect estimation algorithm based on the potential outcome framework. The matching rules of the causal effect algorithms under three causal structures are designed as Figure 1 shown;

[0039] S421: Causal effect estimation based on the backdoor criterion

[0040] Three causal effect estimation algorithms based on the backdoor criterion are provided: direct calculation method, propensity score stratification, and propensity score matching. The usage rules corresponding to the three algorithms are designed as follows:

[0041] S4211: Direct calculation method: The set of backdoor adjustment variables is an empty set;

[0042] S4212: Propensity score stratification: The set of backdoor adjustment variables is not an empty set and the number exceeds 5. The dependent variable is one-dimensional and follows a binary distribution or can be binary discretized. The number of propensity score strata is defaulted to 5 layers, and the user can uniformly change the number of strata as needed;

[0043] S4213: Propensity score matching: The set of backdoor adjustment variables is not an empty set and the number does not exceed 5. The dependent variable is one-dimensional and follows a binary distribution or can be binary discretized.

[0044] S422: Causal effect estimation based on the frontdoor criterion

[0045] The causal effect evaluation algorithm based on the frontdoor criterion supports the two-layer linear regression algorithm based on the frontdoor criterion.

[0046] S423: Causal effect estimation based on instrumental variables

[0047] The causal effect evaluation algorithm based on instrumental variables supports the two-stage least squares method and the binary instrument / Wald estimation algorithm based on instrumental variables. The matching rules corresponding to the algorithms are designed as follows:

[0048] S4231: Binary instrument / Wald estimation: Both the set of instrumental variables and the dependent variable are one-dimensional. Among them, if the instrumental variable follows a binary distribution, the Wald estimation expression is used to calculate the causal effect; if the instrumental variable does not follow a binary distribution, the variance expression is used to calculate the causal effect;

[0049] S4232: Two-stage least squares method: The dimension of the instrumental variable or the dependent variable is greater than one-dimensional. The instrumental variable is used to solve the endogeneity bias caused by confounding variables.

[0050] S43: Test of causal effect evaluation results

[0051] Since the process of drawing causal diagrams depends partly on experts' subjective experience, in order to objectively verify the reliability of causal inference analysis in complex systems, the method of refutation is used to verify the credibility of the calculation results. If the test is passed, it means that the calculation results of causal effects have a high degree of credibility; otherwise, the calculation results of causal effects are considered untrustworthy.

[0052] S431: Verification of data subset

[0053] Use data subset verification to re-establish the evaluation data and recalculate the causal effect estimation results. If the test results are not much different from the calculation results and the test confidence level ≥ 0.9, it can be considered that the test is passed.

[0054] S432: Adding random confounding factors

[0055] Add a randomly generated confounding variable. For a causal relationship X on Y, add a random variable that is independent of both causal variables as a confounding factor to this causal relationship and recalculate the causal effect of X on Y. If the test results are not much different from the calculation results and the test confidence level ≥ 0.9, it can be considered that the test is passed.

[0056] S433: Placebo intervention

[0057] Randomly delete a part of the data. The new data is a random subset of the original data. Use the subset to recalculate the causal effect of X on Y. If the test results are very different from the calculation results and the test confidence level ≥ 0.9, it can be considered that the test is passed.

[0058] S5: Counterfactual inference calculation

[0059] Use the counterfactual calculation algorithm to calculate the counterfactual results, which can solve the sub-problem of predicting intervention effects in the process of causal inference. This algorithm needs to set the input data (sample data of a certain experiment) and the set value of the dependent variable (set the value of the dependent variable according to user needs). After giving the input data, causal diagram structure and effect calculation results, the counterfactual calculation results are automatically given according to the processing flow. The counterfactual calculation function in the effect estimation module of the prototype system is realized by calling the counterfactual calculation algorithm through the model library.

[0060] The best application scenario of counterfactual inference is after obtaining the experimental results. If the experimental results are not ideal, by keeping other variables unchanged, analyze how to adjust the parameter variables so that the result variable reaches the expected value, that is, under the premise of limited resources, accurately adjust the relevant parameter variables to simulate the intervention effect that can only be seen by re-conducting the experiment.

[0061] S6: Causal traceability analysis

[0062] To further analyze the association between the causal effect estimation results and relevant decision-making elements, it is necessary to further analyze the calculation results of the causal effect. Conduct a layer-by-layer analysis of the single-experiment data in the scenario, find the key reasons affecting the target outcome variable, and display the results of the retrospective analysis in the form of a causal quantitative decision tree. The causal traceability analysis uses the regression coefficient as the basis for determining the importance of variables, and at the same time considers multiple independent variables to analyze the influence degree of multiple independent variables on the outcome variable. There are two recommended usage scenarios for causal traceability analysis: under the condition of limited resources, causal traceability analysis can give suggestions for the initial setting of parameters; when the experimental results are not ideal, it can find the variables that cause the result deviation and give optimization suggestions for subsequent experiments. The schematic diagram of causal traceability analysis is as Figure 2 shown.

[0063] The present invention provides a causal inference method based on problem decomposition and algorithm matching, which has the following

[0064] beneficial effects:

[0065] This causal inference method based on problem decomposition and algorithm matching can effectively solve the complex problems in the causal inference process by decomposing the causal inference problem into four types of sub-problems and designing a complete set of causal inference algorithm matching rules. Through six major steps: causal inference problem definition, causal inference data set generation, causal relationship identification and testing, causal effect estimation and testing, counterfactual inference calculation, and causal traceability analysis, it realizes full coverage of the four major types of sub-problems in causal inference, can provide a variety of means for discovering causal relationships, automatically match a suitable causal effect evaluation algorithm for each pair of causal variables, ensure the accuracy of causal inference results such as causal relationship identification and causal effect calculation, and have good adaptability. BRIEF DESCRIPTION OF THE DRAWINGS

[0066] Figure 1 is the design diagram of the causal effect algorithm matching rule of the present invention;

[0067] Figure 2 is the causal traceability analysis diagram of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0068] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments.

[0069] Please refer to Figures 1 to 2 , the present invention provides a technical solution: a causal inference method based on problem decomposition and algorithm matching, including the following steps:

[0070] S1: Causal inference problem definition

[0071] Decompose the causal inference problem into four major categories of phased sub - problems. The solution of different sub - problems corresponds to different steps of the method of the present invention, which are specifically as follows:

[0072] S11: Causal relationship identification sub - problem

[0073] The causal relationship identification sub - problem is to solve the problem of who is the cause and who is the effect. It is necessary to determine independent variables, dependent variables, confounding variables, etc., and construct a causal diagram, corresponding to S1 to S3 proposed by the present invention;

[0074] S12: Causal effect estimation sub - problem

[0075] After determining the causal relationship, the causal effect estimation sub - problem needs to analyze the degree of influence of the independent variable on the dependent variable, that is, calculate the causal effect value of the independent variable on the dependent variable, corresponding to S1 to S4 of the method of the present invention;

[0076] S13: Intervention effect prediction sub - problem

[0077] The intervention effect prediction sub - problem is a counterfactual reasoning problem, which is a consideration of the results of causal relationship identification and causal effect calculation, and analyzes the possible impact of changes in the independent variable on the dependent variable, corresponding to S1 to S5 of the method of the present invention;

[0078] S14: Traceability analysis sub - problem

[0079] After determining the causal relationship and calculating the causal effect, the traceability analysis sub - problem needs to retrospectively analyze the key indicators affecting the dependent variable to form the analysis conclusion of the current experimental plan and the optimization suggestions for the subsequent experimental plan, corresponding to S4 and S6 of the method of the present invention;

[0080] S2: Generation of causal inference dataset

[0081] Import and parse the original experimental data, provide data pre - processing operations such as outlier removal, normalization, data interpolation, etc., to generate the dataset used for this causal inference analysis to ensure data quality and consistency;

[0082] S3: Causal relationship discovery and testing

[0083] Based on the generated causal inference dataset, draw a causal diagram to solve the causal relationship identification sub - problem in the causal inference process. The inspection and identification of the causal diagram are as follows:

[0084] S31: Causal diagram construction

[0085] Provide three methods for drawing causal diagrams, which can be selected according to the actual usage scenario:

[0086] S311: Manual drawing

[0087] Manual drawing is applicable when experts have a full grasp of the causal relationships between research objects in the current scenario and the number of variables is not large. The causal relationships of each variable are artificially assigned and the causal diagram is drawn. If the relationships between variables in the scenario are clear and the number is small, the causal diagram can be manually drawn based on expert experience.

[0088] S312: Semi-automatic drawing

[0089] Semi-automatic drawing is applicable when experts have a full grasp of the causal relationships between research objects in the current scenario and the number of variables is relatively large. It can automatically identify variables with correlation relationships, and only the direction of the causal relationship needs to be determined manually. If the relationships between variables in the scenario are clear but the number is relatively large, manual drawing is time-consuming at this time. The potential causal relationships between each variable are extracted through correlation calculation to obtain an undirected association relationship diagram, and then the expert adds the direction of the causal relationship and modifies and improves the obtained causal diagram.

[0090] S313: Fully automatic drawing

[0091] Fully automatic drawing is applicable when experts are not very familiar with the research objects in this study. It can automatically discover causal relationships and draw a causal structure diagram. The fully automatic drawing provides the PC algorithm and the causal entropy algorithm. These two algorithms automatically complete the discovery of causal relationships and the determination of directions, and automatically draw them into a causal diagram.

[0092] S32: Causal relationship test

[0093] For the drawn causal diagram and the causal relationships it contains, it is necessary to test the correctness of the structure and causal relationships of the drawn causal diagram, as follows:

[0094] S321: Test for the correctness of the causal diagram structure

[0095] Perform a directed acyclic graph detection (DAG completeness test) on the drawn causal diagram. If a cycle is detected, the cycle will be marked in red on the causal diagram drawing and editing interface. With the support of experts, change the direction of the causal pairs in the cycle or remove the causal relationships in the cycle.

[0096] S322: Test for the correctness of causal relationships

[0097] For any pair of causal relationships x - y in the causal diagram, calculate the absolute value of the difference in probability distributions of this pair of causal relationships based on the intervention calculus (Do calculus) as the Do calculus result and judge whether it is greater than the threshold. If it is greater than the threshold, retain this causal pair; if it is less than the threshold, remove this causal pair and automatically update and adjust the causal diagram.

[0098] S4: Causal effect estimation and test

[0099] Combining causal structure identification based on the structural causal model and causal effect estimation based on the potential outcomes framework can solve the sub-problem of causal effect estimation in the causal inference process and achieve quantitative calculation of causal relationships.

[0100] S41: Causal Structure Identification

[0101] The definitions of the applicability criteria for causal variables using the structural causal model include three types: the backdoor criterion, the frontdoor criterion, and instrumental variables. These three criteria cover the three major characteristics of causal structures, and thus define three causal variable structures: causal structures that conform to the backdoor criterion, causal structures that conform to the frontdoor criterion, and causal structures that conform to instrumental variables. Combining means such as d-separation and intervention calculus to identify the causal structure to which each pair of causal variables belongs can effectively understand the specific causal relationship composition between causal variables, control the bias brought by confounding variables, and ensure the accuracy and credibility of the subsequent causal effect calculation results.

[0102] S42: Causal Effect Calculation

[0103] Use the causal effect estimation algorithm of the potential outcomes framework to calculate the causal effect value. The matching rules of the causal effect algorithms under the three causal structures are designed as Figure 1 shown.

[0104] S421: Causal Effect Estimation Based on the Backdoor Criterion

[0105] Provide three causal effect estimation algorithms based on the backdoor criterion: direct calculation method, propensity score stratification, and propensity score matching. The usage rules corresponding to the three algorithms are designed as follows:

[0106] S4211: Direct Calculation Method: The backdoor adjustment variable set is an empty set;

[0107] S4212: Propensity Score Stratification: The backdoor adjustment variable set is not an empty set and the number exceeds 5. The dependent variable is one-dimensional and follows a binary distribution or can be binary discretized. The number of propensity score strata is defaulted to 5 layers, and the user can uniformly change the number of strata according to needs;

[0108] S4213: Propensity Score Matching: The backdoor adjustment variable set is not an empty set and the number does not exceed 5. The dependent variable is one-dimensional and follows a binary distribution or can be binary discretized.

[0109] S422: Causal Effect Estimation Based on the Frontdoor Criterion

[0110] The causal effect evaluation algorithm based on the frontdoor criterion supports the two-layer linear regression algorithm based on the frontdoor criterion.

[0111] S423: Causal Effect Estimation Based on Instrumental Variables

[0112] The causal effect evaluation algorithm based on instrumental variables supports the two-stage least squares method and the binary instrument / Wald estimation algorithm based on instrumental variables. The corresponding matching rules of the algorithm are designed as follows:

[0113] S4231: Binary instrument / Wald estimation: Both the instrumental variable set and the dependent variable are one-dimensional. Among them, if the instrumental variable follows a binary distribution, the Wald estimation expression is used to calculate the causal effect; if the instrumental variable does not follow a binary distribution, the variance expression is used to calculate the causal effect;

[0114] S4232 Two-stage least squares method: The dimension of the instrumental variable or the dependent variable is greater than one-dimensional, and the instrumental variable is used to solve the endogeneity bias caused by confounding variables.

[0115] S424: Causal effect evaluation result test

[0116] Since part of the causal graph drawing process depends on experts' subjective experience, in order to objectively verify the reliability of causal inference analysis of complex systems, the method of refutation is used to verify the credibility of the calculation results. If the test is passed, it means that the calculation result of the causal effect has a high credibility; otherwise, the calculation result of the causal effect is considered untrustworthy.

[0117] S4241: Data subset verification

[0118] The data subset verification is used to re-establish the evaluation data and re-calculate the causal effect estimation result. If the test result is not much different from the calculation result and the test confidence level ≥ 0.9, it can be considered that the test is passed.

[0119] S4242: Adding a random confounding factor

[0120] Add a randomly generated confounding variable. For a causal relationship X on Y, a random variable that is independent of both causal variables is added as a confounding factor to this causal relationship, and the causal effect of X on Y is recalculated. If the test result is not much different from the calculation result and the test confidence level ≥ 0.9, it can be considered that the test is passed.

[0121] S4243: Placebo intervention

[0122] Randomly delete a part of the data. The new data is a random subset of the original data. Use the subset to recalculate the causal effect of X on Y. If the test result is very different from the calculation result and the test confidence level ≥ 0.9, it can be considered that the test is passed.

[0123] S5: Counterfactual inference calculation

[0124] The counterfactual calculation algorithm is adopted to calculate the counterfactual results, which can solve the sub-problem of predicting the intervention effect in the causal inference process. This algorithm needs to set the input data (sample data of a certain experiment) and the set value of the dependent variable (the value of the dependent variable is set according to the user's needs). After the input data, the causal graph structure and the effect calculation results are given, the counterfactual calculation results are automatically given according to the processing flow, and the counterfactual calculation function in the effect estimation module of the prototype system is realized by calling the counterfactual calculation algorithm through the model library.

[0125] The best application scenario of counterfactual inference is after obtaining the experimental results. If the experimental results are not ideal, by keeping other variables unchanged, analyze how to adjust the parameter variables so that the result variable reaches the expected value. That is, under the premise of limited resources, accurately adjust the relevant parameter variables to simulate the intervention effect that can only be seen by re-performing the experiment.

[0126] S6: Causal traceability analysis

[0127] To further analyze the association between the causal effect estimation results and relevant decision-making elements, it is necessary to further analyze the causal effect calculation results. Layer-by-layer analysis is carried out on the single-experiment data in the scenario to find the key reasons affecting the target result variable, and the retrospective analysis results are presented in the form of a causal quantitative decision tree. Causal traceability analysis uses the regression coefficient as the basis for judging the importance of variables, and at the same time considers multiple dependent variables to analyze the influence degree of multiple dependent variables on the result variable. There are two recommended application scenarios for causal traceability analysis: under the condition of limited resources, causal traceability analysis can give suggestions for the initial setting of parameters; when the experimental results are not ideal, it can find the variables that cause the result deviation and give optimization suggestions for subsequent experiments. The schematic diagram of causal traceability analysis is as Figure 2 shown.

[0128] The above is only the preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, makes equivalent replacements or changes, and all should be covered within the protection scope of the present invention.

Claims

1. A causal inference method based on problem decomposition and algorithm matching, characterized in that: It includes the following steps: S1: Causal inference problem definition To determine the problems to be solved in causal inference, according to the common problems of causal inference, it is divided into four problems: causal relationship identification sub-problem, causal effect estimation sub-problem, intervention effect prediction sub-problem, and traceability analysis sub-problem. The four problems correspond to the four stages of causal inference, and through the in-depth study of causal relationships from shallow to deep, conclusions and suggestions based on causal inference are finally formed; S2: Generation of causal inference dataset After the user uploads the original data to the system, perform data preprocessing operations such as data cleaning, outlier removal, normalization, and data interpolation on the original data to obtain the dataset required for causal relationship inference; S3: Causal relationship discovery and verification For the different requirements of application scenarios, three causal diagram drawing methods, namely manual drawing, semi-automatic drawing, and full-automatic drawing, can be selected. For the three causal diagram drawing methods applicable to different scenarios, DAG completeness test and Do calculus are required to verify the accuracy of the causal diagram structure and causal relationship. By adopting a complete causal relationship verification method, ensure that all variables with causal relationships are accurately identified, screen out the variables with causal relationships in the dataset and draw the causal structure diagram; S4: Causal effect estimation and verification Due to the complex characteristics of the causal diagram structure, this step combines the structural causal model and the potential outcome framework, uses the relevant rules in the structural causal model to identify the causal structure, uses the potential outcome framework to integrate multiple causal effect estimation algorithms, and selects a suitable combination of causal effect estimation algorithms through parameter configuration and logical design, so as to complete the calculation of the causal effects of all causal variables in the causal diagram, and provides calculation models such as direct calculation method, propensity score stratification, propensity score matching-based, two-layer linear regression, binary instrument / Wald estimation, two-stage least squares estimation, etc. to handle causal effect calculations under different causal structures and parameter conditions. Finally, use the method of refutation to test the accuracy and reliability of the causal effect calculation results, including three methods: data subset verification, adding random confounding factors, and placebo intervention. One or several of these methods can be selected for testing to obtain the causal effect values between variables with causal relationships; S5: Counterfactual inference calculation Calculate the counterfactual calculation results under specific conditions, which can predict the effect of the intervention. Taking causal variables and specific samples as the research objects, while keeping the values of covariates affecting the causal variables unchanged, calculate the change in the value of the outcome variable when only the value of the independent variable is changed, and the impact brought by the parameter adjustment change can be predicted without re-performing the experiment; S6: Causal traceability analysis Further analyze the results of causal relationship identification and causal effect estimation. Based on the causal effect calculation results, discuss the rationality of the relevant parameter settings during the experimental operation process in the relevant scenario, and form a retrospective analysis conclusion based on the above analysis results.

2. The causal inference method based on problem decomposition and algorithm matching according to claim 1, wherein: The process of the above-mentioned S1: Causal inference problem definition is: clarify the specific problems to be solved and the set goals of this causal inference.

3. The causal inference method based on problem decomposition and algorithm matching according to claim 1, characterized in that: The process of the above-mentioned S2: Generation of causal inference dataset is: perform preprocessing on the original data to generate a causal inference evaluation dataset.

4. A causal inference method based on problem decomposition and algorithm matching according to claim 1, characterized in that: The process of the said S3: Causality discovery and verification step is as follows: Select a causal graph method to draw a causal graph, and verify the correctness of the causal graph structure and causal relationship.

5. A causal inference method based on problem decomposition and algorithm matching according to claim 1, characterized in that: The process of the said S4: Causal effect estimation and verification step is as follows: Automatically match a causal effect evaluation algorithm, calculate the causal effect value and the verification result.

6. The causal inference method based on problem decomposition and algorithm matching according to claim 1, characterized in that: The process of the said S5: Counterfactual inference calculation step is as follows: Calculate the counterfactual result according to the scenario requirements, and predict the impact brought by parameter adjustment changes.

7. A causal inference method based on problem decomposition and algorithm matching according to claim 1, characterized in that: The process of the said S6: Causal traceability analysis step is as follows: Trace back the key variables and form a conclusion of causal traceability analysis.