Causal discovery method and device and electronic equipment

By screening stable core dependent variables in the time dimension and performing quantitative regression, the accuracy and computational quantities of the causal discovery method in digital operations are solved, and the accurate quantification and efficient analysis of multi-objective causal relationships are achieved, supporting operational strategy optimization in complex scenarios.

CN120338079APending Publication Date: 2025-07-18青岛聚看云科技有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510239446.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing causal discovery methods have problems with low accuracy and high computational volume in digital operation scenarios, especially in multi-objective scenarios, which is difficult to accurately capture the correlation between factors and targets, resulting in inefficiency of operation strategies.

Method used

By counting the frequency of candidate dependent variables in the time dimension, the stable core dependent variables are selected, and the factor coefficient is determined using quantitative regression technology to achieve a multi-objective causal discovery process, reducing the number of regressions, and improving the accuracy and completeness of causal relationships.

Benefits of technology

It improves the accuracy and computational efficiency of causal discovery, is suitable for complex multi-objective scenarios, quantifies the contribution of factors to the goals, and supports the formulation and implementation of refined operation strategies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120338079A_ABST
    Figure CN120338079A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data mining, provides a causal discovery method and device and electronic equipment, and can be applied to a digital operation scene. According to the method, candidate dependent variable sets having a causal relationship with at least one target variable are respectively positioned in a plurality of time periods, so that interference dependent variables irrelevant to each target variable can be filtered out by counting the occurrence frequency of each candidate dependent variable in the time dimension; the multiple core dependent variables having the direct causal relationship with the target variables are obtained through the target variables of the business to be processed, so that the causal discovery accuracy is improved, meanwhile, the multi-target causal discovery process is achieved in the mode that at least one target variable of the business to be processed shares the same business index, and when the causal relationship is quantitatively analyzed through regression, the accuracy of causal discovery is improved. The regression frequency is reduced, and the calculation amount of quantitative analysis of the causal relationship is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0002] The formulation of digital operation strategies requires clarifying the association between business factors and strategic goals. This association relationship in complex scenarios not only requires qualitative but also quantitative expression, so as to estimate the contribution of factor changes to the goal before the implementation of the strategy.

[0003] Currently, the industry usually uses association analysis matrices to measure the correlation between factors. However, if the correlation conclusion is obtained based on a biased data source, there may be large errors, such as the Sum of Squares Between (SSB) error in the recommendation scenario, which affects the formulation and implementation effect of operation strategies; and if the data source is unbiased, there may still be spurious associations between factors, resulting in unreliable correlation conclusions and affecting the effect of operation strategies.

[0004] In recent years, causal inference techniques have been gradually widely applied in digital operation scenarios because they can solve the problems of data bias and spurious associations. Causal inference techniques are represented by the Structural Causal Model (SCM) and the Potential Outcome Framework (POF). Among them, SCM represents causal discovery and mines the causal topological relationship between variables in the form of a graph, while POF is based on observed and counterfactual results, focusing on estimating the impact of intervention means on the goal and not paying attention to the specific causal structure.

[0005] In actual digital operations, more often, the changes in observed indicators are used to sort out the key factors affecting the goal, and the role of intervention means has been reflected in the indicator values. Therefore, causal discovery methods are mainly used in digital operations to capture variable relationships. However, the current causal discovery methods have problems such as low accuracy in mining causal relationships and large computational amounts, resulting in inaccurate association relationships between factors and goals, and it is difficult to unbiasedly evaluate the impact of factors on goals. Especially in the case of complex associations between factors, it is more likely to cause inefficiencies in operation strategies. In addition, digital operations often consider multiple goals, and the current causal discovery algorithms are mainly for single goals, making it more difficult to accurately capture the association relationship between factors and goals through conventional data analysis methods. Summary of the Invention

[0006] Embodiments of the present application provide a causal discovery method, device, and electronic device for improving the accuracy of causal relationship analysis and reducing the computational amount of causal relationship analysis.

[0007] In a first aspect, embodiments of the present application provide a causal discovery method, including:

[0008] For each time period, the following operations are performed separately: according to the pre-set association relationships between the at least one target variable and the shared business metrics, locate a set of candidate dependent variables that have a causal relationship with the at least one target variable during the corresponding time period;

[0009] According to the frequencies of the candidate dependent variables in the T sets of candidate dependent variables, filter out multiple core dependent variables;

[0010] Regress the exposure estimates of the multiple core dependent variables on the business metrics respectively, and determine the factor coefficients between the multiple core dependent variables and the target variables with causal relationships respectively.

[0011] The beneficial effects of the above technical solution are as follows: during multiple time periods, locate the sets of candidate dependent variables that have causal relationships with at least one target variable respectively. In this way, by statistically counting the frequencies of each candidate dependent variable in the time dimension, the interfering dependent variables irrelevant to each target variable can be filtered out, and multiple core dependent variables that have direct causal relationships with each target variable can be obtained, thereby improving the accuracy of causal discovery. Moreover, by sharing the same business metrics for at least one target variable of the business to be processed, the causal discovery process for multiple targets is realized, so as to be applicable to more complex multi-target causal discovery scenarios and have a wider application range.

[0012] On the other hand, by regressing the exposure estimates of the core dependent variables on the business metrics, determine the factor coefficients of the core dependent variables corresponding to each of the at least one target variable, thereby quantifying the contribution of the core dependent variables to the target variables, improving the completeness of the causal relationship, and reducing the number of regressions and the computational amount of the quantitative analysis of the causal relationship compared with performing regressions separately according to the at least one target variable.

[0013] Optionally, the filtering out of multiple core dependent variables according to the frequencies of the candidate dependent variables in the multiple sets of candidate dependent variables includes:

[0014] Statistically count the frequencies of each candidate dependent variable in the multiple sets of candidate dependent variables;

[0015] According to each frequency and the total number of variables in the multiple sets of candidate dependent variables, determine the score of each candidate dependent variable respectively, and the score is positively correlated with the frequency;

[0016] In the order of the scores from high to low, filter out the top K candidate dependent variables not less than the preset score threshold from the multiple sets of candidate dependent variables, where K is an integer greater than 1;

[0017] Take the top K candidate dependent variables as multiple core dependent variables.

[0018] The beneficial effects of the above technical solution are as follows: By counting the frequencies of each candidate dependent variable among multiple candidate dependent variables, the aggregation operation of each candidate dependent variable in the time dimension is realized. The higher the frequency, the higher the score, indicating the higher the stability of the dependent variable. Therefore, the accuracy of the core dependent variables screened out in descending order of scores and not less than the preset score threshold is higher, thereby improving the accuracy of causal discovery.

[0019] Optionally, the calculation formula of the score is:

[0020]

[0021] where f i represents the i-th candidate dependent variable, c i represents the frequency of the i-th candidate dependent variable, |S| represents the total number of variables in the multiple candidate dependent variable sets, and r i represents the score of the i-th candidate dependent variable.

[0022] The beneficial effects of the above technical solution are as follows: By introducing an exponential function to calculate the score, the discrimination degree of multiple core dependent variables is increased.

[0023] Optionally, the step of regressing the exposure values of the multiple core dependent variables on the business indicators respectively to determine the factor coefficients between the multiple core dependent variables and the target variables with causal relationships includes:

[0024] For each core dependent variable, perform the following operations respectively: Obtain L historical data of the core dependent variable and L historical data of the business indicator, perform regression on the L historical data of the core dependent variable and the L historical data of the business indicator to determine L initial exposure values of the core dependent variable on the business indicator, and take the mean of the L initial exposure values as the target exposure value of the core dependent variable on the business indicator;

[0025] Perform a single regression on the target exposure values of the multiple core dependent variables in cross-section to obtain the factor coefficients between the multiple core dependent variables and the target variables with causal relationships.

[0026] The beneficial effects of the above technical solution are as follows: On the basis of qualitative analysis of causal relationships, through the cross-section regression method in econometrics, the factor coefficients between the core dependent variables and the target variables with causal relationships are quantified, realizing the quantitative analysis of causal relationships, ensuring the completeness of the core dependent variables, and since the factor coefficients can reflect the contribution of the core dependent variables to the corresponding target variables, the influence of the core dependent variables can be better controlled.

[0027] Optionally, the step of performing a single regression on the target exposure values of the multiple core dependent variables in a cross-section to obtain the factor coefficients between each of the multiple core dependent variables and the target variables having a causal relationship includes:

[0028] Detecting the value of the attribution error in the cross-section regression equation;

[0029] If the value of the attribution error is 0, then according to the association relationship between the at least one target variable and the business indicator, in the cross-section regression equation, the factor coefficients between each of the multiple core dependent variables and the business indicator are used as the factor coefficients between each of the multiple core dependent variables and the target variables having a causal relationship.

[0030] The beneficial effects of the above technical solution are as follows: Since at least one target variable shares a business indicator, therefore, in the cross-section regression process, by quantifying the factor coefficients between each of the multiple core dependent variables and the business indicator, and combining the association relationship between the at least one target variable and the business indicator, the causal qualitative analysis between the core dependent variable and the corresponding target variable is realized. Compared with performing cross-section regression using at least one target variable separately, the regression calculation problem caused by multi-dimensional explosion analysis is solved, and the calculation amount in the regression process is reduced.

[0031] Optionally, the step of performing regression on the exposure values of the multiple core dependent variables on the business indicator respectively to determine the factor coefficients between each of the multiple core dependent variables and the target variables having a causal relationship includes:

[0032] For each core dependent variable, respectively perform: obtaining T historical data of the core dependent variable and T historical data of the business indicator, and performing regression on the T historical data of the core dependent variable and the T historical data of the business indicator to determine the T initial exposure values of the core dependent variable on the business indicator;

[0033] Performing T regressions on the T initial exposure values of each of the multiple core dependent variables in a cross-section to obtain T initial coefficients between each core dependent variable and the target variable having a causal relationship;

[0034] The mean value of the T initial coefficients corresponding to each of the multiple core dependent variables is used as the factor coefficient between the corresponding core dependent variable and the target variable having a causal relationship.

[0035] The beneficial effects of the above technical solution are as follows: Based on the qualitative analysis of causal relationships, the factor coefficients between the core dependent variable and the target variables with causal relationships are quantified through the Fama-MacBeth regression method in econometrics, realizing the quantitative analysis of causal relationships, ensuring the completeness of the core dependent variable. Compared with the traditional single cross-sectional regression method, the Fama-MacBeth regression excludes the influence of the cross-sectional correlation of residuals on the standard deviation of factor coefficients, improving the accuracy of the quantitative analysis of causal relationships.

[0036] In a second aspect, an embodiment of the present application provides an electronic device, including a processor, a memory, and a communication interface, where the communication interface, the memory, and the processor are connected through a bus;

[0037] The communication interface is used to communicate with other devices;

[0038] The memory stores a computer program, and the processor performs the following operations according to the computer program:

[0039] In T time periods, at least one target variable of the service to be processed is obtained respectively, where T is an integer greater than 1;

[0040] For each time period, the following is performed respectively: According to the preset association relationships between the at least one target variable and the shared service metrics, a set of candidate independent variables that have causal relationships with the at least one target variable in the corresponding time period is located;

[0041] According to the frequencies of the candidate independent variables in the T sets of candidate independent variables, multiple core independent variables are screened out;

[0042] The exposure estimates of the multiple core independent variables on the service metrics are regressed respectively to determine the factor coefficients between the multiple core independent variables and the target variables with causal relationships.

[0043] Optionally, when the processor screens out multiple core independent variables according to the frequencies of the candidate independent variables in the multiple sets of candidate independent variables, the specific operation is as follows:

[0044] Count the frequencies of each candidate independent variable in the multiple sets of candidate independent variables;

[0045] According to each frequency and the total number of variables in the multiple sets of candidate independent variables, the scores of each candidate independent variable are determined respectively, and the scores are positively correlated with the frequencies;

[0046] The top K candidate independent variables not less than a preset score threshold are screened out from the multiple sets of candidate independent variables in descending order of the scores, where K is an integer greater than 1;

[0047] Take the top K candidate dependent variables as multiple core dependent variables.

[0048] Optionally, the calculation formula for the score is:

[0049]

[0050] where f i represents the i-th candidate dependent variable, c i represents the frequency of the i-th candidate dependent variable, |S| represents the total number of variables in the set of multiple candidate dependent variables, and r i represents the score of the i-th candidate dependent variable.

[0051] Optionally, the processor regresses the exposure estimates of the multiple core dependent variables on the business metric respectively to determine the factor coefficients between each of the multiple core dependent variables and the target variable with which there is a causal relationship. The specific operation is as follows:

[0052] For each core dependent variable, perform the following respectively: Obtain T historical data of the core dependent variable and T historical data of the business metric, perform regression on the T historical data of the core dependent variable and the T historical data of the business metric to determine T initial exposure values of the core dependent variable on the business metric, and take the mean of the T initial exposure values as the target exposure value of the core dependent variable on the business metric;

[0053] Perform a single regression on the target exposure values of the multiple core dependent variables across sections to obtain the factor coefficients between each of the multiple core dependent variables and the target variable with which there is a causal relationship.

[0054] Optionally, the processor performs a single regression on the target exposure values of the multiple core dependent variables across sections to obtain the factor coefficients between each of the multiple core dependent variables and the target variable with which there is a causal relationship. The specific operation is as follows:

[0055] Detect the value of the attribution error in the cross-sectional regression equation;

[0056] If the value of the attribution error is 0, then according to the association relationship between the at least one target variable and the business metric, take the factor coefficients between each of the multiple core dependent variables and the business metric in the cross-sectional regression equation as the factor coefficients between each of the multiple core dependent variables and the target variable with which there is a causal relationship.

[0057] Optionally, the processor regresses the exposure values of the multiple core dependent variables on the business metric respectively to determine the factor coefficients between each of the multiple core dependent variables and the target variable with which there is a causal relationship. The specific operation is as follows:

[0058] For each core dependent variable, the following operations are respectively performed: obtain T historical data of the core dependent variable and T historical data of the business metric, perform regression on the T historical data of the core dependent variable and the T historical data of the business metric, and determine T initial exposure values of the core dependent variable on the business metric;

[0059] Perform T regressions on the T initial exposure values of each of the multiple core dependent variables cross-sectionally to obtain T initial coefficients between each core dependent variable and the target variable having a causal relationship;

[0060] Take the mean of the T initial coefficients corresponding to each of the multiple core dependent variables as the factor coefficient between the corresponding core dependent variable and the target variable having a causal relationship.

[0061] In a third aspect, an embodiment of the present application provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of any of the above-mentioned causal discovery methods are implemented.

[0062] The technical effects brought by any implementation manner in the second aspect to the third aspect can be referred to the technical effects brought by the corresponding implementation manner in the first aspect, which will not be elaborated here. Description of the Drawings

[0063] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0064] Figure 1 It is a flowchart of a causal discovery method provided by an embodiment of the present application;

[0065] Figure 2A It is a relationship diagram of multiple target variables and a business metric provided by an embodiment of the present application;

[0066] Figure 2B It is an example diagram of the association between multiple target variables and a business metric provided by an embodiment of the present application;

[0067] Figure 3 It is a causal relationship topology diagram provided by an embodiment of the present application;

[0068] Figure 4 It is a flowchart of a method for screening dependent variables provided by an embodiment of the present application;

[0069] Figure 5Flowchart of a causal relationship regression method provided by an embodiment of the present application;

[0070] Figure 6 Flowchart of a method for determining factor coefficients provided by an embodiment of the present application;

[0071] Figure 7A Flowchart of a causal relationship regression method provided by an embodiment of the present application;

[0072] Figure 7B Schematic diagram of the verification result of the effectiveness of a causal relationship provided by an embodiment of the present application;

[0073] Figure 8 Structural diagram of a causal discovery device provided by an embodiment of the present application;

[0074] Figure 9 Structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0075] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the technical solutions of the present application. Based on the embodiments described in this application document, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the technical solutions of the present application.

[0076] Based on the exemplary embodiments shown in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present application. In addition, although the disclosure in the present application is introduced according to one or several exemplary instances, it should be understood that each aspect of these disclosures can also constitute a complete technical solution alone.

[0077] In addition, the terms "including" and "having" and any variations thereof are intended to cover but not exclude inclusion. For example, a product or device including a series of components does not necessarily have to be limited to those components clearly listed, but may include other components not clearly listed or inherent to these products or devices.

[0078] The term "module" used in the present application refers to any known or later-developed hardware, software, firmware, artificial intelligence, fuzzy logic, or a combination of hardware or / and software code that can perform functions related to that element.

[0079] The design concept of the embodiments of the present application is outlined below in combination with the application scenario.

[0080] With the development of Artificial Intelligence (AI) technology, the development of some large-screen service platforms related to video depends on the continuous enhancement of user insight capabilities. Taking Netflix as an example, the company has successfully transformed into the field of self-produced TV series, and the root cause of the success of the operation strategy of self-produced TV series lies in the successful application of video recommendation capabilities and big data mining technology accumulated over the years.

[0081] Therefore, in the digital operation scenario, positioning the core factors that affect business goals and quantifying the contribution of core factors to business goals will greatly enhance user insight capabilities and operation efficiency.

[0082] Currently, the causal discovery methods for establishing the correlation between business goals and core factors are mainly divided into the following three categories:

[0083] First, constraint-based causal discovery algorithms. Such algorithms mainly use conditional independence in data to obtain causal graphs. Representative algorithms include the Peter-Clark algorithm (PC algorithm), the Inferred Causation (IC) algorithm, etc.

[0084] Second, score-based causal discovery algorithms. Such algorithms mainly achieve the maximum convergence state by iterating the score S formed by the data set and the graph structure, so as to obtain the optimal graph structure of causal relationships. Representative algorithms include the Greedy Equivalence Search (GES) algorithm based on Bayesian networks, the NOTEARS algorithm, the LEAST algorithm, etc.

[0085] Third, causal discovery algorithms based on Functional Causal Models (FCMs). Its main principle is that a variable can be written as a function of its set of direct causes and a noise term. Such algorithms need to sort the variables in causal order in advance, and then obtain the form of the function through learning to obtain causal relationships. Representative algorithms include ICA-LiNGAM, ANMs, etc.

[0086] However, the first two types of causal discovery algorithms can generate causal topology graphs, but the weights of the edges in the graphs only characterize a certain degree of conditional correlation and cannot be used as a measure of the influence between variables; the third type of algorithm requires a large amount of computation and a large number of variables with known causal relationships. At the same time, it is also assumed that there are no interference factors in the model during the calculation, and the implementation environment is relatively ideal.

[0087] In view of this, the embodiments of the present application provide a causal discovery algorithm that combines at least one target causal discovery technique and a measurement regression technique. Among the at least one target causal discovery technique, at least one target variable shares the same business metric, and the pre-set association relationships between the at least one target variable and the business metric are respectively set. In this way, based on the association relationships between the at least one target variable and the business metric respectively, a candidate set of dependent variables that have a causal relationship with the at least one target variable can be mined at one time, realizing a multi-objective causal discovery process, thereby being applicable to more complex multi-objective causal discovery scenarios, having a wider application scope. Moreover, by screening the multiple candidate sets of dependent variables that have a causal relationship with the at least one target variable mined in T time periods, irrelevant interfering dependent variables can be filtered out, improving the accuracy of causal discovery; in the measurement regression technique, by regressing the exposure estimate value of the core dependent variable on the business metric, the factor coefficients of the core dependent variable corresponding to each of the at least one target variable are determined, thereby quantifying the contribution of the core dependent variable to the target variable, ensuring the completeness of the causal relationship, and reducing the number of regressions and the computational amount of quantitative analysis of the causal relationship compared with performing regression separately according to the at least one target variable.

[0088] The causal discovery algorithm of the embodiments of the present application can solve the problem of factor contribution measurement in operation target analysis in the digital operation scenario, making up for the shortcoming of the traditional causal discovery algorithm in causal relationship quantification and the regression calculation problem brought by the multi-factor dimension explosion, and is suitable for the formulation and implementation of refined digital operation strategies with complex scenarios and high interpretability requirements.

[0089] It should be noted that the digital operation scenario is only an example of the application scenario of the causal discovery method and is not a restrictive requirement. For example, the causal discovery method can be applied to fault analysis, or the causal discovery method can be cited for the analysis of environmental pollution sources, or the causal discovery method can also be applied to assist in formulating medical plans, formulating sales strategies, etc.

[0090] The following describes the preferred embodiments of the present application with reference to the accompanying drawings of the specification. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present application and are not used to limit the present application. And without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other.

[0091] See Figure 1 , which provides a process of a causal discovery method for the embodiments of the present application. This process mainly includes the following steps:

[0092] S101: Obtain at least one target variable of the business to be processed respectively in T time periods.

[0093] In the business to be processed, there may be one or more analysis objectives, and each objective may change dynamically, with different contents at different times. Therefore, for the business to be processed, at least one target variable in T time periods can be obtained, where T is an integer greater than 1.

[0094] Taking the video recommendation business as an example, there are two target variables: click (label) and payment (pay_label). On the first day, the target variable "click" for video A is 5000, and the target variable "payment" is 1 yuan. On the second day, the target variable "click" for video A is 20,000, and the target variable "payment" is 2 yuan.

[0095] S102: For each time period, perform the following respectively: According to the association relationships between at least one preset target variable and shared business metrics, locate a candidate set of dependent variables that has a causal relationship with at least one target variable in the corresponding time period.

[0096] In some embodiments, when the number of target variables is multiple, to reduce the complexity and computational amount of multi-objective causal discovery, multiple target variables can be associated through one business metric (also called a dummy metric), that is, multiple target variables share one business metric, and the association relationships between each preset target variable and the business metric are as follows Figure 2A shown.

[0097] Generally, the selection of business metrics can be determined based on business experience.

[0098] Still taking the video recommendation business as an example, the business metric (dummy metric) shared by the two target variables "click" and "payment" is "revenue". The association relationship between the target variable "click" and the business metric "revenue" is set to 100% influence, which is represented by the value "1.0" for the relationship between the two. Also, the association relationship between the target variable "payment" and the business metric "revenue" is set to 100% influence, which is also represented by the value "1.0" for the relationship between the two, as Figure 2B shown.

[0099] After establishing the association relationships between each target variable and the business metric, for the at least one target variable obtained for each time period, with the business metric shared by each target variable as the ultimate goal, combined with the association relationships between at least one target variable and the business metric, a candidate set of dependent variables that has a causal relationship with at least one target variable simultaneously in the corresponding time period is mined.

[0100] In some embodiments, the process of locating causal relationships may adopt a causal discovery algorithm based on Hierarchical Target-Orientated (HTO). The business metrics and the correlation relationships between the business metrics and each target variable are used as the input of the algorithm. For a specific business metric, based on the conditional independence property of the basic structure of the Bayesian network, the algorithm takes a falsification approach, starting from the business metric, and progressively obtains the dependent variables with causal relationships level by level, thereby establishing a causal relationship topology diagram between at least one target variable and the dependent variables.

[0101] Taking the video recommendation service as an example, based on the correlation relationships between the business metric "revenue" and the target variables "click" and "pay", the mined causal relationship topology diagram is as Figure 3 shown. In Figure 3 , the causal discovery locates that there are 2 candidate dependent variables, is_single and vod_pit_7d, that have a causal relationship with the target variable "pay", and 1 candidate dependent variable, is_vip, that has a causal relationship with the target variable "click".

[0102] In some embodiments, within a time period, a candidate dependent variable may appear once, may appear multiple times, or may not appear at all.

[0103] S103: Screen out multiple core dependent variables according to the frequencies of each candidate dependent variable in the T sets of candidate dependent variables.

[0104] To ensure the stability of the causal relationship in the time dimension, stability verification can be performed based on the frequencies of each candidate dependent variable discovered by the causal discovery to filter out irrelevant interfering dependent variables and retain accurate core dependent variables.

[0105] In some embodiments, the stability verification process of the dependent variables is as Figure 4 shown and mainly includes the following steps:

[0106] S1031: Count the frequencies of each candidate dependent variable in the T sets of candidate dependent variables.

[0107] Suppose the set of candidate dependent variables in a time period is represented as F = {f1, f2,..., f n}, where each f represents a candidate dependent variable. Then the T sets of candidate dependent variables in T time periods are represented as F τ = {f 1,τ , f 2, τ,..., f n,τ}, τ ∈ {1, 2,..., T}. Among them, the higher the frequency of each candidate dependent variable appearing in the T time periods, the stronger its stability. Therefore, after obtaining the T sets of candidate dependent variables F τAfter that, a convergence operation is performed in the time dimension to count the frequency of occurrence of each candidate dependent variable in the T candidate dependent variable sets in total.

[0108] Specifically, the frequency of each candidate dependent variable can be expressed as: S = {(f1, c1), (f2, c2), …, (f n , c n i)}, where each c represents the frequency of occurrence of the corresponding candidate dependent variable in the T candidate dependent variable sets.

[0109] S1032: Determine the score of each candidate dependent variable respectively according to each frequency and the total number of variables in the multiple candidate dependent variable sets.

[0110] Among them, the score is positively correlated with the frequency.

[0111] In some embodiments, the score can be calculated using an exponential function, specifically expressed as:

[0112]

[0113] where f i represents the i-th candidate dependent variable, c i represents the frequency of the i-th candidate dependent variable, |S| represents the total number of variables in the multiple candidate dependent variable sets, and r i represents the score of the i-th candidate dependent variable.

[0114] By introducing an exponential function to calculate the score, the discrimination of multiple core dependent variables is increased.

[0115] It should be noted that the embodiments of the present application do not impose restrictive requirements on the calculation method of the score. For example, the ratio of the frequency to the total number of variables can also be directly used as the score of the corresponding candidate dependent variable.

[0116] S1033: Screen out the top K candidate dependent variables not less than the preset score threshold from the multiple candidate dependent variable sets in the order of scores from high to low.

[0117] Among them, the higher the score, the stronger the stability of the candidate dependent variable.

[0118] S1034: Use the top K candidate dependent variables as multiple core dependent variables.

[0119] Among them, the formula for the multiple core dependent variables is expressed as: is an integer greater than 1, and m represents the preset score threshold.

[0120] By counting the frequencies of each candidate dependent variable among multiple candidate dependent variables, the aggregation operation of each candidate dependent variable in the time dimension is realized. The higher the frequency, the higher the score, indicating that the stability of the dependent variable is higher. Therefore, the higher the accuracy of the core dependent variables selected in descending order of scores and not less than the preset score threshold, thus improving the accuracy of causal discovery.

[0121] S104: Regress the exposure estimates of multiple core dependent variables on business metrics respectively to determine the factor coefficients between each core dependent variable and the target variable with which there is a causal relationship.

[0122] After screening out multiple core dependent variables with which there is a causal relationship with at least one target variable, it is also necessary to quantitatively analyze the causal relationship between the core dependent variable and the corresponding target variable.

[0123] In some embodiments, the cross-sectional regression method in econometrics can be used to quantitatively analyze the causal relationship, so as to obtain the contribution degree of each core dependent variable to the target variable with which there is a causal relationship.

[0124] See Figure 5 , which is a method flow for quantitative analysis of causal relationship provided by the embodiments of the present application, mainly including the following steps:

[0125] S501: For each core dependent variable, respectively execute: obtain T historical data of the core dependent variable and T historical data of the business metric, and regress the T historical data of the core dependent variable and the T historical data of the business metric to determine the T initial exposure values of the core dependent variable on the business metric.

[0126] For each core dependent variable, at the past T time points, perform time series regression on the T historical data of the core dependent variable and the T historical data of the business metric to obtain the exposure value of the core dependent variable at the corresponding time point.

[0127] Taking the i-th core dependent variable among the K core dependent variables as an example, the formula of the regression process is expressed as:

[0128]

[0129] where h i represents the business metric, represents the initial exposure value of the i-th core dependent variable f i , α i represents the intercept, and ε i represents the error.

[0130] Among them, time series regression includes but is not limited to linear regression and non-linear regression.

[0131] S502: For each core dependent variable, perform the following separately: Use the mean of the T initial exposure values of the core dependent variable as the target exposure value of the core dependent variable in terms of business metrics.

[0132] For each core dependent variable, use the mean of the initial exposure values of the core dependent variable at T time points as the target exposure value of the core dependent variable in terms of business metrics.

[0133] S503: Conduct a single regression on the target exposure values of multiple core dependent variables across sections to obtain the factor coefficients between each core dependent variable and the target variable with which there is a causal relationship.

[0134] Through cross-sectional regression, find the relationship between the business metrics and the target exposure values of the core dependent variables across sections, so as to obtain the factor coefficients between the core dependent variables and the target variables with which there is a causal relationship.

[0135] In specific implementation, the cross-sectional regression process is as Figure 6 shown, mainly including the following steps:

[0136] S5031: Detect the value of the attribution error in the cross-sectional regression equation.

[0137] Specifically, the formula of the cross-sectional regression equation is expressed as:

[0138]

[0139] where η t represents the attribution error of cross-section t, represents the target exposure value of the core dependent variable, and λ t represents the factor coefficient of cross-section t, representing the quantitative value of attribution.

[0140] S5032: Determine whether the value of the attribution error is 0. If so, execute S5033; otherwise, execute S5034.

[0141] S5033: According to the correlation relationship between at least one target variable and the business metrics, use the factor coefficients between each of the multiple core dependent variables and the business metrics in the cross-sectional regression equation as the factor coefficients between each of the multiple core dependent variables and the target variable with which there is a causal relationship.

[0142] Assume that the correlation relationship between the target variable and the business metrics is 100% influence. In the quantitative analysis of causal relationship, each target variable can be regarded as the business metrics. Therefore, the factor coefficient between each core dependent variable and the business metrics can be used as the factor coefficient between the core dependent variable and the target variable with which there is a causal relationship.

[0143] S5034: Mine more core dependent variables.

[0144] When the attributable error is not 0, it is necessary to return to the causal discovery process of at least one target variable to add more core dependent variables, that is, to increase the number of time periods and start over. Figure 1 The entire process shown.

[0145] On the basis of qualitative causal analysis, through the cross-sectional regression method in econometrics, the factor coefficients between the core dependent variables and the target variables with causal relationships are quantified, realizing the quantitative analysis of causal relationships, ensuring the completeness of the core dependent variables, and since the factor coefficients can reflect the contribution of the core dependent variables to the corresponding target variables, the influence of the core dependent variables can be better controlled.

[0146] In some embodiments, the Fama-MacBeth regression method in econometrics can be used to quantitatively analyze causal relationships, so as to obtain the contribution degree of each core dependent variable to the target variable with causal relationships.

[0147] See Figure 7A , which is another method flow for quantitative causal analysis provided by the embodiments of the present application, mainly including the following steps:

[0148] S701: For each core dependent variable, perform respectively: obtain T historical data of the core dependent variable and T historical data of the business indicator, and perform regression on the T historical data of the core dependent variable and the T historical data of the business indicator to determine T initial exposure values of the core dependent variable on the business indicator.

[0149] Among them, the relevant description of S701 refers to the cross-sectional regression process and will not be elaborated here.

[0150] S702: Perform T regressions on the T initial exposure values of multiple core dependent variables on the cross-section to obtain T initial coefficients between each core dependent variable and the target variable with causal relationships.

[0151] Among them, the process of each cross-sectional regression refers to the relevant description of formula 3 above.

[0152] S703: The mean of the T initial coefficients corresponding to each core dependent variable is used as the factor coefficient between the corresponding core dependent variable and the target variable with causal relationships.

[0153] For each core dependent variable, the mean of the T initial coefficients of the T cross-sectional regressions is used as the factor coefficient between the core dependent variable and the target variable with causal relationships.

[0154] Specifically, the quantified value of the final attribution is the mean calculated from the initial coefficients of the T cross-sectional regressions:

[0155]

[0156] Among them, η represents the attribution error, and λ is the quantization value of attribution.

[0157] Based on the qualitative analysis of causality, through the Fama-MacBeth regression method in econometrics, the factor coefficient between the core dependent variable and the target variable with causality is quantified to realize the quantitative analysis of causality, ensuring the completeness of the core dependent variable. Compared with the traditional one-time cross-sectional regression method, the Fama-MacBeth regression excludes the influence of the cross-sectional correlation of residuals on the standard deviation of the factor coefficient, improving the accuracy of the quantitative analysis of causality.

[0158] In the embodiments of the present application, in multiple time periods, candidate dependent variable sets that have a causal relationship with at least one target variable are respectively located. In this way, by counting the frequencies of the appearance of each candidate dependent variable in the time dimension, the interference dependent variables irrelevant to each target variable can be filtered out, and multiple core dependent variables that have a direct causal relationship with each target variable can be obtained, thereby improving the accuracy of causal discovery. Moreover, by sharing the same business indicator for at least one target variable of the service to be processed, the causal discovery process for multiple targets is realized, so as to be applicable to more complex multi-target causal discovery scenarios and have a wider application range.

[0159] On the other hand, by regressing the exposure estimated value of the core dependent variable on the business indicator, the factor coefficients of the core dependent variables corresponding to at least one target variable are determined, thereby quantifying the contribution of the core dependent variable to the target variable, improving the completeness of the causal relationship, and reducing the number of regressions and the computational amount of the quantitative analysis of the causal relationship compared with performing regressions separately according to at least one target variable.

[0160] In some embodiments, the regression method is used to quantitatively analyze the causal relationship, and the effectiveness of the causal relationship can be evaluated by statistically verifying the intercept. During the regression process, when there is a significant intercept, it proves that there are still certain omissions in the dependent variable mining of at least one current target variable, and more potential dependent variables need to be mined for causal relationship analysis.

[0161] Taking the video recommendation service as an example, the two core dependent variables of the target variable "payment" are "is_single" and "vod_pit_7d", and the results of the two regression methods are as Figure 7B shown, through Figure 7BAccording to the data, the factor coefficients and significance conclusions obtained by the two regression methods are basically the same, indicating that the causal discovery method of the embodiments of the present application can effectively express causal relationships. In the digital operation scenario, the causal discovery algorithm provided by the embodiments of the present application constructs a complete and robust quantitative attribution intelligent operation system by combining multi-objective causal discovery and econometric regression, which can solve the problem of factor contribution measurement in operation target analysis, make up for the shortcoming of relationship quantification in the causal discovery algorithm and the regression calculation problem caused by the explosion of multi-factor dimensions, can more accurately and effectively improve operation measures, enhance the effect of business objectives, and is suitable for the formulation and implementation of refined operation strategies with complex scenarios and high requirements for interpretability.

[0162] Based on the same technical concept, the embodiments of the present application provide a causal discovery device, which can implement the steps of the above causal discovery method and achieve the same technical effect.

[0163] See Figure 8 , the causal discovery device includes an acquisition module 801, a relationship qualitative module 802, a factor screening module 803 and a relationship quantitative module 804, where:

[0164] The acquisition module 801 is used to respectively acquire at least one target variable of the business to be processed in T time periods, where T is an integer greater than 1;

[0165] The relationship qualitative module 802 is used to respectively execute for each time period: according to the pre-set association relationships between the at least one target variable and the shared business metrics, locate a candidate set of cause variables that have a causal relationship with the at least one target variable in the corresponding time period;

[0166] The factor screening module 803 is used to screen out multiple core cause variables according to the frequencies of the candidate cause variables in the T candidate sets of cause variables;

[0167] The relationship quantitative module 804 is used to perform regression on the exposure estimates of the multiple core cause variables on the business metrics respectively, and determine the factor coefficients between the multiple core cause variables and the target variables with causal relationships respectively.

[0168] Optionally, the factor screening module 803 is specifically used for:

[0169] Count the frequencies of each candidate cause variable in the multiple candidate sets of cause variables;

[0170] According to each frequency and the total number of variables in the multiple candidate sets of cause variables, determine the score of each candidate cause variable respectively, and the score is positively correlated with the frequency;

[0171] Filter out the top K candidate dependent variables that are not less than a preset score threshold from the multiple candidate dependent variable sets in descending order of the scores, where K is an integer greater than 1;

[0172] Use the top K candidate dependent variables as multiple core dependent variables.

[0173] Optionally, the calculation formula for the score is:

[0174]

[0175] where f i represents the i-th candidate dependent variable, c i represents the frequency of the i-th candidate dependent variable, |S| represents the total number of variables in the multiple candidate dependent variable sets, and r i represents the score of the i-th candidate dependent variable.

[0176] Optionally, the relationship quantification module 804 is specifically configured to:

[0177] For each core dependent variable, respectively perform: obtain T historical data of the core dependent variable and T historical data of the business metric, perform regression on the T historical data of the core dependent variable and the T historical data of the business metric, determine T initial exposure values of the core dependent variable on the business metric, and use the mean of the T initial exposure values as the target exposure value of the core dependent variable on the business metric;

[0178] Perform a regression on the target exposure values of the multiple core dependent variables in cross-section to obtain the factor coefficients between the multiple core dependent variables and the target variables with causal relationships respectively.

[0179] Optionally, the relationship quantification module 804 is specifically configured to:

[0180] Detect the value of the attribution error in the cross-section regression equation;

[0181] If the value of the attribution error is 0, then according to the association relationship between the at least one target variable and the business metric, use the factor coefficients between the multiple core dependent variables and the business metric in the cross-section regression equation as the factor coefficients between the multiple core dependent variables and the target variables with causal relationships respectively.

[0182] Optionally, the relationship quantification module 804 is specifically configured to:

[0183] For each core dependent variable, the following operations are performed separately: obtain T historical data of the core dependent variable and T historical data of the business metric, perform regression on the T historical data of the core dependent variable and the T historical data of the business metric, and determine T initial exposure values of the core dependent variable on the business metric;

[0184] Perform T regressions on the T initial exposure values of each of the multiple core dependent variables in cross-section to obtain T initial coefficients between each core dependent variable and the target variable with which there is a causal relationship;

[0185] The mean of the T initial coefficients corresponding to each of the multiple core dependent variables is used as the factor coefficient between the corresponding core dependent variable and the target variable with which there is a causal relationship. For the convenience of description, the above parts are divided into various modules (or units) according to their functions and described separately. Of course, when implementing the present application, the functions of the various modules (or units) can be implemented in the same or multiple software or hardware.

[0186] After introducing the causal discovery method and causal discovery device of the exemplary embodiment of the present application, next, an electronic device according to another exemplary embodiment of the present application is introduced.

[0187] In one embodiment, the electronic device can be a server or a terminal device. As Figure 9 shown, the structure of the electronic device includes a processor 901, a memory 902, and a communication interface 903, and the communication interface 903, the memory 902, and the processor 901 are connected through a bus 904;

[0188] The communication interface 904 is used to communicate with other devices;

[0189] The memory 902 stores a computer program, and the processor 901 executes the steps of any causal discovery method according to the computer program.

[0190] In the embodiments of the present application, the memory 902 may mainly include a program storage area and a data storage area. Among them, the program storage area may store an operating system and programs required to run the instant messaging function, etc.; the data storage area may store various instant messaging information and operation instruction sets, etc. The memory 902 may be a volatile memory, such as a random-access memory (RAM); the memory 902 may also be a non-volatile memory, such as a read-only memory, a flash memory, a hard disk drive (HDD), or a solid-state drive (SSD); or the memory 902 is any other medium that can be used to carry or store a desired computer program in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 902 may be a combination of the above memories.

[0191] The processor 901 may include one or more central processing units (CPUs), GPUs, or be a digital processing unit, etc. The processor 901 is used to implement the steps of any of the above causal discovery methods when calling the computer program stored in the memory 902.

[0192] It should be noted that Figure 9 is only an example, giving the necessary hardware for the electronic device to execute the steps of the causal discovery method provided in the embodiments of the present application. Those not shown, the electronic device may also include conventional hardware such as a display screen, a power supply, and operation buttons.

[0193] In the embodiments of the present application, the specific connection medium between the communication interface 903, the memory 902, and the processor 901 is not limited. In the embodiments of the present application, the bus 904 connecting the communication interface 903, the memory 02, and the processor 901 is Figure 9 described in thick lines in Figure 9 The connection manners between other components are only for illustrative purposes and are not to be construed as limiting. The bus 904 may be divided into an address bus, a data bus, a control bus, etc. For ease of description,

[0194] Those skilled in the art to which the present application pertains can understand that various aspects of the present application can be implemented as a system, a method, or a program product. Therefore, various aspects of the present application can be specifically implemented in the following forms, namely: a complete hardware implementation manner, a complete software implementation manner (including firmware, microcode, etc.), or an implementation manner combining hardware and software aspects, which can be collectively referred to as "circuits", "modules", or "systems" here.

[0195] The embodiments of the present application also provide a computer-readable storage medium for storing some instructions, which, when executed, can complete the steps of any of the causal discovery methods in the foregoing embodiments.

[0196] The embodiments of the present application also provide a computer program product for storing a computer program, and the computer program is used to execute the steps of any of the causal discovery methods in the foregoing embodiments.

[0197] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0198] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be realized by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for realizing the functions specified in one Figure 1 flow or multiple flows and / or blocks Figure 1 block or multiple blocks.

[0199] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device realizes the functions specified in one Figure 1 flow or multiple flows and / or blocks Figure 1 block or multiple blocks.

[0200] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for realizing the functions specified in one Figure 1 flow or multiple flows and / or blocks Figure 1 block or multiple blocks.

[0201] Obviously, those skilled in the art can make various changes and modifications to this application without departing from the spirit and scope of this application. Thus, if these modifications and variations of this application fall within the scope of the claims of this application and their equivalent technologies, this application is also intended to include these modifications and variations.

Claims

1. A causal discovery method, characterized in that, The method includes: Obtaining at least one target variable of the business to be processed respectively in T time periods, where T is an integer greater than 1; For each time period, respectively perform: according to the pre-set correlation relationships between the at least one target variable and the shared business metrics, locate a set of candidate dependent variables that have a causal relationship with the at least one target variable in the corresponding time period; Screen out multiple core dependent variables according to the frequencies of the candidate dependent variables in the T sets of candidate dependent variables; Regress the exposure estimates of the multiple core dependent variables on the business metrics respectively to determine the factor coefficients between the multiple core dependent variables and the target variables with causal relationships.

2. The method according to claim 1, characterized in that, The screening out multiple core dependent variables according to the frequencies of the candidate dependent variables in the multiple sets of candidate dependent variables includes: Counting the frequencies of each candidate dependent variable in the multiple sets of candidate dependent variables; Determining the score of each candidate dependent variable respectively according to each frequency and the total number of variables in the multiple sets of candidate dependent variables, and the score is positively correlated with the frequency; Screening out the top K candidate dependent variables not less than a preset score threshold from the multiple sets of candidate dependent variables in the order from high score to low score, where K is an integer greater than 1; Taking the top K candidate dependent variables as multiple core dependent variables.

3. The method according to claim 2, wherein The calculation formula of the score is: Among them, f i represents the i-th candidate dependent variable, c i represents the frequency of the i-th candidate dependent variable, |S| represents the total number of variables in the multiple candidate dependent variable sets, r i represents the score of the i-th candidate dependent variable.

4. The method according to claim 1, characterized in that, The regressing the exposure values of the multiple core dependent variables on the business metrics respectively to determine the factor coefficients between the multiple core dependent variables and the target variables with causal relationships includes: For each core dependent variable, respectively perform: obtaining T historical data of the core dependent variable and T historical data of the business metric, regressing the T historical data of the core dependent variable and the T historical data of the business metric to determine T initial exposure values of the core dependent variable on the business metric, and taking the mean of the T initial exposure values as the target exposure value of the core dependent variable on the business metric; Performing a single regression on the target exposure values of the multiple core dependent variables across sections to obtain the factor coefficients between the multiple core dependent variables and the target variables with causal relationships.

5. The method according to claim 4, characterized in that The performing a single regression on the target exposure values of the multiple core dependent variables across sections to obtain the factor coefficients between the multiple core dependent variables and the target variables with causal relationships includes: Detecting the value of the attribution error in the cross-sectional regression equation; If the value of the attribution error is 0, then according to the correlation relationship between the at least one target variable and the business metric, taking the factor coefficients between the multiple core dependent variables and the business metric in the cross-sectional regression equation as the factor coefficients between the multiple core dependent variables and the target variables with causal relationships.

6. The method according to claim 1, wherein The regressing the exposure values of the multiple core dependent variables on the business metrics respectively to determine the factor coefficients between the multiple core dependent variables and the target variables with causal relationships includes: For each core dependent variable, the following operations are respectively performed: obtaining T historical data of the core dependent variable and T historical data of the business indicator, performing regression on the T historical data of the core dependent variable and the T historical data of the business indicator, and determining T initial exposure values of the core dependent variable on the business indicator; Performing T regressions on the T initial exposure values of each of the multiple core dependent variables in cross-section to obtain T initial coefficients between each core dependent variable and the target variable having a causal relationship; Taking the mean of the T initial coefficients corresponding to each of the multiple core dependent variables as the factor coefficient between the corresponding core dependent variable and the target variable having a causal relationship.

7. An electronic device, characterized in that, Comprising a processor, a memory, and a communication interface, the communication interface, the memory, and the processor are connected through a bus; The communication interface is used for communicating with other devices; The memory stores a computer program, and the processor performs the following operations according to the computer program: Obtaining at least one target variable of the business to be processed respectively in T time periods, where T is an integer greater than 1; For each time period, the following operations are respectively performed: according to the pre-set association relationships between the at least one target variable and the shared business indicator respectively, locating a set of candidate dependent variables having a causal relationship with the at least one target variable in the corresponding time period; Filtering out multiple core dependent variables according to the frequencies of the candidate dependent variables in the T sets of candidate dependent variables; Performing regression on the exposure estimates of the multiple core dependent variables on the business indicator respectively to determine the factor coefficients between the multiple core dependent variables and the target variables having a causal relationship respectively.

8. The electronic device according to claim 7, wherein The processor filters out multiple core dependent variables according to the frequencies of the candidate dependent variables in the multiple sets of candidate dependent variables, and the specific operation is as follows: Counting the frequencies of each candidate dependent variable in the multiple sets of candidate dependent variables; Determining the score of each candidate dependent variable respectively according to each frequency and the total number of variables in the multiple sets of candidate dependent variables, and the score is positively correlated with the frequency; Sorting the multiple sets of candidate dependent variables in descending order of the scores, and filtering out the top K candidate dependent variables not less than a preset score threshold from the multiple sets of candidate dependent variables, where K is an integer greater than 1; Taking the top K candidate dependent variables as multiple core dependent variables.

9. The electronic device according to claim 7, wherein The processor performs regression on the exposure estimates of the multiple core dependent variables on the business indicator respectively to determine the factor coefficients between the multiple core dependent variables and the target variables having a causal relationship respectively, and the specific operation is as follows: For each core dependent variable, the following operations are respectively performed: obtaining T historical data of the core dependent variable and T historical data of the business indicator, performing regression on the T historical data of the core dependent variable and the T historical data of the business indicator, determining T initial exposure values of the core dependent variable on the business indicator, and taking the mean of the T initial exposure values as the target exposure value of the core dependent variable on the business indicator; Perform a single regression on the target exposure values of the multiple core dependent variables across the cross-section to obtain the factor coefficients between each of the multiple core dependent variables and the target variables with which there is a causal relationship.

10. The electronic device according to claim 7, characterized in that, The processor performs a regression on the exposure values of the multiple core dependent variables on the business metrics respectively, and determines the factor coefficients between each of the multiple core dependent variables and the target variables with which there is a causal relationship. The specific operation is as follows: For each core dependent variable, perform the following respectively: obtain T historical data of the core dependent variable and T historical data of the business metric, and perform a regression on the T historical data of the core dependent variable and the T historical data of the business metric to determine the T initial exposure values of the core dependent variable on the business metric; Perform T regressions on the T initial exposure values of each of the multiple core dependent variables across the cross-section to obtain the T initial coefficients between each core dependent variable and the target variable with which there is a causal relationship; The mean of the T initial coefficients corresponding to each of the multiple core dependent variables is used as the factor coefficient between the corresponding core dependent variable and the target variable with which there is a causal relationship.