Attribution analysis method, computer equipment and computer storage medium

By training multiple alternative attribution models and selecting the optimal model, the impact of unstructured factors on business metrics is quantified, solving the problem that existing technologies cannot quantify the degree of impact and achieving more objective and interpretable attribution analysis.

CN121435170APending Publication Date: 2026-01-30KINGDEE SOFTWARE(CHINA) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511642528.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-10
Publication Date
2026-01-30

AI Technical Summary

Technical Problem

Existing value tree attribution analysis methods rely on the computational relationships between indicators, which cannot quantify the impact of unstructured factors on top-level indicators. Furthermore, manual judgment is highly subjective and lacks interpretability.

Method used

By acquiring multiple sets of business datasets, training multiple candidate attribution models, learning model performance based on hyperparameter combinations, selecting the optimal model, quantifying the impact of unstructured factors on business indicators, and using model- and data-driven methods for attribution analysis.

Benefits of technology

It significantly improves the objectivity and interpretability of attribution analysis, reduces subjective bias caused by human judgment, and enhances the accuracy and reliability of the analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121435170A_ABST
    Figure CN121435170A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses an attribution analysis method, computer equipment and a computer storage medium. The alternative attribution model carries out attribution analysis of the business indexes and multiple influence factors of the business indexes on the index data and the influence factor data on the basis of each hyper-parameter combination in sequence, the optimal hyper-parameter combination of the alternative attribution model is determined according to the model performance expressed by the alternative attribution model on the basis of each hyper-parameter combination learning business data set, and the optimal hyper-parameter combination of the alternative attribution model is obtained. And determining a target attribution model according to the model performance represented by each alternative attribution model based on the optimal hyper-parameter combination. The influence degree of the influence factors on the indexes is determined through the model, quantitative analysis can be carried out on unstructured factors, and therefore potential influence factors of business index fluctuation can be revealed more comprehensively. On the basis of model and data driving instead of manual subjective judgment of the influence degree, the objectivity and interpretability of attribution analysis can be remarkably improved, subjective deviation caused by manual judgment is reduced, and the reliability and accuracy of attribution analysis are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing, specifically to an attribution analysis method, a computer device, and a computer storage medium. Background Technology

[0002] Current mainstream value tree attribution analysis methods rely on the calculation relationship between indicators in the indicator value tree. By calculating the contribution of each bottom indicator change to the top indicator change, and then analyzing the proportion of contribution, the degree of influence of each bottom indicator on the top indicator is determined.

[0003] However, attribution analysis relies on the calculation relationship between top-level and bottom-level indicators. For top-level and bottom-level indicators without a clear calculation relationship, it is difficult to determine the degree of influence of the bottom-level indicators on the top-level indicators, and the degree of influence cannot be quantified. The only recourse is to assign a corresponding influence value to these bottom-level indicators based on experience, but this method is highly subjective, making it difficult to objectively and accurately evaluate the degree of influence and lacking interpretability. Summary of the Invention

[0004] This application provides an attribution analysis method, computer equipment, and computer storage medium to quantify the impact of unstructured factors and more comprehensively reveal the potential influencing factors of business indicator fluctuations.

[0005] A first aspect of this application provides an attribution analysis method, the method comprising:

[0006] Obtain multiple sets of business datasets, each set of business datasets including indicator data of business metrics and influence factor data of multiple influencing factors corresponding to the business metrics;

[0007] For each alternative attribution model, multiple sets of the business datasets are input into the alternative attribution model, so that the alternative attribution model performs attribution analysis of the business indicators and their multiple influencing factors on the indicator data and the influencing factor data based on each hyperparameter combination in turn.

[0008] Based on the model performance of the candidate attribution models on the business dataset learned by each of the hyperparameter combinations, the optimal hyperparameter combination of the candidate attribution models is determined.

[0009] Based on the model performance of each of the candidate attribution models according to its optimal combination of hyperparameters, determine the target attribution model with the best performance among the multiple candidate attribution models;

[0010] The target attribution model is used to calculate the degree of influence of multiple influencing factors on the business indicator to be analyzed.

[0011] A second aspect of this application provides a computer device, the computer device comprising:

[0012] The acquisition unit is used to acquire multiple sets of business datasets, each set of business datasets including indicator data of business indicators and influence factor data of multiple influencing factors corresponding to the business indicators;

[0013] The training unit is used to input multiple sets of the business datasets into the candidate attribution model for each candidate attribution model, so that the candidate attribution model performs attribution analysis of the business indicators and their multiple influencing factors on the indicator data and the influencing factor data based on each hyperparameter combination in turn.

[0014] The first determining unit is used to determine the optimal hyperparameter combination of the candidate attribution models based on the model performance of the business dataset learned by the candidate attribution models based on each hyperparameter combination.

[0015] The second determining unit is used to determine the target attribution model with the best performance among the multiple candidate attribution models based on the model performance of each candidate attribution model based on its optimal hyperparameter combination.

[0016] The target attribution model is used to calculate the degree of influence of multiple influencing factors on the business indicator to be analyzed.

[0017] A third aspect of this application provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the method of the first aspect described above.

[0018] A fourth aspect of this application provides a computer storage medium storing instructions that, when executed on a computer, cause the computer to perform the method described in the first aspect.

[0019] A fifth aspect of this application provides a computer program product that, when run on a computer device, causes the computer device to perform the method described in the first aspect.

[0020] As can be seen from the above technical solutions, the embodiments of this application have the following advantages:

[0021] The candidate attribution models learn the relationships between business indicators and their multiple influencing factors based on each hyperparameter combination. The optimal hyperparameter combination is determined based on the model performance of each candidate attribution model on the business dataset. Finally, the target attribution model is determined based on the model performance of each candidate attribution model on its optimal hyperparameter combination. Therefore, by determining the degree of influence of influencing factors on indicators through the target attribution model, quantitative analysis of unstructured factors can be performed. Specifically, for business indicators and influencing factors without clear computational relationships, the degree of influence can be quantified by learning the inherent patterns in the data through the model. This avoids the subjectivity and lack of interpretability caused by manually assigning numerical values ​​to the degree of influence, thus revealing more comprehensively the potential influencing factors of business indicator fluctuations. Based on model and data-driven approaches, rather than subjective human judgment of the degree of influence, the objectivity and interpretability of attribution analysis can be significantly improved, reducing subjective bias caused by human judgment and enhancing the reliability and accuracy of attribution analysis. Attached Figure Description

[0022] Figure 1 This is a schematic diagram of the network framework in an embodiment of this application;

[0023] Figure 2 This is a flowchart illustrating the attribution analysis method in an embodiment of this application;

[0024] Figure 3 for Figure 2 A flowchart illustrating a specific implementation of step 202 in the illustrated embodiment;

[0025] Figure 4 This is a schematic diagram illustrating the correspondence between various indicators and their influencing factors in the interface of this application embodiment;

[0026] Figure 5 This is an exemplary schematic diagram of the attribution analysis result display interface in an embodiment of this application;

[0027] Figure 6 This is another exemplary schematic diagram of the attribution analysis result display interface in the embodiments of this application;

[0028] Figure 7 This is a schematic diagram of the structure of a computer device in an embodiment of this application;

[0029] Figure 8 This is another schematic diagram of the structure of the computer device in the embodiments of this application. Detailed Implementation

[0030] This application provides an attribution analysis method, computer equipment, and computer storage medium to quantify the impact of unstructured factors and more comprehensively reveal the potential influencing factors of business indicator fluctuations.

[0031] Please see Figure 1 The network framework in this embodiment includes:

[0032] The business server 100 and the terminal cluster; the terminal cluster may include: terminal devices 200a, terminal devices 200b, terminal devices 200c, ..., terminal devices 200n and other terminal devices.

[0033] The aforementioned business server 100 can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud databases, cloud services, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The terminal devices (including terminal devices 200a, 200b, 200c, ..., 200n) can be smartphones, tablets, laptops, desktop computers, PDAs, mobile internet devices (MIDs), wearable devices (such as smartwatches, smart bracelets, etc.), smart computers, smart in-vehicle systems, and other intelligent terminals.

[0034] The service server 100 can establish communication connections with each terminal device in the terminal cluster, and the terminal devices in the terminal cluster can also establish communication connections with each other. In other words, the service server 100 can establish communication connections with each terminal device among terminal devices 200a, 200b, 200c, ..., 200n. For example, terminal device 200a can establish a communication connection with the service server 100. Terminal devices 200a and 200b can establish a communication connection, and terminal devices 200a and 200c can also establish a communication connection. The communication connection method is not limited; it can be established directly or indirectly through wired communication or wireless communication, etc., depending on the actual application scenario. This application does not impose any restrictions on this.

[0035] It should be understood that, such as Figure 1Each terminal device in the terminal cluster shown can have an application client installed. When the application client runs on each terminal device, it can interact with the business server 100, allowing the business server 100 to receive business data from each terminal device (such as financial management data uploaded by users through the terminal device). This application client can be a financial management application, enterprise affairs management application, browser application, social application, instant messaging application, live streaming application, game application, short video application, video application, music application, shopping application, novel application, payment application, or any other application client with the ability to display text, images, audio, and video data. The specific application client can be determined based on the actual application scenario requirements and is not limited here. This application client can be a standalone client or an embedded sub-client integrated into a client (such as a financial management client, enterprise affairs management client, etc.), depending on the actual application scenario and is not limited here.

[0036] The current mainstream value tree attribution analysis method relies on the calculation relationship between indicators and influencing factors in the indicator value tree. It calculates the contribution of each bottom-level influencing factor to the change in the top-level indicator, and then analyzes the percentage of contribution to identify the impact of the bottom-level factors on the top-level indicator. For example, if an insurance company's gross profit decreases by 10 million, and the influencing factors for the gross profit indicator are sales revenue and sales cost (gross profit = sales revenue - sales cost), with sales revenue decreasing by 7 million and sales cost increasing by 3 million, then sales revenue contributes 7 million to the decrease in gross profit, a positive impact of 70%; sales cost contributes 3 million to the decrease in gross profit, a positive impact of 30%.

[0037] Current mainstream attribution analysis methods can only rely on pre-defined calculation relationships for attribution, and cannot attribute factors that are not calculated, nor can they quantify the contribution of these non-calculated factors to the top-level indicator. For example, the calculation relationship "sales revenue = unit price × sales volume" can identify the factors affecting the sales revenue indicator, but for other non-calculated factors, such as sales discounts and contract conversion rates, it is impossible to clearly determine their degree of influence on the sales revenue indicator.

[0038] For example, the sales of milk powder companies in a certain region are affected by factors such as the local birth rate, market share, and CPI. The higher the local birth rate, the more positive it is for milk powder sales. However, this effect cannot be constructed through simple calculations. For instance, it is impossible to explain how much of the 15% increase in sales is caused by a 2% increase in the birth rate.

[0039] Furthermore, when there are many non-calculated factors influencing complex fluctuations in indicators (such as a sudden 10% drop in premium income), it is necessary to manually rely on experience to judge the degree of influence of each factor. However, this approach is highly subjective, lacks data-driven and quantitative basis, makes it difficult to objectively and accurately evaluate the degree of influence, and lacks interpretability.

[0040] For example, when insurance companies discover a decline in quarterly insurance business profits, existing systems can only display the impact of calculated influencing factors (such as rising costs and declining sales) on the fluctuation of this indicator. However, they cannot quantify the impact of unstructured factors such as the rate of change in insurance product mix, the concentration of business sources, and the proportion of group / individual insurance premiums. For instance, if the proportion of group insurance premiums decreases from 58% to 56%, it cannot explain the specific impact of this factor's fluctuation on the profit decline. If the degree of impact is judged manually, there is a lack of data support, which introduces a large degree of subjectivity, resulting in low accuracy and reliability of attribution analysis.

[0041] To address the aforementioned technical problems, an attribution analysis method according to an embodiment of this application is proposed. The following will combine... Figure 1 The network framework described herein is used to illustrate the attribution analysis method in the embodiments of this application.

[0042] Please see Figure 2 One embodiment of the attribution analysis method in this application includes:

[0043] 201. Obtain multiple sets of business datasets, each set of business datasets including indicator data of business indicators and influence factor data of multiple influencing factors corresponding to the business indicators;

[0044] The method of this embodiment can be applied to a computer device, which may be... Figure 1 The network framework shown includes a business server 100 or various terminal devices. After acquiring multiple sets of business datasets, the next step is to process and analyze these data to build a model suitable for attribution analysis. Specifically, the computer equipment will train multiple candidate attribution models based on these business datasets and select the optimal model through a series of performance evaluation steps. This process includes, but is not limited to, key steps such as data preprocessing, feature engineering, model training, and performance metric calculation.

[0045] The multiple influencing factors corresponding to the business metrics can be preset by the program or set by the user according to the actual business scenario; there is no limitation here.

[0046] This business dataset can be acquired through various means. For example, relevant data can be extracted from the company's internal database, including historical business records, financial statements, and market research data; supplementary information can also be obtained from external data sources, such as industry reports, publicly available statistical data, or third-party data services. Furthermore, user-inputted custom data can be incorporated to enhance the diversity and relevance of the business data.

[0047] 202. For each alternative attribution model, multiple sets of the business datasets are input into the alternative attribution model, so that the alternative attribution model performs attribution analysis of the business indicators and their multiple influencing factors on the indicator data and the influencing factor data based on each hyperparameter combination in turn.

[0048] In this embodiment, these attribution models can be built on different algorithmic architectures, such as the GLM (Generalized Linear Model), GAM (Generalized Additive Model), Random Forest, and XGBoost (Extreme Gradient Boosting). Each model learns from the input business dataset, performing attribution analysis on the business metrics and their multiple influencing factors based on the metric data and influencing factor data in the business dataset; that is, learning the relationship between the business metrics and their multiple influencing factors.

[0049] For example, the model learning process involves optimizing multiple hyperparameter combinations. Each alternative attribution model tries various hyperparameter configurations to explore the best fit. By learning from multiple business datasets, the model can gradually capture the complex correlation patterns between business metrics and influencing factors, including linear relationships, nonlinear relationships, and interaction effects. Interaction effects refer to the combined impact of potential interactions between different influencing factors on business metrics. For example, when analyzing sales revenue, sales discounts and market share may have an interaction effect; that is, adjusting sales discounts has a more significant effect on increasing sales revenue when market share is higher. Capturing this interaction effect helps to more comprehensively understand the mechanisms of change in business metrics. This learning mechanism enables the model to effectively model factors that lack direct mathematical expression, thereby achieving a quantitative analysis of the impact of unstructured factors on metric fluctuations.

[0050] Multiple hyperparameter combinations can be pre-defined by the program or dynamically adjusted according to actual needs. These hyperparameters may include learning rate, regularization parameter, tree depth, number of leaf nodes, etc., depending on the algorithm architecture and the complexity of the business scenario. By trying different hyperparameter combinations, the model can find the optimal hyperparameter configuration during training, thereby improving the accuracy and stability of attribution analysis.

[0051] 203. Based on the model performance of the candidate attribution models on the business dataset learned by each of the hyperparameter combinations, determine the optimal hyperparameter combination of the candidate attribution models;

[0052] At this stage, the model's performance can be evaluated through cross-validation to ensure its generalization ability. Specifically, the business dataset can be divided into training and test sets. The model is trained using the training set and its prediction accuracy is tested on the test set. By comparing the performance metrics of different candidate models on the business dataset, such as mean squared error (MSE), mean absolute error (MAE), or coefficient of determination (R²), the optimal hyperparameter combination for each candidate attribution model can be selected. These performance metrics reflect the model's ability to explain fluctuations in business indicators and can reflect the model's fitting effect and predictive ability from different perspectives, providing a reliable foundation for subsequent attribution analysis.

[0053] 204. Based on the model performance of each of the candidate attribution models according to its optimal combination of hyperparameters, determine the target attribution model with the best performance among the multiple candidate attribution models;

[0054] In this step, based on the model performance of each candidate attribution model according to its optimal hyperparameter combination, the target attribution model with the best model performance is determined from among multiple candidate attribution models. The target attribution model can be used to calculate the degree of influence of multiple influencing factors on the business indicator to be analyzed.

[0055] The process of determining a target attribution model can comprehensively consider multiple performance dimensions to ensure its applicability and accuracy in practical applications. In addition to the aforementioned metrics such as mean squared error, mean absolute error, and coefficient of determination, other evaluation criteria can be introduced, such as the model's computational efficiency, interpretability, and reliability and stability against outliers. A comprehensive consideration of these factors helps to select the target attribution model most suitable for a specific business scenario.

[0056] For example, in scenarios with high real-time requirements, the computational efficiency of a model may become a key evaluation factor; while in scenarios where the analysis results need to be presented to decision-makers, the interpretability of the model is particularly important. Model interpretability refers to the ability to present the results to users in an intuitive and easy-to-understand way, enabling them to comprehend how the model arrives at specific conclusions. This is especially important for scenarios involving complex business logic, as decision-makers typically need to clearly understand the specific contribution of each influencing factor to the final result in order to formulate appropriate strategies. Furthermore, the reliability and stability of the model against outliers is also a crucial factor, especially when data quality varies or noise is present. The model needs to possess a certain degree of robustness to interference to ensure the stability of the analysis results.

[0057] By scoring or weighting the performance of candidate attribution models across these dimensions, the selection range can be further narrowed down, ultimately determining the optimal target attribution model.

[0058] In this embodiment, each candidate attribution model sequentially learns the relationship between business indicators and their multiple influencing factors based on each hyperparameter combination. The optimal hyperparameter combination is determined based on the model performance of each candidate attribution model on the business dataset. Finally, the target attribution model is determined based on the model performance of each candidate attribution model on its optimal hyperparameter combination. Therefore, by determining the degree of influence of influencing factors on indicators through the target attribution model, unstructured factors can be quantitatively analyzed, thus revealing more comprehensively the potential influencing factors of business indicator fluctuations. Based on model and data-driven approaches, rather than subjective human judgment of the degree of influence, the objectivity and interpretability of attribution analysis can be significantly improved, reducing subjective bias caused by human judgment and enhancing the reliability and accuracy of attribution analysis.

[0059] based on Figure 2 In one optional implementation of the illustrated embodiment, during the training and performance testing of the candidate attribution models, a search algorithm can be used to traverse each hyperparameter combination to test the degree to which the hyperparameter combination improves model performance. Specifically, one process step is as follows: Figure 3 As shown:

[0060] 2021. Determine the combination of hyperparameters to be evaluated;

[0061] 2022. Input multiple sets of the business datasets into the alternative attribution model, so that the alternative attribution model performs attribution analysis on the indicator data and the influencing factor data based on the hyperparameter combination to be evaluated, comparing the business indicators with their multiple influencing factors.

[0062] In this step, the hyperparameter combinations to be evaluated can be pre-defined or dynamically generated. In practice, methods such as grid search, random search, or Bayesian optimization can be used to determine the range of hyperparameter combinations to be evaluated. These methods can effectively narrow the hyperparameter search space and improve model tuning efficiency. For example, grid search finds the optimal configuration by setting discrete values ​​for each hyperparameter and exhaustively combining them; while Bayesian optimization uses the previous evaluation results of hyperparameter combinations to build a probabilistic model to guide the selection of hyperparameters that are more likely to improve performance.

[0063] The initially determined combination of hyperparameters to be evaluated can be randomly selected from the hyperparameter search space (range). For example, in Bayesian optimization, the initial hyperparameter combination is usually determined by random selection (random sampling). This initial step is called the initialization phase or prior construction phase. Subsequent optimization processes adjust the hyperparameter configuration step by step based on these prior initial evaluation results. In the initialization phase, the randomly selected hyperparameter combination provides a starting point for the model, allowing the optimization algorithm to iteratively improve upon it. As the evaluation process progresses, methods such as Bayesian optimization use the performance data of the evaluated hyperparameter combinations to construct a probabilistic model reflecting the relationship between hyperparameters and model performance. This model can predict which hyperparameter combinations are more likely to improve model performance, thus guiding the subsequent search direction.

[0064] Hyperparameter combinations are set before the candidate attribution model begins training, determining the model's architecture and learning process. The business dataset serves as the foundation for model training and testing, and its quality and diversity directly impact the effectiveness of hyperparameter tuning. For example, a large number of missing or outlier values ​​in the business dataset may mislead the direction of hyperparameter optimization, causing the model to perform well on the training set but poorly on the test set or in real-world applications. Therefore, rigorous data cleaning and preprocessing are necessary before inputting the business dataset, including missing value imputation, outlier detection and handling, and data standardization or normalization, to ensure data quality and consistency.

[0065] Specifically, the combination of hyperparameters indirectly determines the model's ability to fit the indicator data and influencing factor data in the business dataset by influencing the model's weight update rules, regularization strength, or tree structure complexity.

[0066] Therefore, after inputting multiple sets of business datasets into the candidate attribution models, the models are trained based on the current combination of hyperparameters to be evaluated, learning the complex relationships between business metrics and multiple influencing factors. This process involves extracting data features, recognizing patterns, and mining potential regularities. During the model's learning of the business dataset, the internal weight parameters are gradually adjusted according to a pre-defined algorithm architecture to better fit the potential relationships in the data. For example, the Generalized Additive Model (GAM) can capture nonlinear relationships through a smoothing function, while the Random Forest model can reveal complex interaction effects by constructing multiple decision trees.

[0067] Simultaneously, after the model is trained on the business dataset, its performance based on the evaluated hyperparameter combination can be further tested on the test set. During performance evaluation on the test set, the model's predictions are compared with the true values ​​to quantify its fit and generalization ability. The core objective of this stage is to verify whether the model's performance on unseen data is stable and reliable. For example, the presence of systematic bias in the model can be analyzed by calculating the residual distribution between the predicted and true values. If the residuals exhibit a random distribution without a clear pattern, it indicates that the model has learned the business data sufficiently; conversely, if the residuals show a certain regularity, it may indicate that the model has failed to fully capture the potential relationship between business indicators and influencing factors, or that there is a risk of overfitting.

[0068] 2023. Determine whether the convergence condition is met. If not, proceed to step 2024. If yes, end the model training process.

[0069] Convergence criteria are typically set based on preset performance thresholds or iteration limits. For example, if the model performance does not show a significant improvement in a number of consecutive iterations, it can be determined that the convergence criterion has been met; or, when the maximum number of iterations is reached, the search process is terminated regardless of whether the performance reaches the expected level. This mechanism can effectively avoid wasting computational resources while ensuring the controllability of the model tuning process.

[0070] If the convergence condition is not met when the evaluation of the hyperparameter combination to be evaluated is completed, the next hyperparameter combination to be evaluated is determined and the training process is repeated until the convergence condition is met.

[0071] 2024. Determine the next hyperparameter combination to be evaluated based on the model performance of the business dataset learned by the alternative attribution model based on the hyperparameter combination to be evaluated, and return to step 2022.

[0072] After determining the model performance based on the hyperparameter combinations to be evaluated, if the convergence condition has not yet been met, a specific search algorithm can be used to dynamically adjust and generate the next hyperparameter combination to be evaluated. For example, when using Bayesian optimization, the algorithm constructs a probabilistic model based on the currently evaluated hyperparameter combinations and their corresponding model performance to predict which hyperparameter combinations are more likely to improve model performance. Compared to traditional grid search or random search, this method can find a near-optimal hyperparameter configuration in a shorter time, thus significantly improving tuning efficiency. By systematically testing and evaluating the performance of each hyperparameter combination to be evaluated, the optimal hyperparameter configuration can be gradually approximated.

[0073] After determining the next combination of hyperparameters to be evaluated, multiple sets of business datasets can be input into the candidate attribution model, and the training and evaluation process in step 2022 can be repeated. This iterative process will continue until the preset convergence condition is met. In each iteration, the model's performance will be recorded in detail and used to guide subsequent hyperparameter tuning. In this way, not only can the model's fitting ability be optimized, but the risk of getting trapped in local optima can also be effectively avoided.

[0074] Therefore, in this embodiment, the next hyperparameter combination to be evaluated is determined by the model performance corresponding to the already evaluated hyperparameter combinations. Based on this method, the model performance based on each hyperparameter combination is evaluated sequentially. After multiple iterations, the optimal hyperparameter configuration can be gradually selected. This method not only improves the efficiency of model tuning but also ensures the accuracy and reliability of model attribution analysis training. Through detailed evaluation of each hyperparameter combination, the impact of different configurations on model performance can be more comprehensively understood, providing a more scientific basis for subsequent attribution analysis.

[0075] One possible implementation for determining the next hyperparameter combination to be evaluated based on model performance is to use a Bayesian search algorithm. During the Bayesian search, the model performance index of the candidate attribution model on the evaluated hyperparameter combination can be calculated based on the prediction results output by the candidate attribution model learning from the business dataset, thus obtaining the model performance index corresponding to the evaluated hyperparameter combination. Further, a probabilistic model is constructed. Multiple model performance indices corresponding to the evaluated hyperparameter combinations can be input into this probabilistic model. The probabilistic model can then predict the model performance index corresponding to the unevaluated hyperparameter combination based on the model performance indices corresponding to the multiple evaluated hyperparameter combinations, and calculate the confidence level of the predicted model performance index.

[0076] In this process, the probabilistic model calculates the confidence level corresponding to its predicted performance index, which is equivalent to quantifying the uncertainty of the model's performance index prediction and representing the reliability of the predicted value. A higher confidence level indicates that the probabilistic model's prediction of the performance index of the unevaluated hyperparameter combination is more reliable, thus providing a more scientific basis for selecting the next hyperparameter combination to be evaluated. By combining the predicted performance index and its confidence level, the Bayesian search algorithm can prioritize hyperparameter combinations with high potential performance improvement and low prediction uncertainty as the next round of evaluation targets. This method not only improves search efficiency but also maximizes the effect of model tuning with limited computational resources.

[0077] In practical applications, Gaussian Process Regression (GPR) and other methods can be used to construct probabilistic models. The core of GPR lies in using the correlation between evaluated data points to infer the performance of unevaluated data points. Furthermore, to further improve search accuracy, an acquisition function, such as Expected Improvement (EI) or Upper Confidence Bound (UCB), can be introduced to quantify the potential value of each unevaluated hyperparameter combination. For example, when predicting the model performance index and its corresponding confidence level for each unevaluated hyperparameter combination, an acquisition function can be constructed. Based on this acquisition function, the expected improvement of each unevaluated hyperparameter combination can be calculated, and the unevaluated hyperparameter combination with the highest expected improvement can be selected as the next hyperparameter combination to be evaluated.

[0078] These acquisition functions comprehensively consider prediction performance metrics and confidence levels, helping the algorithm find a balance between exploring new data points and utilizing existing information, thereby achieving efficient global optimization. Furthermore, during iteration, as more hyperparameter combinations are evaluated and added to the training set, the probabilistic model continuously updates its predictive capabilities. This dynamic adjustment mechanism allows the Bayesian search algorithm to gradually focus on hyperparameter configurations with better performance and ultimately converge to a solution close to the optimum.

[0079] Therefore, in this embodiment, evaluating the performance of different hyperparameter combinations sequentially through Bayesian search significantly improves the efficiency and accuracy of model tuning. This method not only reduces reliance on manual intervention but also quickly locates potential optimal configurations in a complex hyperparameter space. By dynamically adjusting the search strategy, Bayesian optimization achieves a good balance between exploration and exploitation, thereby avoiding the risk of getting trapped in local optima.

[0080] based on Figure 2In the illustrated embodiment, when determining the optimal hyperparameter combination of candidate attribution models, one optional implementation method is to calculate the model performance index of each hyperparameter combination based on the prediction results output by the candidate attribution model after learning the business dataset using that hyperparameter combination. Here, cross-validation methods (such as k-fold cross-validation) can be used to divide the business dataset into a training set and a test set. The training set is used for the model to learn the correlation between business indicators and influencing factors, while the test set is used to evaluate the model's generalization ability to unseen data.

[0081] After obtaining the model performance index corresponding to each hyperparameter combination, the hyperparameter combination with the best model performance index can be determined from multiple hyperparameter combinations as the optimal hyperparameter combination for the candidate attribution model.

[0082] Taking the mean squared error (MSE) as an example, its calculation logic is to average the squared differences between the model's predicted values ​​and the actual values ​​on the test set; the smaller the value, the more accurate the model's prediction. The mean absolute error (MAE) is the average of the absolute values ​​of these differences, more intuitively reflecting the average level of prediction error. The coefficient of determination (R²) quantifies the proportion of fluctuations in business indicators explained by the model by comparing the degree of variation between the model's predicted values ​​and the actual values; the closer its value is to 1, the better the model's fit. Model performance indicators obtained through this standardized calculation method can objectively and quantitatively reflect the learning effect of alternative attribution models under corresponding hyperparameter combinations, avoiding the bias of subjective judgment and providing a reliable numerical basis for the subsequent selection of the optimal hyperparameter combination.

[0083] based on Figure 2 In the illustrated embodiment, when determining the optimal target attribution model among multiple candidate attribution models, one optional implementation method is to calculate the model performance index of each candidate attribution model on its optimal hyperparameter combination based on the prediction results output by the candidate attribution model after learning the business dataset based on its optimal hyperparameter combination. Furthermore, among the multiple candidate attribution models, the candidate attribution model with the optimal model performance index can be determined as the target attribution model.

[0084] When determining the target attribution model, the performance of the optimal hyperparameter combination of all candidate models can be compared across multiple dimensions. For example, in the attribution analysis of user conversion rate on an e-commerce platform, candidate models include GAM, Random Forest (RF), and XGBoost. After optimization through Bayesian search, the performance metrics of various models based on their optimal hyperparameter combinations are as follows: XGBoost has the lowest MSE (0.07), RF has the highest R² (0.93), while GAM has the highest interpretability score (measured by the average absolute contribution of SHAP values) (0.85). At this point, the priority of "accurate prediction" versus "interpretability" can be weighed: if the goal is to provide operators with the specific contribution of each influencing factor (such as page load time, recommendation algorithm performance) to formulate optimization strategies, then GAM can become the target model based on its interpretability advantage; if the goal is to provide input features for machine learning models (such as feature engineering for conversion rate prediction models), then XGBoost is more suitable due to its high prediction accuracy.

[0085] For example, the e-commerce platform may choose GAM as the target attribution model because it can clearly show conclusions such as "for every 1 second increase in page load time, the conversion rate decreases by 2.3%" and "if the personalization of the recommendation algorithm is improved by 10%, the conversion rate increases by 1.8%", which is easy for operations personnel to understand and implement. Operations personnel can formulate corresponding optimization strategies based on the impact of the above multiple influencing factors on user conversion rates.

[0086] Through the above process, the attribution analysis method can select the target model that best meets business needs from the candidate models, which not only ensures the accuracy and stability of prediction, but also takes into account interpretability and decision-making practicality, providing a scientific and objective basis for the formulation of business strategies.

[0087] based on Figure 2 In the illustrated embodiment, when acquiring multiple sets of business datasets, one optional implementation method is to receive user operations on setting influencing factors for business metrics, and determine multiple influencing factors corresponding to the business metrics based on these operations. These influencing factor setting operations may include one or more of adding, deleting, or modifying influencing factors. Furthermore, for each business metric, the metric data and the influencing factor data of the multiple influencing factors corresponding to that business metric can be acquired to obtain the business dataset.

[0088] like Figure 4The interface displays the correspondence between various indicators and their influencing factors. As shown in the figure, multiple influencing factors corresponding to each indicator can be displayed, and these factors can be flexibly configured by the user according to actual business needs. For example, users can add influencing factors such as "opportunity responsiveness," "discount rate," and "opportunity amount" to the "conversion rate" node. The impact of these factors on the conversion rate is difficult to express through calculation relationships but rather exhibits a correlational effect.

[0089] In addition, users can also intuitively add influencing factors through the interface, such as adding "number of customers" to the "sales" metric, or delete irrelevant or redundant factors, such as removing factors that have no significant impact on current business objectives. Furthermore, existing influencing factors can be modified, such as adjusting their weights or redefining their quantification methods to more accurately reflect the actual business situation.

[0090] After setting the influencing factors, the system automatically collects data related to these factors and the corresponding indicator data, integrating them into the business dataset. This process ensures the comprehensiveness and relevance of the data, thus providing high-quality input for subsequent model training.

[0091] Furthermore, the ability to support user-defined influencing factors provides greater flexibility for attribution analysis in different scenarios. For example, in assessing loan default risk in the financial sector, factors such as "borrower income level," "credit score," and "debt ratio" can be set as core factors based on specific business needs; while in the healthcare sector, when studying patient recovery probabilities, factors such as "age," "treatment plan," and "lifestyle habits" can be included in the considerations. This modular design makes the attribution analysis method highly versatile and adaptable to the needs of various industries and application scenarios.

[0092] based on Figure 2 In the illustrated embodiment, when acquiring multiple sets of business datasets, an optional implementation method is to preprocess the business datasets, i.e., acquire multiple sets of original business datasets, each set including indicator data of business metrics and influence factor data of multiple influencing factors corresponding to the business metrics. Preprocessing can be performed on the indicator data and / or influence factor data in each set of original business datasets to obtain multiple sets of business datasets.

[0093] This preprocessing operation can involve, for example, adding the mean of the variable from other business records to a missing business record in the indicator data and / or influencing factor data, and / or removing the missing business record. For instance, suppose the variable "discount rate" has 10 values, 8 of which are present (e.g., 5%, 10%, 8%, etc.), and 2 are empty. The average discount rate of these 8 present values ​​(let's say 7.5%) can be calculated first, and then 7.5% can be added to the 2 empty values. This method is simple and fast, maintaining the same sample size in the dataset. Alternatively, if the "opportunity amount" value in a sales record is empty, the entire sales record can be removed from the training data. This method avoids introducing false information and ensures that the preprocessed data is complete.

[0094] This preprocessing operation can also involve normalizing the variables in the indicator data and / or influencing factor data based on the mean and standard deviation of the variable's multiple values. This is because different business indicators or influencing factors can have vastly different values ​​and units (dimensions). For example, the opportunity amount might be hundreds of thousands or millions, the discount rate might be a decimal between 0 and 1 (e.g., 0.1 represents a 10% discount), and the conversion rate might also be a decimal between 0 and 1. If such data is directly fed into the model, the model might consider the opportunity amount to be huge and extremely important, with a higher impact, while the discount rate, being small, is insignificant and has a low impact. This is clearly inconsistent with business logic. Therefore, it is necessary to normalize all variables so that the value range of each variable is limited to the same interval (e.g., [0,1] or [-1,1]), thereby eliminating the interference of differences in dimensions and values ​​on model learning.

[0095] For example, Z-score normalization can be calculated using the following formula: subtract the mean of each variable from its value, then divide by its standard deviation to obtain the standardized value. This method not only makes the data more comparable but also accelerates the model's convergence speed and improves training efficiency.

[0096] Therefore, through normalization, the model can truly focus on the intrinsic relationship between each factor and the indicator, rather than being misled by their numerical magnitude. Normalized data can significantly accelerate model training, help the algorithm find the optimal solution more quickly, and also improve the model's numerical stability and generalization ability.

[0097] After model training and selection are completed, the computer equipment further analyzes the new business data to be processed using the selected target attribution model. The core of this step is applying the model's learning outcomes to real-world business scenarios, helping users understand the key drivers behind changes in specific business metrics. Therefore, the computer equipment can acquire the business data to be processed, which may include the metric data of the business metric to be analyzed and the impact factor data of multiple influencing factors corresponding to that business metric. The business data to be processed can be input into the target attribution model, which then calculates the degree of influence of the multiple influencing factors on the business metric to be analyzed based on the metric data and impact factor data in the business data.

[0098] Subsequently, the computer equipment can display the degree of influence of multiple influencing factors output by the target attribution model on the business indicator to be analyzed. In this way, the actual impact of non-calculated relationship factors that are difficult to handle by traditional methods can be revealed, thereby providing more comprehensive and accurate data support for enterprise decision-making. Optionally, improvement suggestions for the business indicator to be analyzed can also be displayed based on the influencing factors and their degree of influence.

[0099] like Figure 5 The interface displays the attribution analysis results, showing various factors influencing net profit growth, including revenue growth and cost reduction. Gross profit growth and expense reduction further support profit growth.

[0100] In addition, the interface can also display improvement suggestions, such as maintaining revenue growth, optimizing the supply chain to reduce costs, controlling management expenses, improving cost efficiency, focusing on the long-term balance between gross profit margin and expense ratio, and preventing the risk of cost rebound. These measures can further improve profits, and the company's managers can formulate corresponding management measures based on these suggestions to achieve the company's long-term stable development.

[0101] Of course, besides such Figure 5 The interface shown can also display the attribution analysis results in other forms, such as... Figure 6 As shown, taking the attribution analysis of abnormal year-on-year sales revenue as an example, the attribution analysis results can be visualized using a dynamic tree diagram to display each influencing factor. The tree diagram can display intermediate indicators corresponding to multiple influencing factors and show the various influencing factors that affect the target indicator along the flow path of their contribution. It also displays the degree of influence of each factor on the indicator and its corresponding value, allowing users to intuitively understand the relationships between factors and their specific effects on the target indicator. This path visualization method not only clearly presents the causal chain in complex business scenarios but also helps users quickly locate key driving factors, thereby formulating more targeted optimization strategies.

[0102] The attribution analysis method in the embodiments of this application has been described above. The computer device in the embodiments of this application is described below. Please refer to [link / reference]. Figure 7 One embodiment of the computer device in this application includes:

[0103] The acquisition unit is used to acquire multiple sets of business datasets, each set of business datasets including indicator data of business indicators and influence factor data of multiple influencing factors corresponding to the business indicators;

[0104] The training unit is used to input multiple sets of the business datasets into the candidate attribution model for each candidate attribution model, so that the candidate attribution model performs attribution analysis of the business indicators and their multiple influencing factors on the indicator data and the influencing factor data based on each hyperparameter combination in turn.

[0105] The first determining unit is used to determine the optimal hyperparameter combination of the candidate attribution models based on the model performance of the business dataset learned by the candidate attribution models based on each hyperparameter combination.

[0106] The second determining unit is used to determine the target attribution model with the best performance among the multiple candidate attribution models based on the model performance of each candidate attribution model based on its optimal hyperparameter combination.

[0107] The target attribution model is used to calculate the degree of influence of multiple influencing factors on the business indicator to be analyzed.

[0108] In a preferred embodiment of this invention, the training unit is specifically used for:

[0109] Determine the hyperparameter combination to be evaluated, and input multiple sets of the business datasets into the candidate attribution model so that the candidate attribution model performs attribution analysis on the indicator data and the influence factor data based on the hyperparameter combination to be evaluated.

[0110] Based on the model performance of the business dataset learned by the candidate attribution model on the hyperparameter combination to be evaluated, the next hyperparameter combination to be evaluated is determined, and the process of inputting multiple sets of the business datasets into the candidate attribution model is resumed until the convergence condition is met.

[0111] In a preferred embodiment of this invention, the training unit is specifically used for:

[0112] Based on the prediction results output by the candidate attribution model learning the business dataset, calculate the model performance index of the candidate attribution model on the hyperparameter combination to be evaluated, and obtain the model performance index corresponding to the evaluated hyperparameter combination.

[0113] Construct a probabilistic model and input the model performance index corresponding to the evaluated hyperparameter combination into the probabilistic model so that the probabilistic model can predict the model performance index corresponding to the unevaluated hyperparameter combination and the confidence level of the model performance index based on the model performance index corresponding to the evaluated hyperparameter combination.

[0114] Construct a data acquisition function, and calculate the expected improvement of the unevaluated hyperparameter combinations based on the model performance index and its confidence level for each unevaluated hyperparameter combination using the data acquisition function;

[0115] The unevaluated hyperparameter combination with the highest expected improvement is selected as the next hyperparameter combination to be evaluated.

[0116] In a preferred embodiment of this invention, the first determining unit is specifically used for:

[0117] For each hyperparameter combination, the model performance index of the candidate attribution model on the hyperparameter combination is calculated based on the prediction results output by the candidate attribution model after learning the business dataset based on the hyperparameter combination.

[0118] Among the multiple hyperparameter combinations, the hyperparameter combination that corresponds to the optimal model performance index is determined as the optimal hyperparameter combination for the candidate attribution model.

[0119] In a preferred embodiment of this invention, the second determining unit is specifically used for:

[0120] For each of the candidate attribution models, the model performance index of the candidate attribution model on its optimal hyperparameter combination is calculated based on the prediction results output by the candidate attribution model after learning the business dataset based on its optimal hyperparameter combination.

[0121] Among the various candidate attribution models, the candidate attribution model with the best model performance index is determined as the target attribution model.

[0122] In a preferred embodiment of this invention, the acquisition unit is specifically used for:

[0123] Receive user's operation to set the influencing factors of the business indicator, and determine multiple influencing factors corresponding to the business indicator based on the influencing factor setting operation;

[0124] For each of the business metrics, obtain the metric data of the business metrics and the influence factor data of the multiple influencing factors corresponding to the business metrics to obtain the business dataset.

[0125] In a preferred embodiment of this invention, the acquisition unit is specifically used for:

[0126] Obtain multiple sets of original business datasets. Each set of original business datasets includes indicator data of business metrics and influence factor data of multiple influencing factors corresponding to the business metrics.

[0127] Preprocessing is performed on the indicator data and / or the influencing factor data in each set of the original business datasets to obtain multiple sets of the business datasets; wherein, the preprocessing includes:

[0128] For the missing values ​​of business records in the indicator data and / or the impact factor data, the average value of the variable in other business records is added to the business record, and / or the business record is removed;

[0129] And / or,

[0130] For the variables in the indicator data and / or the influence factor data, the values ​​of the variable are normalized according to the mean and standard deviation of the multiple values ​​of the variable.

[0131] In a preferred embodiment of this invention, the computer device further includes an analysis unit, used for:

[0132] Acquire business data to be processed, which includes indicator data of business indicators to be analyzed and influence factor data of multiple influencing factors corresponding to the business indicators to be analyzed.

[0133] The business data to be processed is input into the target attribution model so that the target attribution model can calculate the degree of influence of multiple influencing factors on the business indicator to be analyzed based on the indicator data and influencing factor data in the business data to be processed.

[0134] The system displays the degree of influence of multiple influencing factors output by the target attribution model on the business indicator to be analyzed, and / or displays suggestions for improving the business indicator based on the influencing factors and their degree of influence.

[0135] In this embodiment, the operations performed by each unit in the computer device are the same as described above. Figure 2 The embodiments and their various alternative implementations are similar to those described in the embodiments, and will not be repeated here.

[0136] The computer device in the embodiments of this application is described below. Please refer to [link / reference]. Figure 8One embodiment of the computer device in this application includes:

[0137] The computer device 800 may include one or more central processing units (CPUs) 801 and a memory 805, in which one or more applications or data are stored.

[0138] The memory 805 can be volatile or persistent storage. The program stored in the memory 805 can include one or more modules, each module including a series of instruction operations on the computer device. Furthermore, the central processing unit 801 can be configured to communicate with the memory 805 and execute the series of instruction operations in the memory 805 on the computer device 800.

[0139] The computer device 800 may also include one or more power supplies 802, one or more wired or wireless network interfaces 803, one or more input / output interfaces 804, and / or one or more operating systems, such as Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™, etc.

[0140] The central processing unit 801 can perform the aforementioned... Figure 2 The operations performed by the computer device in the embodiments and their various alternative embodiments are not detailed here.

[0141] This application also provides a computer storage medium, one embodiment of which includes: the computer storage medium storing instructions, which, when executed on a computer, cause the computer to perform the aforementioned... Figure 2 The operations performed by the computer device in the embodiments and their various alternative embodiments.

[0142] This application also provides a computer program product, one embodiment of which includes: when the computer program product is run on a computer device, causing the computer device to perform the aforementioned... Figure 2 The operations performed by the computer device in the embodiments and their various alternative embodiments.

[0143] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0144] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection between apparatuses or units through some interfaces, and may be electrical, mechanical, or other forms.

[0145] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0146] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0147] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

Claims

1. An attribution analysis method characterized by, The method comprises: acquiring a plurality of groups of business data sets, each group of the business data sets comprising index data of a business index and influence factor data of a plurality of influence factors corresponding to the business index; for each alternative attribution model, inputting the plurality of groups of business data sets into the alternative attribution model, so that the alternative attribution model sequentially performs attribution analysis of the index data and the influence factor data based on each hyperparameter combination, for the business index and the plurality of influence factors thereof; determining an optimal hyperparameter combination of the alternative attribution model according to model performance exhibited by the alternative attribution model based on each of the hyperparameter combinations in learning the business data sets; determining a target attribution model with optimal performance in a plurality of the alternative attribution models according to model performance exhibited by each of the alternative attribution models based on the optimal hyperparameter combination thereof; the target attribution model is used to calculate the influence degree of a plurality of influence factors of a to-be-analyzed business index on the to-be-analyzed business index.

2. The method of claim 1, wherein, The inputting of the plurality of groups of business data sets into the alternative attribution model comprises: determining a to-be-evaluated hyperparameter combination, inputting the plurality of groups of business data sets into the alternative attribution model, so that the alternative attribution model performs attribution analysis of the index data and the influence factor data based on the to-be-evaluated hyperparameter combination, for the business index and the plurality of influence factors thereof; determining a next to-be-evaluated hyperparameter combination according to model performance exhibited by the alternative attribution model based on the to-be-evaluated hyperparameter combination in learning the business data sets, and returning to perform the inputting of the plurality of groups of business data sets into the alternative attribution model until a convergence condition is met to stop.

3. The method of claim 2, wherein, The determining of the next to-be-evaluated hyperparameter combination according to the model performance exhibited by the alternative attribution model based on the to-be-evaluated hyperparameter combination in learning the business data sets comprises: calculating a model performance indicator of the alternative attribution model on the to-be-evaluated hyperparameter combination according to a prediction result output by the alternative attribution model in learning the business data sets, to obtain the model performance indicator corresponding to an evaluated hyperparameter combination; constructing a probability model, inputting the model performance indicator corresponding to the evaluated hyperparameter combination into the probability model, so that the probability model predicts a model performance indicator corresponding to an unevaluated hyperparameter combination and a confidence degree of the model performance indicator according to the model performance indicator corresponding to the evaluated hyperparameter combination; constructing a collection function, calculating an expected improvement of an unevaluated hyperparameter combination according to the model performance indicator of each unevaluated hyperparameter combination and the confidence degree thereof through the collection function; taking the unevaluated hyperparameter combination with the highest expected improvement as the next to-be-evaluated hyperparameter combination.

4. The method of claim 1, wherein, The determining of the optimal hyperparameter combination of the alternative attribution model according to the model performance exhibited by the alternative attribution model based on each of the hyperparameter combinations in learning the business data sets comprises: For each of the hyperparameter combinations, a model performance indicator of the candidate attribution model on the hyperparameter combination is calculated according to a prediction result output by the candidate attribution model based on learning of the business data set according to the hyperparameter combination; Among the plurality of hyperparameter combinations, a hyperparameter combination corresponding to the model performance indicator optimal among the model performance indicators is determined as an optimal hyperparameter combination of the candidate attribution model.

5. The method of claim 1, wherein, The determining, among the plurality of candidate attribution models, a target attribution model with optimal performance according to a model performance of each of the candidate attribution models based on the optimal hyperparameter combination of the candidate attribution model, comprises: For each of the candidate attribution models, a model performance indicator of the candidate attribution model on the optimal hyperparameter combination of the candidate attribution model is calculated according to a prediction result output by the candidate attribution model based on learning of the business data set according to the optimal hyperparameter combination of the candidate attribution model; Among the plurality of candidate attribution models, a candidate attribution model corresponding to the model performance indicator optimal among the model performance indicators is determined as the target attribution model.

6. The method of claim 1, wherein, The obtaining of the plurality of groups of business data sets comprises: Obtaining a plurality of groups of original business data sets, each of the groups of original business data sets comprising index data of a business index and influence factor data of a plurality of influence factors corresponding to the business index; Preprocessing the index data and / or the influence factor data in each of the groups of original business data sets to obtain the plurality of groups of business data sets, wherein the preprocessing comprises: For a value of a missing variable of a business record in the index data and / or the influence factor data, supplementing a mean value of the variable in other business records to the business record, and / or, eliminating the business record; and / or, For a variable in the index data and / or the influence factor data, performing normalization calculation on each value of the variable according to a mean value and a standard deviation of a plurality of values of the variable.

7. The method according to any one of claims 1 to 6, characterized in that, The method further comprises: Obtaining to-be-processed business data, the to-be-processed business data comprising index data of a to-be-analyzed business index and influence factor data of a plurality of influence factors corresponding to the to-be-analyzed business index; Inputting the to-be-processed business data into the target attribution model, so that the target attribution model calculates, according to the index data and the influence factor data in the to-be-processed business data, influence degrees of the plurality of influence factors on the to-be-analyzed business index; Displaying the influence degrees of the plurality of influence factors on the to-be-analyzed business index output by the target attribution model, and / or displaying business index improvement suggestions based on the influence factors and the influence degrees.

8. A computer device, comprising: The computer device comprises: An obtaining unit, configured to obtain a plurality of groups of business data sets, each of the groups of business data sets comprising index data of a business index and influence factor data of a plurality of influence factors corresponding to the business index; A training unit, configured to, for each of the candidate attribution models, input the plurality of groups of business data sets into the candidate attribution model, so that the candidate attribution model performs, according to each of the hyperparameter combinations, attribution analysis of the index data and the influence factor data on the business index and the plurality of influence factors. A first determining unit is configured to determine an optimal hyperparameter combination of each of the candidate attribution models according to model performances of the candidate attribution models based on each of the hyperparameter combinations. A second determining unit is configured to determine a target attribution model with optimal performance from the candidate attribution models according to model performances of the candidate attribution models based on the optimal hyperparameter combinations of the candidate attribution models. The target attribution model is used to calculate influence degrees of a plurality of influence factors of a business index to be analyzed on the business index to be analyzed. 9.A computer device, comprising a memory and a processor, wherein the memory stores a computer program, and the computer device is configured to perform the method according to any one of claims 1-8 when the computer program is executed by the processor. The processor implements the method in any one of claims 1 to 7 when executing the computer program.

10. A computer storage medium, characterized in that, The computer storage medium stores instructions, and the instructions cause the computer to execute the method in any one of claims 1 to 7 when executed on the computer.

11. A computer program product, characterised in that, The computer program product causes the computer device to execute the method in any one of claims 1 to 7 when running on the computer device.