Business data attribution analysis method and system based on artificial intelligence
By introducing artificial intelligence technologies such as graph neural network, Transformer, PageRank and GBDT in attribution analysis, the traditional methods have solved the shortcomings in high-dimensional data analysis and rapid adaptability, and achieved more accurate and flexible attribution analysis and generated detailed business decision support reports.
Patent Information
- Application Number
- CN202510170376.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-17
- Publication Date
- 2025-06-06
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Traditional attribution analysis methods seem unscrupulous when processing high-dimensional data, and it is difficult to quickly adapt to data structure changes or business logic updates. The analysis results are limited to single-dimensional or fixed-dimensional combinations, making it difficult to generate hierarchical analysis conclusions.
Using an artificial intelligence-based method, the interactive relationship between dimensions is modeled through graph neural networks to form a dimensional interaction network, and the self-attention mechanism of Transformer dynamically adjusts the interaction weight between dimensions, combines the PageRank algorithm to evaluate the influence and propagation path of dimensions, and uses the GBDT model to predict the contribution of each dimension to the business results, and finally generates a detailed attribution report through the large language AI model.
It realizes comprehensive and accurate analysis of high-dimensional data, improves the adaptability and flexibility of the system, can more accurately reflect the importance of dimensions in different business scenarios, quantifies the contribution of each dimension to business results, and generates a detailed attribution report to provide comprehensive and in-depth information support for business decisions.
Smart Images

Figure CN120106658A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of business data attribution analysis, and in particular to a business data attribution analysis method and system based on artificial intelligence. Background Art
[0002] With the rapid development of big data and artificial intelligence technologies, the reliance on data-driven decision-making in business management is increasing. In actual business, various influencing factors interweave and cause fluctuations in key indicators (such as sales, profit margins, etc.). Traditional analysis methods are difficult to quickly and accurately locate the core causes of these fluctuations. Therefore, the market urgently needs an analysis tool that can efficiently attribute and output intuitive insights to meet the urgent needs of enterprises in data-driven decision-making.
[0003] At present, common attribution analysis methods usually rely on manually constructed analysis frameworks or traditional statistical models (such as regression analysis, variance analysis, etc.). Although these methods can help companies reveal the factors that affect indicator changes to a certain extent, they still have many shortcomings in practical applications:
[0004] 1. Dimension limitation: In actual business data, there is often a large amount of dimensional information. When the data dimensions are large, traditional attribution analysis methods are unable to cope with high-dimensional data and it is difficult to comprehensively and effectively analyze the contribution of all dimensional combinations to business results.
[0005] 2. Poor dynamic adaptability: For complex business scenarios, traditional attribution methods are usually based on fixed models or frameworks, which are difficult to quickly adapt to changes in data structure or updates to business logic. When facing new business scenarios or data changes, enterprises need to rebuild the analysis framework or adjust model parameters, which increases the cost and complexity of analysis.
[0006] 3. Limitations of analysis results: Traditional attribution analysis tools focus more on a single dimension or a fixed combination of dimensions, and it is difficult to flexibly generate hierarchical analysis conclusions, such as the contribution and correlation of each dimension. Summary of the invention
[0007] Based on this, the purpose of the present invention is to propose a business data attribution analysis method and system based on artificial intelligence to solve the above-mentioned problems.
[0008] According to an artificial intelligence-based business data attribution analysis method proposed in the present invention, the method includes:
[0009] Obtain business data and perform preprocessing;
[0010] Identify key attribution dimensions associated with business outcomes;
[0011] Based on the business data, the interactive relationship between dimensions is modeled through a graph neural network to form a dimensional interaction network;
[0012] The self-attention mechanism of Transformer is used to dynamically adjust the interaction weights between dimensions in the dimensional interaction network;
[0013] The PageRank algorithm is used to evaluate the influence and propagation path of the dimension in the dimension interaction network;
[0014] Use the GBDT model to predict and evaluate the contribution of each dimension to business results and rank them;
[0015] Based on the business data, dimensional interaction network, dimensional influence and propagation path, as well as the contribution of each dimension to business results and its ranking, the big language AI model is called to generate an attribution report leading to business results.
[0016] Furthermore, the step of modeling the interactive relationship between dimensions through a graph neural network based on the business data to form a dimension interaction network includes:
[0017] Convert the business data into a node feature matrix and an edge list;
[0018] Define the graph structure, taking each dimension as a node in the graph and the interaction between dimensions as edges;
[0019] Initialize the feature vector of each node v according to the business data;
[0020] In the graph convolution layer, the representation of the current node is updated by aggregating the information of neighboring nodes. The update formula is:
[0021]
[0022] Among them, H (l+1) is the node representation matrix of the l+1th layer, A is the adjacency matrix of the graph, D is the degree matrix, and W (l) is the weight matrix of the lth layer, σ is the activation function;
[0023] Extracting a representation vector of each node as a dimension vector, wherein the dimension vector contains interaction information between dimensions;
[0024] Based on the dimension vectors, a dimension interaction network is constructed to visualize the complex connections between dimensions.
[0025] Furthermore, the step of dynamically adjusting the interaction weights between dimensions in the dimensional interaction network using the self-attention mechanism of the Transformer includes:
[0026] Arrange each dimension vector into a sequence as the input sequence;
[0027] For each dimensional vector in the input sequence, a corresponding query vector, a key vector and a value vector are obtained by linear transformation;
[0028] The attention weight is obtained by calculating the dot product between the query vector and the key vector and normalizing it using the softmax function;
[0029] The attention weight is multiplied by the value vector and weighted summed to obtain a new dimensional representation.
[0030] Furthermore, the step of evaluating the influence and propagation path of a dimension in the dimension interaction network by using the PageRank algorithm includes:
[0031] Assign an initial PageRank value to each node in the dimensional interaction network;
[0032] Apply the PageRank formula to calculate the PageRank value of each node. The calculation formula is:
[0033]
[0034] Where PR(X) is the PageRank value of node X, PR(Y) is the PageRank value of node Y, d is the damping factor, N is the total number of nodes, B(X) is the set of all nodes linked to node X, L(Y) is the number of outbound links of node Y, Indicates the PageRank value passed through the link;
[0035] Repeat the PageRank formula for iterative calculation until the PageRank values of all nodes converge;
[0036] According to the PageRank value of each node, find the nodes with higher influence in the dimensional interaction network;
[0037] The distribution of edges and PageRank values between nodes is analyzed to identify the paths by which influence spreads in the network.
[0038] Furthermore, the steps of using the GBDT model to predict and evaluate the contribution of each dimension to the business results include:
[0039] Initialize an importance score for each dimension;
[0040] In the process of building each decision tree of the GBDT model, the reduction of the Gini index when each dimension is used as a split point is calculated, and the dimension with the largest reduction in the Gini index is selected as the split point of the current node. The calculation formula of the Gini index of dimension X is:
[0041]
[0042] Among them, G(X) is the Gini index of dimension X, k X is the number of different values of dimension X, c i is the number of samples when dimension X takes the i-th value, K is the total number of samples, is the sample proportion when dimension X takes the i-th value. The reduction of Gini index is obtained by calculating the difference between Gini index before and after splitting.
[0043] After each decision tree of the GBDT model is trained, the importance score of each dimension is updated according to the reduction of the Gini index, the influence of the dimension in the dimension interaction network, and the interaction weight between dimensions. The update formula is:
[0044] F t (X) = F t-1 (X)+α·ΔG t (X)+β·PR(X)+γ·∑W(X,Y),
[0045] Among them, F t (X) is the importance score of dimension X after the t-th decision tree is trained, F t-1 (X) is the importance score of dimension X after the training of the t-1th decision tree, α, β, γ are weight coefficients, which are used to adjust the reduction of the Gini index, the influence of the dimension in the dimension interaction network, and the contribution of the interaction weight between dimensions in the importance score of the dimension, W(X, Y) is the interaction weight factor, which represents the interaction weight between dimension X and dimension Y;
[0046] After all decision trees are trained, the final dimension importance scores are normalized to obtain the contribution of each dimension to the business results.
[0047] Furthermore, the step of calling the big language AI model to generate an attribution report leading to business results based on the business data, the dimension interaction network, the influence and propagation path of the dimension, and the contribution of each dimension to the business results and the ranking thereof includes:
[0048] Convert the key information of business data, dimension interaction network information, dimension influence and propagation path information, contribution of each dimension to business results and its ranking information into JSON or XML format, and submit them through the API interface provided by the big language AI model. The key information of business data includes key indicators, time range and data source of business data, and the dimension interaction network information includes adjacency matrix and dimension vector.
[0049] The interactive relationship between dimensions and the contribution of dimensions to business results are intuitively displayed in the form of charts;
[0050] Design the logical structure of the attribution report, including introduction, methodology, key findings, detailed analysis, conclusions and recommendations;
[0051] Write detailed prompts and submit them through the API interface provided by the Big Language AI model to describe the introduction, methodology, key findings, detailed analysis, conclusions and recommendations of the attribution report to the Big Language AI model.
[0052] The present invention also proposes an artificial intelligence-based business data attribution analysis system, which is used to implement the above-mentioned artificial intelligence-based business data attribution analysis method, and the system includes:
[0053] Data acquisition module: used to acquire business data and perform preprocessing;
[0054] Dimension determination module: used to determine the key attribution dimensions related to business results;
[0055] Interaction network module: used to model the interaction relationship between dimensions through graph neural network based on the business data to form a dimensional interaction network;
[0056] Interaction adjustment module: used to dynamically adjust the interaction weights between dimensions in the dimensional interaction network using the self-attention mechanism of Transformer;
[0057] Influence module: used to evaluate the influence and propagation path of a dimension in the dimension interaction network through the PageRank algorithm;
[0058] Contribution evaluation module: used to predict and evaluate the contribution of each dimension to business results using the GBDT model;
[0059] Attribution reporting module: used to call the big language AI model to generate an attribution report leading to business results based on the business data, dimension interaction network, dimension influence and propagation path, and the contribution of each dimension to the business results and its ranking.
[0060] In summary, the business data attribution analysis method based on artificial intelligence of the present invention models the interaction relationship between dimensions through graph neural network to form a dimensional interaction network to capture the complex connections and potential patterns between dimensions; and dynamically adjusts the interaction weights between dimensions through the self-attention mechanism of Transformer, so that the interaction weights between dimensions can be automatically adjusted according to the data characteristics in different business scenarios, so as to more accurately reflect the importance of dimensions in different business scenarios, thereby improving the adaptability and flexibility of the system; and then evaluates the influence and propagation path of the dimension in the dimensional interaction network through the PageRank algorithm, so that the system understands the dependency relationship and influence propagation mode between dimensions; then, the GBDT model is used to predict and evaluate the importance score of each dimension, thereby quantifying the contribution of each dimension to the business results, making the contribution evaluation more accurate and reliable; based on business data, dimensional interaction network, influence and propagation path of dimensions, and contribution of each dimension to business results and their ranking, the large language AI model is called to generate reports, so as to generate detailed attribution reports leading to business results with the help of the deep understanding ability of the large language model, providing comprehensive and in-depth information support for business decisions.
[0061] Additional aspects and advantages of the present invention will be given in part in the following description and in part will be obvious from the following description or will be learned through embodiments of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0062] The above and / or additional aspects and advantages of the present invention will become apparent and easily understood from the description of the embodiments in conjunction with the following drawings, in which:
[0063] Figure 1 This is a flow chart of a business data attribution analysis method based on artificial intelligence according to the first embodiment of the present invention;
[0064] Figure 2 This is a system block diagram of an artificial intelligence-based business data attribution analysis system according to Embodiment 2 of the present invention. DETAILED DESCRIPTION
[0065] In order to facilitate the understanding of the present invention, the present invention will be described more fully below with reference to the relevant drawings. Several embodiments of the present invention are given in the drawings. However, the present invention can be implemented in many different forms and is not limited to the embodiments described herein. On the contrary, the purpose of providing these embodiments is to make the disclosure of the present invention more thorough and comprehensive.
[0066] It should be noted that when an element is referred to as being "fixed to" another element, it may be directly on the other element or there may be a central element. When an element is considered to be "connected to" another element, it may be directly connected to the other element or there may be a central element at the same time. The terms "vertical", "horizontal", "left", "right" and similar expressions used herein are for illustrative purposes only.
[0067] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art of the present invention. The terms used herein in the specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. The term "and / or" used herein includes any and all combinations of one or more related listed items.
[0068] Embodiment 1
[0069] See also Figure 1 , the present invention proposes a business data attribution analysis method based on artificial intelligence, the method comprising steps S101 to S107:
[0070] S101, obtain business data and perform preprocessing.
[0071] It should be noted that to obtain business data, taking e-commerce as an example, it is necessary to obtain sales data, user behavior data, product information, market feedback, etc. Sales data, user registration information, order data, etc. can be obtained through the company's CRM system, ERP system, etc. You can also use the API interface provided by the e-commerce platform to obtain external data such as product information, user comments, sales rankings, etc.
[0072] The acquired data is preprocessed, including data cleaning (such as removing outliers, missing values, etc.) and data transformation (such as normalization, standardization, etc.).
[0073] S102, determine the key attribution dimensions related to business results.
[0074] It should be noted that, according to business needs and goals, some dimensions that have a relatively large impact on business results are determined as key attribution dimensions, which will be the focus of subsequent analysis and modeling, such as time (such as season, month, date, etc.), region (such as country, city, region, etc.), user behavior (such as browsing, clicking, purchasing, etc.), product characteristics (such as price, quality, brand, etc.), etc.
[0075] S103, based on the business data, modeling the interactive relationship between dimensions through a graph neural network to form a dimensional interaction network.
[0076] It should be noted that by modeling the interactive relationship between dimensions through graph neural networks, a dimension interaction network is formed, which can capture the complex connections between different dimensions in business data, including nonlinear relationships and high-order interactions. Graph neural networks can flexibly handle different types of business data and dimensions and have strong adaptability. By building a dimension interaction network, the connection between dimensions can be intuitively displayed, so that the interaction between dimensions can be more deeply understood, providing a more accurate and in-depth information basis for business data analysis and decision-making in subsequent attribution reports.
[0077] Further optionally, the step of modeling the interactive relationship between dimensions through a graph neural network based on the business data to form a dimensional interaction network includes:
[0078] Convert the business data into a node feature matrix and an edge list;
[0079] Define the graph structure, taking each dimension as a node in the graph and the interaction between dimensions as edges;
[0080] Initialize the feature vector of each node v according to the business data;
[0081] In the graph convolution layer, the representation of the current node is updated by aggregating the information of neighboring nodes. The update formula is:
[0082]
[0083] Among them, H (l+1) is the node representation matrix of the l+1th layer, A is the adjacency matrix of the graph, D is the degree matrix, and W (l) is the weight matrix of the lth layer, σ is the activation function;
[0084] Extracting a representation vector of each node as a dimension vector, wherein the dimension vector contains interaction information between dimensions;
[0085] Based on the dimension vectors, a dimension interaction network is constructed to visualize the complex connections between dimensions.
[0086] Understandably, the business data is converted into a node feature matrix and an edge list, where the node feature matrix describes the features of each dimension and the edge list defines the interaction between dimensions. The graph structure is defined, with each dimension as a node in the graph and the interaction between dimensions as edges.
[0087] Based on the business data, the feature vectors of each node v are initialized, and these feature vectors will be used as the input of the graph convolution layer. In the graph convolution layer, the representation of the current node is updated by aggregating the information of neighboring nodes. This process can capture the local structural information between nodes, and through multiple layers of graph convolution, more global structural information can be gradually captured. Then the representation vector of each node is extracted as a dimension vector, which contains the interaction information between dimensions. And based on the dimension vector, a dimension interaction network is constructed to visualize the complex connections between dimensions and help understand the dimensional interactions in business data.
[0088] S104, using the self-attention mechanism of Transformer to dynamically adjust the interaction weights between dimensions in the dimensional interaction network.
[0089] It should be noted that the Transformer model's self-attention mechanism is used to dynamically adjust the interaction weights between dimensions in the dimension interaction network to reflect the importance of dimensions in different business scenarios, thereby adapting to the needs of different business scenarios. This dynamic adjustment capability makes the model more flexible and adaptable. Dynamically adjusting the interaction weights between dimensions also helps the model to more accurately predict business results, thereby improving the performance of the model.
[0090] Further optionally, the step of dynamically adjusting the interaction weights between dimensions in the dimensional interaction network by using the self-attention mechanism of the Transformer includes:
[0091] Arrange each dimension vector into a sequence as the input sequence;
[0092] For each dimensional vector in the input sequence, a corresponding query vector, a key vector and a value vector are obtained by linear transformation;
[0093] By calculating the dot product between the query vector and the key vector and normalizing it using the softmax function, we get the attention weight, i.e. the interaction weight.
[0094] The attention weight is multiplied by the value vector and weighted summed to obtain a new dimensional representation.
[0095] It is understandable that each dimension vector is first arranged into a sequence as the input sequence of the Transformer model. These dimension vectors are obtained through graph neural network modeling and contain interaction information between dimensions. For each dimension vector in the input sequence, the corresponding query vector, key vector and value vector are obtained through linear transformation (i.e. multiplication by the weight matrix). Then, by calculating the dot product between the query vector and the key vector and normalizing it using the softmax function, the attention weight, also known as the interaction weight, is obtained, which reflects the interaction intensity between different dimensions, that is, which interactions between dimensions are more important in the current business scenario. By calculating the attention weight, the model can capture the complex relationship between dimensions, including nonlinear relationships and high-order interactions. Then, the attention weight is multiplied by the value vector and weighted summed to obtain a new dimension representation. The new dimension representation obtained is a dynamic adjustment of the interaction between dimensions, so that the model can more accurately capture the importance of dimensions in different business scenarios.
[0096] S105, evaluating the influence and propagation path of the dimension in the dimension interaction network through the PageRank algorithm.
[0097] It should be noted that the PageRank algorithm is used to evaluate the influence and propagation path of dimensions in the dimension interaction network based on the dimensional interaction network modeled by the graph neural network. The PageRank algorithm can quantify the influence of dimensions in the interaction network, providing an important information basis for the subsequent dimension contribution prediction and attribution report generation. By tracking the transmission process of PageRank values, we can clearly see which dimensions have a significant impact on other dimensions and how this impact spreads in the network, which is of great significance for understanding the causal relationship behind business results. We can also deeply explore the potential relationships in business data, which may involve complex interactions between multiple dimensions and their joint impact on business results. Incorporating this information into the attribution report can greatly enhance the depth and breadth of the report and provide more comprehensive information support for decision makers.
[0098] Further optionally, the step of evaluating the influence and propagation path of a dimension in the dimension interaction network by using the PageRank algorithm includes:
[0099] Assign an initial PageRank value to each node in the dimensional interaction network;
[0100] Apply the PageRank formula to calculate the PageRank value of each node. The calculation formula is:
[0101]
[0102] Where PR(X) is the PageRank value of node X, PR(Y) is the PageRank value of node Y, d is the damping factor, N is the total number of nodes, B(X) is the set of all nodes linked to node X, L(Y) is the number of outbound links of node Y, Indicates the PageRank value passed through the link;
[0103] Repeat the PageRank formula for iterative calculation until the PageRank values of all nodes converge;
[0104] According to the PageRank value of each node, find the nodes with higher influence in the dimensional interaction network;
[0105] The distribution of edges and PageRank values between nodes is analyzed to identify the paths by which influence spreads in the network.
[0106] Understandably, in this embodiment, the ageRank algorithm is cleverly applied to the dimension interaction network to evaluate the importance of each dimension in the business data. First, an initial PageRank value is assigned to each node (i.e., dimension) in the dimension interaction network, which can be assigned based on the initial importance of the node, or the PageRank values of all nodes are initialized to the same value (such as 1 / N, where N is the total number of nodes). The PageRank formula is then applied to calculate the PageRank value of each node, which reflects the influence of the node in the network. The formula takes into account the damping factor d (usually 0.85) and the process of transferring influence between nodes through links. Specifically, the PageRank value received by node X from other nodes Y is the PageRank value PR (Y) of node Y divided by the number of outbound links L (Y) of Y, and then multiplied by the damping factor d. Repeat the application of the PageRank formula for iterative calculation until the PageRank values of all nodes converge (i.e., the change is very small). After each iteration, the PageRank values of all nodes need to be normalized to ensure that their sum is always 1.
[0107] According to the size of the PageRank value of each node, nodes with high influence in the dimensional interaction network can be identified. These nodes are nodes with higher PageRank values, which are usually nodes that other nodes rely on more or that spread information faster in the network.
[0108] By analyzing the distribution of the edges and PageRank values between nodes, we can identify the paths of influence propagation in the network, that is, find the paths starting from high-influence nodes, passing through multiple nodes and finally reaching low-influence nodes. These paths reflect the propagation process of influence and help understand the dependencies between dimensions and the propagation mode of influence. The specific implementation steps can be as follows:
[0109] According to the edges in the dimensional interaction network, an adjacency matrix is constructed, where the elements in the matrix represent the connection relationship between nodes. If there is an edge connecting node X and node Y, the value of the corresponding position in the matrix is 1, otherwise it is 0; starting from the high-influence node, use the breadth-first search (BFS) or depth-first search (DFS) algorithm (but in the actual application of large-scale networks, some heuristic algorithms or approximate algorithms can be used to speed up the search process) to search for all possible paths in the network; among all the searched paths, select the paths that pass through multiple nodes and finally reach the low-influence node. These paths represent the propagation paths of influence in the network.
[0110] The path can also be evaluated, such as the length of the path, that is, the number of nodes passed by the path. A shorter path can indicate a higher efficiency in the spread of influence. If the edges between nodes have weights (for example, determined by the interaction strength or frequency between nodes), the weight of the path can be calculated. A path with a larger weight can indicate a higher importance in the spread of influence.
[0111] Based on the analysis results of the path propagation of dimensions in the interactive network, targeted strategies can be formulated, such as strengthening the monitoring and management of key nodes, optimizing the network structure to improve the efficiency of influence propagation, etc.
[0112] S106, use the GBDT model to predict and evaluate the contribution of each dimension to the business results, and rank them.
[0113] It should be noted that the GBDT (Gradient Boosting Decision Tree) model is used to predict and evaluate the importance score of each dimension, thereby reflecting the contribution of each dimension to the business results. This step plays a key role in the entire attribution analysis method because it directly quantifies the impact of different dimensions on business results. Based on the dimension sorting and contribution evaluation in the generated attribution report, decision makers can prioritize the dimensions that have a greater impact on business results, concentrate resources and energy on in-depth analysis and optimization, and thus improve decision-making efficiency.
[0114] Further optionally, the step of using the GBDT model to predict and evaluate the contribution of each dimension to the business result includes:
[0115] Initialize an importance score for each dimension, which can be set to 0 or other initial values;
[0116] In the process of building each decision tree of the GBDT model, the reduction of the Gini index when each dimension is used as a split point is calculated, and the dimension with the largest reduction in the Gini index is selected as the split point of the current node. The calculation formula of the Gini index of dimension X is:
[0117]
[0118] Among them, G(X) is the Gini index of dimension X, k X is the number of different values of dimension X, c i is the number of samples when dimension X takes the i-th value, K is the total number of samples, is the sample proportion when dimension X takes the i-th value. The reduction of Gini index is obtained by calculating the difference between Gini index before and after splitting.
[0119] After each decision tree of the GBDT model is trained, the importance score of each dimension is updated according to the reduction of the Gini index, the influence of the dimension in the dimension interaction network, and the interaction weight between dimensions. The update formula is:
[0120] F t (X) = F t-1 (X)+α·ΔG t (X)+β·PR(X)+γ·∑W(X,Y),
[0121] Among them, F t (X) is the importance score of dimension X after the t-th decision tree is trained, F t-1 (X) is the importance score of dimension X after the training of the t-1th decision tree, α, β, γ are weight coefficients used to adjust the reduction of the Gini index, the influence of the dimension in the dimension interaction network, and the contribution of the interaction weight between dimensions in the importance score of the dimension, W(X, Y) is the interaction weight factor, which represents the interaction weight between dimension X and dimension Y, that is, the interaction weight obtained by dynamic adjustment of the Transformer model in step S104;
[0122] After all decision trees are trained, the final dimension importance scores are normalized to obtain the contribution of each dimension to the business results, so that the total score is 1 (or a fixed value).
[0123] Understandably, the GBDT model is used to predict and evaluate the importance score of each dimension to reflect the contribution of each dimension to the business results. This step mainly relies on the decision tree in the GBDT model to calculate the reduction in the Gini index to evaluate the importance of the dimension. The reduction in the Gini index reflects the degree of reduction in the impurity of the data set before and after the split, and is an important indicator to measure the discrimination of the dimension in the data set. However, although some dimensions are not highly discriminative when viewed alone, they may have an important impact on business results when interacting with other dimensions. Therefore, fully and comprehensively considering the interactive relationship between dimensions and the influence of dimensions in the interactive network can more accurately evaluate the contribution of each dimension to business results.
[0124] The Gini index G(X) is used to measure the impurity of dimension X in a data set. The closer the value is to 0, the higher the purity of the data set. The range of the Gini index G(X) is [0, 1]. When G(X) is close to 0, it means that the distribution of the values of dimension X is very uneven, that is, the proportion of samples under one or some values is much higher than other values; when G(X) is close to 1, it means that the distribution of the values of dimension X is relatively uniform, and the proportion of samples under various values is not much different. However, in practical applications, we usually pay more attention to the situation where the Gini index is low, because it means that a certain value of dimension X has a good degree of discrimination for samples.
[0125] S107, based on the business data, dimension interaction network, dimension influence and propagation path, and the contribution of each dimension to the business results and their ranking, call the big language AI model to generate an attribution report leading to the business results.
[0126] It should be noted that the use of the deep analysis and natural language understanding capabilities of the big language AI model to automatically generate attribution reports can greatly improve the efficiency, accuracy and logic of report generation. The report integrates information such as business data, dimension interaction network, dimension influence and propagation path, and dimension contribution, providing comprehensive and in-depth information support for business decision-making.
[0127] Further optionally, based on the business data, the dimension interaction network, the influence and propagation path of the dimension, and the contribution of each dimension to the business result and its ranking, calling the big language AI model to generate an attribution report leading to the business result includes:
[0128] Convert the key information of business data, dimension interaction network information, dimension influence and propagation path information, contribution of each dimension to business results and its ranking information into JSON or XML format, and submit them through the API interface provided by the big language AI model. The key information of business data includes key indicators, time range and data source of business data, and the dimension interaction network information includes adjacency matrix and dimension vector.
[0129] The interactive relationship between dimensions and the contribution of dimensions to business results are intuitively displayed in the form of charts;
[0130] Design the logical structure of the attribution report, including introduction, methodology, key findings, detailed analysis, conclusions and recommendations;
[0131] Write detailed prompts and submit them through the API interface provided by the Big Language AI model to describe the introduction, methodology, key findings, detailed analysis, conclusions and recommendations of the attribution report to the Big Language AI model.
[0132] Understandably, first of all, the key information of business data, dimension interaction network information, dimension influence and propagation path information, contribution of each dimension to business results and its ranking are integrated. And this information is converted into JSON or XML format, which is compatible with the API interface of the big language AI model to ensure that the information can be accurately and correctly passed to the model.
[0133] The key information of business data includes key indicators of business data (such as sales, user growth rate, etc.), time range (such as the time period of analysis) and data source (such as the data source system or database). The dimension interaction network information includes the adjacency matrix (indicating the connection relationship between dimensions) and the dimension vector (indicating the characteristics or attributes of the dimension).
[0134] Among them, the interactive relationship between dimensions and the ranking of their contribution to business results can be intuitively displayed in the form of charts (such as bar charts, line charts, heat maps, network diagrams, etc.), making complex data and analysis results easier to understand.
[0135] Design a logical structure for the attribution report, including an introduction, methodology, key findings, detailed analysis, and conclusions and recommendations. This structure helps readers systematically understand the report content, from background to methods, to results and recommendations.
[0136] After determining the structure of the attribution report, write detailed prompts to describe the various parts of the attribution report to the Big Language AI model, such as the background introduction of the introduction, the analysis method of the methodology, the main conclusions of the key findings, the specific content of the detailed analysis, and the decision support of the conclusions and recommendations. Then submit these prompts through the API interface provided by the Big Language AI model, and the model generates the corresponding text content based on the prompts to form a complete attribution report.
[0137] Example prompt: Assume that you are a professional data analyst responsible for generating a detailed attribution report. This report needs to be generated based on the provided business data, dimension interaction network information, dimension influence and communication path, and the contribution of each dimension to the business results. The report content should include:
[0138] Introduction: Briefly describe the business background, purpose and importance of the analysis.
[0139] Methodology: An overview of the data, models, algorithms used and the rationale for their selection.
[0140] Key Findings: Summarizes the main findings of PageRank and GBDT models, highlighting the dimensions that have the greatest impact on business results.
[0141] Detailed analysis:
[0142] Dimensional impact analysis: Analyze the specific impact of each key dimension on business results, including both positive and negative impacts.
[0143] Discussion on interactive relationships: Based on the dimension interaction network, explore how the interactions between dimensions jointly affect business results.
[0144] Case study: Select or generate typical cases to conduct in-depth analysis of how specific dimensions or combinations of dimensions play a role in actual business scenarios.
[0145] Conclusion and suggestions: Based on the analysis results, put forward targeted business optimization suggestions, such as adjusting marketing strategies, improving product design, etc.
[0146] In summary, the business data attribution analysis method based on artificial intelligence of the present invention models the interaction relationship between dimensions through graph neural network to form a dimensional interaction network to capture the complex connections and potential patterns between dimensions; and dynamically adjusts the interaction weights between dimensions through the self-attention mechanism of Transformer, so that the interaction weights between dimensions can be automatically adjusted according to the data characteristics in different business scenarios, so as to more accurately reflect the importance of dimensions in different business scenarios, thereby improving the adaptability and flexibility of the system; and then evaluates the influence and propagation path of the dimension in the dimensional interaction network through the PageRank algorithm, so that the system understands the dependency relationship and influence propagation mode between dimensions; then, the GBDT model is used to predict and evaluate the importance score of each dimension, thereby quantifying the contribution of each dimension to the business results, making the contribution evaluation more accurate and reliable; based on business data, dimensional interaction network, influence and propagation path of dimensions, and contribution of each dimension to business results and their ranking, the large language AI model is called to generate reports, so as to generate detailed attribution reports leading to business results with the help of the deep understanding ability of the large language model, providing comprehensive and in-depth information support for business decisions.
[0147] Embodiment 2
[0148] See also Figure 2 The present invention proposes a business data attribution analysis system based on artificial intelligence, the system comprising:
[0149] Data acquisition module: used to acquire business data and perform preprocessing;
[0150] Dimension determination module: used to determine the key attribution dimensions related to business results;
[0151] Interaction network module: used to model the interaction relationship between dimensions through graph neural network based on the business data to form a dimensional interaction network;
[0152] Interaction adjustment module: used to dynamically adjust the interaction weights between dimensions in the dimensional interaction network using the self-attention mechanism of Transformer;
[0153] Influence module: used to evaluate the influence and propagation path of a dimension in the dimension interaction network through the PageRank algorithm;
[0154] Contribution evaluation module: used to predict and evaluate the contribution of each dimension to business results using the GBDT model;
[0155] Attribution reporting module: used to call the big language AI model to generate an attribution report leading to business results based on the business data, dimension interaction network, dimension influence and propagation path, and the contribution of each dimension to the business results and its ranking.
[0156] Further optionally, the interactive network module is also used for:
[0157] Convert the business data into a node feature matrix and an edge list;
[0158] Define the graph structure, taking each dimension as a node in the graph and the interaction between dimensions as edges;
[0159] Initialize the feature vector of each node v according to the business data;
[0160] In the graph convolution layer, the representation of the current node is updated by aggregating the information of neighboring nodes. The update formula is:
[0161]
[0162] Among them, H (l+1) is the node representation matrix of the l+1th layer, A is the adjacency matrix of the graph, D is the degree matrix, and W (l) is the weight matrix of the lth layer, σ is the activation function;
[0163] Extracting a representation vector of each node as a dimension vector, wherein the dimension vector contains interaction information between dimensions;
[0164] Based on the dimension vectors, a dimension interaction network is constructed to visualize the complex connections between dimensions.
[0165] Further optionally, the interaction adjustment module is further used for:
[0166] Arrange each dimension vector into a sequence as the input sequence;
[0167] For each dimensional vector in the input sequence, a corresponding query vector, a key vector and a value vector are obtained by linear transformation;
[0168] The attention weight is obtained by calculating the dot product between the query vector and the key vector and normalizing it using the softmax function;
[0169] The attention weight is multiplied by the value vector and weighted summed to obtain a new dimensional representation.
[0170] Further optionally, the influence module is also used for:
[0171] Assign an initial PageRank value to each node in the dimensional interaction network;
[0172] Apply the PageRank formula to calculate the PageRank value of each node. The calculation formula is:
[0173]
[0174] Where PR(X) is the PageRank value of node X, PR(Y) is the PageRank value of node Y, d is the damping factor, N is the total number of nodes, B(X) is the set of all nodes linked to node X, L(Y) is the number of outbound links of node Y, Indicates the PageRank value passed through the link;
[0175] Repeat the PageRank formula for iterative calculation until the PageRank values of all nodes converge;
[0176] According to the PageRank value of each node, find the nodes with higher influence in the dimensional interaction network;
[0177] The distribution of edges and PageRank values between nodes is analyzed to identify the paths by which influence spreads in the network.
[0178] Further optionally, the contribution assessment module is also used to:
[0179] Initialize an importance score for each dimension;
[0180] In the process of building each decision tree of the GBDT model, the reduction of the Gini index when each dimension is used as a split point is calculated, and the dimension with the largest reduction in the Gini index is selected as the split point of the current node. The calculation formula of the Gini index of dimension X is:
[0181]
[0182] Among them, G(X) is the Gini index of dimension X, k X is the number of different values of dimension X, c i is the number of samples when dimension X takes the i-th value, K is the total number of samples, is the sample proportion when dimension X takes the i-th value. The reduction of Gini index is obtained by calculating the difference between Gini index before and after splitting.
[0183] After each decision tree of the GBDT model is trained, the importance score of each dimension is updated according to the reduction of the Gini index, the influence of the dimension in the dimension interaction network, and the interaction weight between dimensions. The update formula is:
[0184] F t (X) = F t-1 (X)+α·ΔG t (X)+β·PR(X)+γ·∑W(X,Y),
[0185] Among them, F t (X) is the importance score of dimension X after the t-th decision tree is trained, F t-1 (X) is the importance score of dimension X after the training of the t-1th decision tree, α, β, γ are weight coefficients, which are used to adjust the reduction of the Gini index, the influence of the dimension in the dimension interaction network, and the contribution of the interaction weight between dimensions in the importance score of the dimension, W(X, Y) is the interaction weight factor, which represents the interaction weight between dimension X and dimension Y;
[0186] After all decision trees are trained, the final dimension importance scores are normalized to obtain the contribution of each dimension to the business results.
[0187] Further optionally, the attribution reporting module is further used to:
[0188] Convert the key information of business data, dimension interaction network information, dimension influence and propagation path information, contribution of each dimension to business results and its ranking information into JSON or XML format, and submit them through the API interface provided by the big language AI model. The key information of business data includes key indicators, time range and data source of business data, and the dimension interaction network information includes adjacency matrix and dimension vector.
[0189] The interactive relationship between dimensions and the contribution of dimensions to business results are intuitively displayed in the form of charts;
[0190] Design the logical structure of the attribution report, including introduction, methodology, key findings, detailed analysis, conclusions and recommendations;
[0191] Write detailed prompts and submit them through the API interface provided by the Big Language AI model to describe the introduction, methodology, key findings, detailed analysis, conclusions and recommendations of the attribution report to the Big Language AI model.
[0192] The above-mentioned embodiments only express several implementation methods of the present invention, and the description thereof is relatively specific and detailed, but it cannot be understood as limiting the scope of the patent of the present invention. It should be pointed out that, for ordinary technicians in this field, several variations and improvements can be made without departing from the concept of the present invention, which all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention shall be subject to the attached claims.
Claims
1. A business data attribution analysis method based on artificial intelligence, characterized in that: The method comprises: Obtain business data and perform preprocessing; Identify key attribution dimensions associated with business outcomes; Based on the business data, the interactive relationship between dimensions is modeled through a graph neural network to form a dimensional interaction network; The self-attention mechanism of Transformer is used to dynamically adjust the interaction weights between dimensions in the dimensional interaction network; The PageRank algorithm is used to evaluate the influence and propagation path of the dimension in the dimension interaction network; Use the GBDT model to predict and evaluate the contribution of each dimension to business results and rank them; Based on the business data, dimensional interaction network, dimensional influence and propagation path, as well as the contribution of each dimension to business results and its ranking, the big language AI model is called to generate an attribution report leading to business results.
2. The business data attribution analysis method based on artificial intelligence according to claim 1 is characterized in that: The step of modeling the interaction relationship between dimensions through a graph neural network based on the business data to form a dimension interaction network includes: Convert the business data into a node feature matrix and an edge list; Define the graph structure, taking each dimension as a node in the graph and the interaction between dimensions as edges; Initialize the feature vector of each node v according to the business data; In the graph convolution layer, the representation of the current node is updated by aggregating the information of neighboring nodes. The update formula is: Among them, H (l+1) is the node representation matrix of the l+1th layer, A is the adjacency matrix of the graph, D is the degree matrix, and W (l) is the weight matrix of the lth layer, σ is the activation function; Extracting a representation vector of each node as a dimension vector, wherein the dimension vector contains interaction information between dimensions; Based on the dimension vectors, a dimension interaction network is constructed to visualize the complex connections between dimensions.
3. The business data attribution analysis method based on artificial intelligence according to claim 2 is characterized in that: The step of dynamically adjusting the interaction weights between dimensions in the dimensional interaction network by using the Transformer self-attention mechanism includes: Arrange each dimension vector into a sequence as the input sequence; For each dimensional vector in the input sequence, a corresponding query vector, a key vector and a value vector are obtained by linear transformation; The attention weight is obtained by calculating the dot product between the query vector and the key vector and normalizing it using the softmax function; The attention weight is multiplied by the value vector and weighted summed to obtain a new dimensional representation.
4. The business data attribution analysis method based on artificial intelligence according to claim 3 is characterized in that: The step of evaluating the influence and propagation path of a dimension in the dimension interaction network by using the PageRank algorithm includes: Assign an initial PageRank value to each node in the dimensional interaction network; Apply the PageRank formula to calculate the PageRank value of each node. The calculation formula is: Where PR(X) is the PageRank value of node X, PR(Y) is the PageRank value of node Y, d is the damping factor, N is the total number of nodes, B(X) is the set of all nodes linked to node X, L(Y) is the number of outbound links of node Y, Indicates the PageRank value passed through the link; Repeat the PageRank formula for iterative calculation until the PageRank values of all nodes converge; According to the PageRank value of each node, find the nodes with higher influence in the dimensional interaction network; The distribution of edges and PageRank values between nodes is analyzed to identify the paths by which influence spreads in the network.
5. The business data attribution analysis method based on artificial intelligence according to claim 1 is characterized in that: The steps of using the GBDT model to predict and evaluate the contribution of each dimension to the business results include: Initialize an importance score for each dimension; In the process of building each decision tree of the GBDT model, the reduction of the Gini index when each dimension is used as a split point is calculated, and the dimension with the largest reduction in the Gini index is selected as the split point of the current node. The calculation formula of the Gini index of dimension X is: Among them, G(X) is the Gini index of dimension X, kX is the number of different values of dimension X, and c i is the number of samples when dimension X takes the i-th value, K is the total number of samples, is the sample proportion when dimension X takes the i-th value. The reduction of Gini index is obtained by calculating the difference between Gini index before and after splitting. After each decision tree of the GBDT model is trained, the importance score of each dimension is updated according to the reduction of the Gini index, the influence of the dimension in the dimension interaction network, and the interaction weight between dimensions. The update formula is: F t (X)=F t-1 (X)+α·ΔG t (X)+β·PR(X)+γ·∑W(X,Y), Among them, F t (X) is the importance score of dimension X after the t-th decision tree is trained, F t-1 (X) is the importance score of dimension X after the training of the t-1th decision tree, α, β, γ are weight coefficients, which are used to adjust the reduction of the Gini index, the influence of the dimension in the dimension interaction network, and the contribution of the interaction weight between dimensions in the importance score of the dimension, W(X, Y) is the interaction weight factor, which represents the interaction weight between dimension X and dimension Y; After all decision trees are trained, the final dimension importance scores are normalized to obtain the contribution of each dimension to the business results.
6. The business data attribution analysis method based on artificial intelligence according to claim 1 is characterized in that: The step of calling the big language AI model to generate an attribution report leading to the business results based on the business data, the dimension interaction network, the influence and propagation path of the dimension, and the contribution of each dimension to the business results and the ranking thereof includes: Convert the key information of business data, dimension interaction network information, dimension influence and propagation path information, contribution of each dimension to business results and its ranking information into JSON or XML format, and submit them through the API interface provided by the big language AI model. The key information of business data includes key indicators, time range and data source of business data, and the dimension interaction network information includes adjacency matrix and dimension vector. The interactive relationship between dimensions and the contribution of dimensions to business results are intuitively displayed in the form of charts; Design the logical structure of the attribution report, including introduction, methodology, key findings, detailed analysis, conclusions and recommendations; Write detailed prompts and submit them through the API interface provided by the Big Language AI model to describe the introduction, methodology, key findings, detailed analysis, conclusions and recommendations of the attribution report to the Big Language AI model.
7. A business data attribution analysis system based on artificial intelligence, used to implement the business data attribution analysis method based on artificial intelligence according to any one of claims 1 to 6, characterized in that: The system comprises: Data acquisition module: used to acquire business data and perform preprocessing; Dimension determination module: used to determine the key attribution dimensions related to business results; Interaction network module: used to model the interaction relationship between dimensions through graph neural network based on the business data to form a dimensional interaction network; Interaction adjustment module: used to dynamically adjust the interaction weights between dimensions in the dimensional interaction network using the self-attention mechanism of Transformer; Influence module: used to evaluate the influence and propagation path of a dimension in the dimension interaction network through the PageRank algorithm; Contribution evaluation module: used to predict and evaluate the contribution of each dimension to business results using the GBDT model; Attribution reporting module: used to call the big language AI model to generate an attribution report leading to business results based on the business data, dimension interaction network, dimension influence and propagation path, and the contribution of each dimension to the business results and its ranking.
Citation Information
Cited By
Attribution analysis method based on BI platform, medium and equipment
CN120725103A